Overview
10 stories in this issue. The first 3 are today's priorities.
Popular model watch
- Top · Pokee AI ships Isaac 28B with a 10M-token context on a single GPU
- Top · GPT Image 2 narrowly beats Grok Imagine on quality
- Top · DeepSeek Janus-Pro loses an image shoot-out to Fibo
Global AI news 4. Claude Code defaults to Auto Mode from August 14 5. OpenAI acquires NextSlide, pushing ChatGPT into slides 6. Shepherd: an open-source agent runtime you can fork, replay and revert 7. Backflip AI turns 3D scans into editable CAD in minutes 8. Readers can't tell AI short stories from human ones 9. Firebird opens the CIS region's largest AI factory in Armenia
Regional and early signals 10. Apple China briefly published, then pulled, a "Qwen on Mac" support doc (Chinese-language source)

Jiufeng graphic based on the sources cited in this issue.
Popular model watch
Pokee AI ships Isaac 28B with a 10M-token context on a single GPU
Pokee AI says its 28B Isaac model carries a 10-million-token context, runs on one GPU inside the customer boundary, and ships under a license rather than as open weights.
Pokee AI released Pokee-Isaac 28B, a text-only foundation model with a 10M-token context window built to run inside the customer boundary. The Pokee research team reports it stays above 93.3% on the long-context RULER benchmark at every tested length, ending at 93.3% at 10M tokens — on par with cost-optimized cloud models such as GPT-5.6 Luna and Gemini 3.5 Flash Lite — and scores 70.94 on the BFCL v4 function-calling benchmark against Luna's 70.61. Isaac is served through an OpenAI-compatible API, can be deployed in a VPC, on-prem or on-device, reportedly fits on a single GPU, and ships with day-0 vLLM and SGLang support.
Limitations: Isaac is distributed under a license, not as open weights, and is available only through the API and authorized deployments. The scores are Pokee's own; its report calls the slim BFCL v4 edge over Luna "parity rather than a lead."

Image source: GitHub; mirrored on Jiufeng R2.
Source: MarkTechPost · RULER · τ³-bench
GPT Image 2 narrowly beats Grok Imagine on quality
In a task-by-task RuntimeWire comparison, GPT Image 2 API edged Grok Imagine 70.1 to 67.0 on text and composition — but only just.
RuntimeWire ran GPT Image 2 API against Grok Imagine task by task: GPT Image 2 led on aggregate 70.1 to 67.0, winning 4 of 8 tasks with 2 ties, at 77% statistical confidence. It was steadier on attribute binding, composition, legible text, style fidelity (closer to a true ukiyo-e woodblock print) and spatial layout.
Limitations: the article calls it "a win on points, not a rout." Grok Imagine was better on hands and anatomy — a narrower but real strength. The scoring is one outlet's subjective, per-task judgment over a limited prompt set.
Source: RuntimeWire
DeepSeek Janus-Pro loses an image shoot-out to Fibo
In a second matchup, DeepSeek's Janus-Pro lost 45.1 to 61.2 against Fibo Bbq Preview, winning just 1 of 8 tasks.
In another RuntimeWire image comparison, Fibo Bbq Preview beat DeepSeek Janus-Pro 61.2 to 45.1, taking 7 of 8 tasks at 98% confidence. Fibo tracked the brief more closely on reflective materials and optics, soldered circuit-board macros, and action and design detail.
Limitations: the review says Janus-Pro "often nailed the vibe but broke the actual physics or dropped a key prompt element" — for instance turning a circuit-board macro into a keyboard shot. Again a single outlet's subjective test over only 8 prompts.
Source: RuntimeWire
Global AI news
Claude Code defaults to Auto Mode from August 14
From August 14 Anthropic will make Auto Mode the default in Claude Code for Pro, Max and Team plans, with a classifier catching dangerous actions; Enterprise still opts in manually.
Anthropic says that starting August 14, Claude Code will enable Auto Mode by default on Pro, Max and Team plans, with only Enterprise left to opt in. Auto Mode lets the coding agent proceed without approval at every step; a classifier flags dangerous or irreversible operations and asks for confirmation only then. In a controlled study of 1,053 paying testers, manual review caught just 13.6% of dangerous commands versus 89% for Auto Mode, and teams on Auto Mode shipped about 25% more PRs.
Limitations: Anthropic calls Auto Mode "at least as safe as manual approval, and often better," based on the 1,053 paying testers plus internal red-teaming; skipping step-by-step approval means the AI drives more of the workflow on its own.
Source: The Decoder
OpenAI acquires NextSlide, pushing ChatGPT into slides
OpenAI has bought NextSlide, an AI presentation tool founded barely a year ago; the team has joined OpenAI to work on ChatGPT, with terms undisclosed.
OpenAI acquired NextSlide, an AI presentation tool founded just over a year ago; founder Ahmed Beshry and his team have moved to OpenAI to work on ChatGPT, and the deal amount was not disclosed. According to The Information, Beshry previously co-founded the computer-vision smart-cart startup Caper, which Instacart bought in October 2021 for roughly $350 million in cash and stock.
Limitations: the report does not say which staff are joining or which technology will fold into ChatGPT, and the deal terms are undisclosed.
Source: RuntimeWire
Shepherd: an open-source agent runtime you can fork, replay and revert
Researchers at Northeastern and Stanford open-sourced Shepherd, which records an agent run as a Git-like event trace so any past state can be forked, replayed and reverted.
Researchers at Northeastern University and Stanford University released Shepherd, a Python runtime substrate that records an agent run as a Git-like trace of typed events, so any past state can be forked and replayed. The team reports forking is 5× faster than Docker, with over 95% prompt-cache reuse on replay. It targets a real pain point: long agent runs accumulate edited files, a live dev server, installed packages and a warm cache — state Git cannot version, because Git tracks files, not a running process or a cache.
Limitations: Shepherd is in early alpha, and the team says it is not yet production-ready.
Source: MarkTechPost · Shepherd (GitHub) · arXiv
Backflip AI turns 3D scans into editable CAD in minutes
Backflip AI released a second CAD model that converts 3D scans and mesh files into fully editable, parametric CAD models in minutes.
Backflip AI released its second CAD-generation model, which turns 3D scans and mesh files into fully editable, parametric CAD models — a task that normally takes significant time and expertise. Unlike earlier models that only output triangle meshes, it produces real CAD operations such as extrusions and revolves, so results are easy to modify. It runs as an Autodesk Fusion plugin, free for 4 reconstructions, with paid plans from $20/month. CEO Greg Mark says most factories have digital models for under 1% of their parts; the startup, operating since December 2024, has raised $30 million.
Limitations: the speed and quality figures are the company's own, and the free tier is capped at 4 reconstructions.
Source: The Decoder
Readers can't tell AI short stories from human ones
A study of more than 2,500 people found readers did no better than chance at spotting ChatGPT-written short stories — and rated the AI versions higher, until told a machine wrote them.
A study in Judgment and Decision Making found that more than 2,500 participants performed no better than chance at distinguishing ChatGPT-generated short stories from human-written ones. In the first experiment, 1,682 people each read a ~1,000-word story (3 from literary magazines and anthologies, 3 generated by ChatGPT 4.0) and were given an authorship label that was accurate for only half of them. Readers found the AI stories more engaging and higher quality — yet rated any story labeled "human-written" more highly. In two further experiments totaling 905 people, one round guessed the human author correctly only about 40% of the time (below chance) and another 52%.
Limitations: Villanova researchers Sydney Sears and Deena Skolnick Weisberg write there is "no general evidence" that readers can tell the two apart, though people more familiar with AI scored better; Weisberg stresses this does not mean writing should be handed to AI.
Source: The Decoder · IT Home
Firebird opens the CIS region's largest AI factory in Armenia
Firebird launched what it calls the CIS region's largest AI factory in Armenia, planning to deploy more than 70,000 NVIDIA Rubin and Blackwell GPUs at 300 megawatts.
Emerging AI cloud Firebird launched the CIS region's largest AI factory in Armenia, built on NVIDIA accelerated computing and Dell high-performance infrastructure. Firebird plans to deploy more than 70,000 NVIDIA Rubin and Blackwell GPUs at 300 megawatts. Armenian Prime Minister Nikol Pashinyan, Kazakhstan Deputy Prime Minister Zhaslan Madiyev and U.S. Chargé d'Affaires David Allen attended the launch. NVIDIA says an AI factory supplies the compute to train, fine-tune and deploy models, giving the country the ability to build AI for its own language, industries and priorities.
Limitations: the announcement comes from NVIDIA's own promotional blog; the 70,000 GPUs and 300 MW are Firebird's build-out plan, not capacity already in place.
Source: NVIDIA AI Blog
Regional and early signals
Apple China briefly published, then pulled, a "Qwen on Mac" support doc (Chinese-language source)
A support doc titled "Using Qwen with Apple Intelligence on Mac" appeared in Apple's Simplified Chinese Mac guide, then was removed, with the original link no longer reachable.
Per IT Home, on August 8 a support document titled "Using Qwen with Apple Intelligence on Mac" appeared in Apple's Simplified Chinese Mac user guide, explicitly stating that Apple Intelligence can work with Alibaba's Qwen model. The doc said the Qwen extension applies to macOS 26.6 or later for users attributed to mainland China by device, account or location, can be enabled under "Settings → Apple Intelligence & Siri — Extensions," and works with Writing Tools and Siri after signing in to a Qwen account. IT Home reports the Chinese document has since been removed from Apple's site and the original page URL no longer resolves.
Limitations: only a Chinese-language source is available; neither Apple nor Alibaba has commented on the posting or removal, the reason for the takedown is unknown, and any official rollout remains unconfirmed.
Source: IT Home
Get the latest AI model insights and tutorials from Jiufeng.
Explore more


