Overview
9 stories in this issue. The first 3 are today's priorities.
Hot model dynamics
- Top · Four Claude Code agents reach 62.1% on a coding benchmark
- Top · MiniMax H3 team AMA: a 2K-output model will ship, image model on the way
- Top · Google denies training Gemini on private docs after a leak claim
- Gemini Robotics 2 drives three robot bodies; Apollo 2 dexterity 32%–92%
- Alibaba's Qwen app adds five free features; Wan 3.0 enters public test
Global AI news 6. Fields Medalist joins OpenAI to work on AI safety 7. Airbnb says AI cut its feature shipping time by 60% 8. Cohere Health digitizes clinical policies on Amazon Bedrock AgentCore
Regional and early signals 9. China's AI-phone grading clears first batch; L3 is only the starting line

Jiufeng graphic based on the sources cited in this issue.
Hot model dynamics
Four Claude Code agents reach 62.1% on a coding benchmark
Coral AI Labs' AgentRadio let four Opus 4.6 agents exchange discoveries mid-run and solved 62.1% of a codebase-understanding benchmark — at roughly sixfold inference cost.
Coral AI Labs co-founders Caelum Forder and Peter Carroll published the AgentRadio paper on July 30: four Claude Code agents running Anthropic's Opus 4.6, able to exchange discoveries during execution, resolved 62.1% of a demanding 124-task codebase-understanding benchmark. In the same runs a single Opus 4.6 agent scored 32.3%, while a single Opus 4.8 agent scored 57.2% on Scale AI's public leaderboard. The design runs the mention watcher as a background OS process, so incoming messages appear between execution steps without stopping a command in progress.
Limitations: RuntimeWire notes the result is narrower than the framing that circulated after VentureBeat's report — AgentRadio did not show Opus 4.6 is generally better than Opus 4.8, nor that four-agent systems consistently beat newer models; it beat one Opus 4.8 configuration on these 124 tasks. Scale's current leaderboard already lists a single-agent setup above 62.1%, and the multi-agent approach carries about a sixfold inference-cost increase.

Image source: GitHub; mirrored on Jiufeng R2.
Source: RuntimeWire · VentureBeat · GitHub
MiniMax H3 team AMA: a 2K-output model will ship, image model on the way
Chinese-language source. In a Reddit AMA the MiniMax H3 team said the model behind online 2K output "will be released, and not too long," and an image model is on the way.
On the evening of August 7, the MiniMax H3 team held an AMA in Reddit's r/StableDiffusion; participants included H3 research lead dacongya, researchers Luigi/Nero/Kiro, systems engineer Reynor and DevRel lead Ryanlee, and the thread was marked "Finished" after 300-plus comments. H3, released July 31, is a full-modal video generation model that understands text, image, video and audio context and generates up to 15-second, up to 2K-resolution video with native two-channel audio. Asked whether the model used for final 2K output would ship, Ryanlee said "yes, and not too long"; the 2K approach, called In-context Regeneration, avoids a separate super-resolution stage and instead regenerates high-resolution text and textures from the multimodal context.
Limitations: After the open-weights release, the community version still cannot fully reproduce the online 2K capability; the image model, Sparse Attention availability and a 4-step variant all remain "under consideration / on the way," with no firm dates from the AMA.
Source: InfoQ (Chinese-language source)
Google denies training Gemini on private docs after a leak claim
Chinese-language source. A game developer says Gemini named an unpublished character that existed only in a private Google Doc; Google denies scanning or training on private documents.
Reported by Windows Report and relayed by IT Home (IT之家) on August 8: an indie game developer noticed players querying Gemini about the studio's in-development game and found Gemini returning an unpublished character, "Vantage Tripod." The developer says the character existed only in a private Google Docs file, was never uploaded publicly and is not in the game files. Google responded that it does not scan private Google Docs or train Gemini on them, and that Gemini only accesses Workspace files at the user's explicit request; it added that if someone posts a public document link online, those pages may be crawled and indexed, and users control their Workspace permissions.
Limitations: How Gemini learned the unpublished detail is still undetermined and under investigation; this rests on the developer's account and Google's denial, with no third-party verification.
Source: IT Home (Chinese-language source)
Gemini Robotics 2 drives three robot bodies; Apollo 2 dexterity 32%–92%
On July 30 Google drove three robot configurations with one Gemini Robotics 2 checkpoint — two of them Apptronik's Apollo 2 humanoid — with disclosed dexterity from 32% to 92%.
Per RuntimeWire: Google DeepMind introduced Gemini Robotics 2 in a July 30 announcement and used a single model checkpoint to control three different configurations — two are Apptronik's Apollo 2 humanoid (one with SharpaWave hands, one with Inspire hands) and the third is a Franka Duo robot fitted with Robotiq grippers — to show the same model transfers across different robot bodies; the 32%-to-92% dexterity results come from the two Apollo 2 configurations. This delivers on the division of labor the two announced in December 2024: Apptronik supplies the body, deployment sites and physical training data, while Google supplies models to perceive instructions, plan tasks and control movement. The announcement was authored by Google robotics leader Carolina Parada.
Limitations: RuntimeWire notes commercial reliability remains the gating issue — the 32%–92% spread itself shows some tasks are far from reliably usable; this is a demonstration test, not production-deployment data.
Source: RuntimeWire · Google DeepMind
Alibaba's Qwen app adds five free features; Wan 3.0 enters public test
Alibaba shipped five Qwen3.8-MAX-powered features in the Qwen app on Aug 7 (free for now), and the next day Alibaba Cloud opened the Wan 3.0 video model to public testing.
Per Pandaily: on August 7 Alibaba released five new features in the Qwen app, all powered by the new Qwen3.8-MAX flagship and all free for now; the next day (August 8) Alibaba Cloud opened its Wan 3.0 video model to public testing, with an eye on enterprise buyers.
Limitations: This is a thin brief — the exact shape of the five features and the usage limits and enterprise pricing of the Wan 3.0 test are not detailed in the report; "free for now" means commercial terms are still undecided.
Source: Pandaily
Global AI news
Fields Medalist joins OpenAI to work on AI safety
Number theorist and new Fields Medalist Jacob Tsimerman is leaving the University of Toronto for OpenAI's AI-safety work; he published a paper on "omnicide events" last year.
Per The Decoder: newly awarded Fields Medalist Jacob Tsimerman, a number theorist at the University of Toronto, is joining OpenAI to work on AI safety. He argues AI is an extremely transformative technology and that society needs to invest far more in safety, and that mathematicians can contribute because AI still runs largely on an empirical basis with few guarantees about how the systems actually work. Last year he published a paper on "omnicide events" — scenarios in which AI could contribute to human extinction — and said, "Panic isn't the right response, but we need to honestly assess the risks," adding he is convinced AI will soon outperform humans in math research.
Limitations: Whether AI poses extinction-level risk is debated among experts; former DeepMind CEO Demis Hassabis called recent AI math advances progress but not yet a fundamental breakthrough on the order of AlphaGo's "Move 37," which would require cracking problems like the Millennium Prize Problems.
Source: The Decoder · arXiv
Airbnb says AI cut its feature shipping time by 60%
On its earnings call Airbnb said AI shortened the time from idea to shipped feature by 60%, as it tests an AI search with a toggle.
Per TechCrunch: Airbnb co-founder and CEO Brian Chesky said on the latest earnings call that AI has cut the time from conceptualization to shipping features by 60%; earlier this year the company said AI was writing 60% of its code. On the consumer side, Airbnb is testing a new AI-powered search experience with a toggle.
Limitations: The report notes Airbnb has been slow to push AI features into its consumer interface, and the AI search is still a test; the "60%" figures are the company's own, without independent verification.
Source: TechCrunch
Cohere Health digitizes clinical policies on Amazon Bedrock AgentCore
Cohere Health built a multi-tenant agentic architecture on AgentCore to turn prior-authorization clinical policies from static documents into machine-readable structured data.
Per the AWS blog: clinical-intelligence company Cohere Health built a multi-tenant agentic architecture on Amazon Bedrock AgentCore, using AgentCore Runtime's MicroVM isolation and unified tool access through AgentCore Gateway. It targets prior authorization — the approval health plans require before covering certain services or medications — one of healthcare's most manual processes, because the governing policies are trapped in static, unstructured formats and affect hundreds of millions of patients a year. Digitizing those policies into structured, machine-readable data using standard terminologies eases that bottleneck.
Limitations: Policy content varies by clinical area, geography, line of business and health plan and keeps evolving, so deployment must maintain clinical oversight; this is a vendor case-study post from AWS, with no third-party benchmark data.
Source: AWS Machine Learning Blog
Regional and early signals
China's AI-phone grading clears first batch; L3 is only the starting line
China's national AI-terminal grading test cleared 11 devices in July — including nine phones from Huawei, Motorola, Honor, vivo, OPPO, Xiaomi and Stepfun — with L3 framed as the entry level.
Per Pandaily: China's national "AI-terminal intelligence grading" test cleared a first batch of 11 mobile devices in July, including nine smartphones from Huawei, Motorola, Honor, vivo, OPPO, Xiaomi and Stepfun; the report frames L3 as the starting line for "AI-agent phones," not the finish.
Limitations: This is a thin item — the grading criteria, the capability boundaries between L3 and higher tiers, and the test methodology are not detailed; it rests on a single English trade-press report with no primary standard document, so treat it as an early regional signal.
Source: Pandaily
Get the latest AI model insights and tutorials from Jiufeng.
Explore more


