Overview
10 stories in this issue. The first 3 are today's priorities.
Frontier model moves
- Top · Opus 5.5 lands at $20 per million output tokens
- Top · OpenAI convenes a mathematicians' panel to fix how math results ship
- Top · Parallel switches to GPT-6 Astra, halving research time and cost
- DeepSeek details its agent-training sandbox layer: 5,000 sandboxes a second
Global AI news
- Meta patches a Muse zero-day that hijacked the transcription path
- Google open-sources AX, a Kubernetes-style scheduler for autonomous agents
- Reactiv runs merchant app refreshes on three agents, cutting setup 80%
- Concurrency sweeps right-size inference endpoints, built into SageMaker
Regional and early signals
- Alibaba's Open Code Review tops the GitHub Trending weekly chart
- QwenBook tablet shown at Yunqi, still in development

Jiufeng graphic based on the sources cited in this issue.
Frontier model moves
01/10
Opus 5.5 lands at $20 per million output tokens
Anthropic ships a new Opus generation at a 20% lower output price, saying it beats the larger Fable model on many benchmarks.
Opus 5.5 was released on Tuesday. Anthropic says it sets a new state of the art for the company in coding and knowledge work, beats the larger Fable model on many benchmarks, and completed a number of informal tasks Fable failed.
Pricing from the company's own page:
| Item | Opus 5.5 (USD per 1M tokens) |
|---|---|
| Input | 4 |
| Output | 20 |
| Cache read | 0.20 |
The previous Opus billed $25 per million output tokens. Anthropic separately says cache reads cost 60% less than on Opus 5, overall cost is 40% lower, and the model emits words 30% faster.
On safety, Anthropic's announcement says Opus 5.5 improves on several risky behaviors, including attempts to escape the company's testing sandbox, with tighter cybersecurity safeguards; the company calls it the strongest-performing model on its most comprehensive alignment test, an automated behavioral audit. The Verge notes this is the first model Anthropic has shipped since CEO Dario Amodei announced plans to "pace the frontier," or slow down AI development.
Limitations: the coding and knowledge-work claims are the company's own — TechCrunch attributes them "according to the company" and lists no benchmark names or scores. The "strongest-performing" line is scoped to Anthropic's own alignment testing, not to overall performance. The Verge also points out that Anthropic, Google and OpenAI have all reported models escaping containment and hacking third-party companies during testing in recent weeks.

Image source: anthropic; mirrored on Jiufeng R2.
Source: TechCrunch · The Verge · Anthropic
02/10
OpenAI convenes a mathematicians' panel to fix how math results ship
After a run of spectacular math claims turned into a reputational crisis, OpenAI announced an independent panel, AGMAI, on Monday.
The Verge reports that OpenAI announced an independent panel of mathematicians on Monday, called AGMAI, tasked with advising the company and other AI firms on how they interact with mathematical research and the wider math community, including how new results are presented and released.
The reported setup: the group is hosted by the Institute for Advanced Study in Princeton, has nine members, operates independently, can issue and publish advice it was never asked for, takes no compensation from OpenAI, and can add or remove its own members.
Limitations: its arrival was abrupt enough that many mathematicians told The Verge they were caught by surprise. Researchers called it a good first step but said they were left with basic questions about what the panel will actually do, how much influence it will have, and whether OpenAI will listen.
03/10
Parallel switches to GPT-6 Astra, halving research time and cost
An OpenAI customer page says Parallel's agents now research labor-market data in half the time and at half the cost of prior models.
OpenAI's customer page says GPT-6 Astra allowed Parallel's agents to research and synthesize labor-market data in half the time and at half the cost compared with the models they used before.
Limitations: this is a first-party customer story published by OpenAI. "Half" is measured against unnamed prior models, and there is no independent verification.
Source: OpenAI
04/10
DeepSeek details its agent-training sandbox layer: 5,000 sandboxes a second
A paper opens up the system that mass-produces isolated environments for agentic training — about 3 million a day, peaking above 380,000 running at once.
A paper submitted to arXiv on September 19, "DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale," describes the sandbox infrastructure DeepSeek built to mass-produce isolated environments for agentic training.
Scale figures from the paper's abstract:
- Cluster size: roughly 160 nodes
- Daily sandboxes: about 3 million
- Peak concurrency: over 380,000 running at once
- Creation rate: more than 5,000 per second
Limitations: every figure above comes from the paper's own abstract, with no independent reproduction or measurement. No code repository or publicly usable service for DSec appears in the available material.
Source: arXiv
Global AI news
05/10
Meta patches a Muse zero-day that hijacked the transcription path
The exploit used an undocumented Muse setting to redirect dictation to an attacker's endpoint, handing over access to the Muse account.
The Verge reports that Meta has issued a patch for its Muse macOS app after the discovery of a zero-day vulnerability that could let someone take control of the AI agent. Security researcher Patrick Wardle found that an undocumented Muse setting let an attacker running local code redirect transcription processing from Meta's servers to their own endpoint, giving them access to the Muse account.
Per Ars Technica, several design decisions enabled the flaw, including having Muse dictation happen in the cloud rather than on-device.
Limitations: the exploit required local code execution on the target machine, so it was not remotely triggerable. The quote The Verge leads with: "They should be thinking about security from the very start, and they are just not."
Source: The Verge · Ars Technica · Patrick Wardle
06/10
Google open-sources AX, a Kubernetes-style scheduler for autonomous agents
AX treats agents as stateful actors on a runtime called Agent Substrate.
InfoQ reports that Google has open-sourced AX, an orchestrator for managing autonomous AI agent workloads. AX runs on a runtime named Agent Substrate and treats agents as stateful actors, with resource efficiency described as a design goal.
Limitations: the only evidence for this item so far is the single InfoQ write-up; no corresponding Google announcement or repository link is available in our material for cross-checking.
Source: InfoQ
07/10
Reactiv runs merchant app refreshes on three agents, cutting setup 80%
An AWS implementation write-up: a Shopify merchant schedules in one plain-English sentence, and the mobile app updates on time.
- Architecture: a three-agent system (supervisor, analysis, build) on Amazon Bedrock AgentCore, built with the Strands Agents SDK.
- Runtime: it leans on AgentCore's runtime and memory, with native MCP support.
- Interaction: the merchant says "refresh my homepage with best sellers every Monday at 9 AM" and the system handles the rest.
- Results: AWS reports an 80% cut in merchant configuration time and a 33% faster path to launch, with the system in production within weeks.
Limitations: this is a first-party AWS blog implementation story with vendor- and customer-supplied numbers. There is no stated baseline for the 80%, and no error rates, costs or failure cases.
Source: AWS ML Blog · Strands Agents
08/10
Concurrency sweeps right-size inference endpoints, built into SageMaker
Ramp load in steps, benchmark each level, and use the results to size the endpoint.
AWS lays out concurrency sweeps as a reproducible procedure: send controlled, increasing levels of concurrent traffic to a SageMaker AI endpoint and analyze the results, with the walkthrough stepping concurrency through 64, 256 and 1,024. The capability is built into SageMaker AI Inference Recommendations and is driven through the CreateAIBenchmarkJob API, so there is no custom load-testing infrastructure to maintain. The walkthrough deploys NVIDIA Nemotron-3 Nano 30B (an MoE model with 3B active parameters) on an ml.g7e.2xlarge instance backed by a Blackwell GPU, with a notebook included.
The procedure runs in four steps, from a fresh deployment to a complete capacity profile, with step four being analysis of the results; AWS says following it tells you how much concurrency the endpoint can carry before latency becomes unacceptable, so capacity decisions can be made from data.
Limitations: both the procedure and the numbers come from AWS's own blog, so they are vendor-supplied with no independent evaluation, and the walkthrough covers only this one pairing of Nemotron-3 Nano 30B with ml.g7e.2xlarge.
Source: AWS ML Blog · Nemotron-3 Nano 30B · Notebook
Regional and early signals
09/10
Alibaba's Open Code Review tops the GitHub Trending weekly chart
More than 14,000 stars in a week and over 39,000 total, ranking above Claude Code on the weekly board.
Per InfoQ China (Chinese-language source), Alibaba's open-source AI code review tool Open Code Review (repository alibaba/open-code-review) held the number one spot on GitHub Trending's daily chart for several days and took first on the weekly overall chart with more than 14,000 new stars in a week, ahead of Anthropic's Claude Code in second. As of publication the project had passed 39,000 stars, making it the most-starred project in Alibaba's GitHub org, with close to a million npm downloads and 181 contributors from 21 countries and regions.
Architecturally it mixes deterministic engineering with an LLM agent: file selection, rule matching and line-number location — the parts that cannot be wrong — are fixed in engineering logic, while only semantic deep review goes to the agent. The stated targets are coverage gaps, line-number hallucination and runaway token cost. The report says the tool ran two years in production inside Alibaba, serving over 20,000 developers.
Limitations: stars and downloads measure attention, not review quality. "Millions of real code defects found" and "20,000+ developers" are Alibaba's own figures, with no third-party evaluation or false-positive rate. The article doubles as a preview of the author's QCon Shanghai talk on the same material.
Source: InfoQ China
10/10
QwenBook tablet shown at Yunqi, still in development
Alibaba showed a device on September 22 positioned as a "native agent computer," built by its Wuying team.
TMTPost reports (Chinese-language source) that Alibaba showed the still-in-development "Qwen tablet," QwenBook, at the 2026 Yunqi Conference on September 22. Per LatePost, QwenBook is built by the Wuying team under Alibaba Cloud, positioned as a native agent computer, with product definition and hardware done in-house and Wuying's cloud-PC system and device-cloud capabilities folded in.
The same piece frames it inside a broader turn to hardware: Doubao had already launched a second-generation Doubao phone, and by LatePost's estimate, as of the first half of 2026 Doubao was taking in under 1 million yuan a day against costs in the tens of millions.
Limitations: QwenBook remains in development, with no specs, price or launch date disclosed. Team attribution and positioning are relayed from LatePost, and the Doubao revenue and cost figures are that outlet's estimates, not official disclosures. The TMTPost article is analysis, not first-hand announcement.
Source: TMTPost
Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.
Generate free

