Overview
11 stories in this issue. The first 3 are today's priorities.
Hot Model Updates
- Top · GPT-6 Astra tops two agent benchmarks and sweeps all five drone tasks
- Top · Cognition releases SWE-2, post-trained from Kimi K3 at 64% lower cost
Global AI News
- Top · AWS open-sources Pizza Bot, an inbox for background agents
- AgentsDock enters beta for supervising coding agents from a phone
- Higgsfield puts GPT-driven motion graphics inside After Effects
- Study: written reasoning steps map to distinct internal patterns
- Two-year classroom experiment: the AI-free group finished last both years
Regional and Early Signals
- Zhipu closes about $5 billion for the next GLM and "fully self training"
- Figma's security agents cut complex alert handling time by about 70%
- Liangyuan Xinchuang ships three Physical AI releases in just over a month
- A $399 open source robot duck passed $2.6M in orders in 24 hours

Jiufeng graphic based on the sources cited in this issue.
Hot Model Updates
01/11
GPT-6 Astra tops two agent benchmarks and sweeps all five drone tasks
Andon Labs measured a $15,515 average final bank balance for Astra running a vending business, nearly three times Claude Fable 5.1, plus the first clean sweep of Drone-Bench's five subtasks.
Andon Labs uses Vending-Bench and Drone-Bench to measure how well models act independently over long periods and how well they write software for physical systems. In the simulated vending machine business, Astra averaged a $15,515 final bank balance, and even its worst run ended at $13,272. Fable 5.1 averaged $5,422 across six runs, with a best run of $9,874. On Drone-Bench, Astra became the first model to beat the human-AI baseline on all five subtasks, including writing code that lets a drone autonomously find and follow a specific person.
| Subject | Average final balance | Best run | Worst run |
|---|---|---|---|
| GPT-6 Astra | $15,515 | not given | $13,272 |
| Claude Fable 5.1 | $5,422 (6 runs) | $9,874 | not given |
A separate setup, Vending-Bench Arena, puts several agents in competition at the same location. There, Astra refused an illegal price-fixing proposal from GLM-5.3, while Fable 5.1 reached that arrangement with GLM-5.3.
Limitations: The Decoder notes Astra's success rate on the drone tasks remains unreliable. Andon Labs' post on X states only that Astra is the first model to beat humans on each of the five Drone-Bench tasks in at least one attempt — a best-of-runs measure. Both benchmarks are built and run by Andon Labs itself.
Source: The Decoder · Andon Labs on X
02/11
Cognition releases SWE-2, post-trained from Kimi K3 at 64% lower cost
SWE-2 scores 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 at 64% lower cost, but it runs only inside Devin.
Cognition, the company behind the Devin coding agent, released SWE-2 as its most capable coding model to date. It is post-trained with reinforcement learning from Kimi K3, Moonshot AI's 2.8T-parameter open model.
- Score: 50.0% on FrontierCode 1.1 Main, within 1 point of Fable 5.1
- Cost: 64% lower than Fable 5.1
- RL headroom: Cognition says its RL adds 5 to 6 points on many benchmarks on top of K3
- New feature: its first model with selectable reasoning-effort levels, all trained in a single RL run
- Versus previous generation: SWE-1.7 was post-trained from Kimi K2.7; this base model has almost 3x the parameters
Limitations: SWE-2 has no open weights and no standalone API. It runs only inside Devin — Desktop and CLI today, with Devin Web and Fusion rolling out. The 50.0% figure is self-reported by Cognition.
Source: MarkTechPost · Kimi K3 paper
Global AI News
03/11
AWS open-sources Pizza Bot, an inbox for background agents
Pizza Bot manages agent tasks that keep running in the background through an email-style inbox, licensed Apache 2.0, after earlier versions served more than 2,000 people inside Amazon.
AWS rebuilt an internal agent application as the open source project Pizza Bot: self-hosted, built on DeepAgents and LangGraph, sorting completed results and pending decisions into an inbox. Threads are split into All (thread history), Unread (completed work awaiting review) and Action (paused for approval or an answer). Users can organize threads into folders and inspect delegated workers in the Activity panel. Tasks start manually, through cron schedules or through webhooks, with the server owning scheduling; after downtime, missed cron intervals produce one catch-up run. Earlier versions served more than 2,000 people inside Amazon for meeting preparation, email drafting, Slack summaries, CRM logging and research.
Availability: macOS, Windows and Linux desktop builds, plus browser and terminal clients connected to a local or standalone backend. The code is licensed under Apache 2.0.
Limitations: this is a self-hosted application, so the backend is yours to deploy; missed schedules only produce a single catch-up run rather than one per interval; work that needs human sign-off stops in the Action state until someone handles it.

Image source: GitHub; mirrored on Jiufeng R2.
Source: MarkTechPost · GitHub · AWS open source blog
04/11
AgentsDock enters beta for supervising coding agents from a phone
Nvidia senior research scientist Zhengyi Luo turned his robotics-training operations into AgentsDock, connecting Claude Code, Codex and Cursor to the machines researchers already use.
AgentsDock is a desktop and mobile workspace for supervising coding agents across a researcher's existing computers. The current desktop build is version 0.2.13-beta.33, with clients for macOS, Windows and Linux alongside beta access for iPhone, iPad and Android. Its interface gives plots, videos, files and rendered experiment rollouts the same billing as chat and source code: a researcher can inspect a training result on mobile, reconnect to a persistent terminal and send another instruction without returning to the machine running the job. The architecture has two pieces — a client on a computer, phone or tablet, and the server from the public AgentsServer repository. Zhengyi "Zen" Luo, a senior research scientist at Nvidia, originally built the system to orchestrate robotics-training jobs and now runs much of that work from his phone.
Limitations: the software is still beta (0.2.13-beta.33) and mobile is beta access only; the report gives no general-availability timeline and no pricing.
Source: RuntimeWire · AgentsServer repository
05/11
Higgsfield puts GPT-driven motion graphics inside After Effects
AI Motion Designer delivers its output as an editable After Effects composition, at the price of holding Higgsfield, ChatGPT and Adobe subscriptions at once.
Higgsfield introduced AI Motion Designer on September 10th, and the company's account promoted the workflow in a post on X. Users can ask the agent to animate a title, assemble a logo reveal or add a transition. Higgsfield says the output arrives in an open After Effects composition as editable layers, keyframes and timing, letting a designer revise the project instead of regenerating a flattened video. As RuntimeWire describes the split, the model handles planning, code writing and computer use while Higgsfield supplies the media-generation and editing tools. The Adobe workflow is separate from 3D Jutsu, the standalone browser-based 3D workspace Higgsfield presented on September 5th, and its Blender plugin is a separate product again.
Limitations: the workflow requires Higgsfield, ChatGPT and Adobe subscriptions together, and RuntimeWire notes its value depends on whether that revision model saves enough work to justify three subscriptions.
Source: RuntimeWire · Higgsfield on X · OpenAI
06/11
Study: written reasoning steps map to distinct internal patterns
KAIST and Naver AI Lab found that operations such as extraction, formula recall and deduction are separable inside a model's internal states, with the strongest signal in the middle layers.
The team defined eight recurring reasoning operations, including extraction, decomposition, formula recall, deduction and computation. Qwen2.5-7B, Qwen3-8B and Gemma4-31B solved math problems, the solution paths were split into segments, and GPT-5 labeled each segment with one of those operations. The steps turned out to be separable in the models' numerical representations, and the signal was strongest in the middle layers. The separability replicated on Llama-3-8B, and a classifier trained on Qwen3-8B transferred successfully to GPQA-Diamond and MATH-500. The Decoder notes this matters for AI safety, because it shows which step a model is internally performing.
Limitations: the per-segment operation labels come from GPT-5 rather than human annotators.
Source: The Decoder
07/11
Two-year classroom experiment: the AI-free group finished last both years
A law course at Vrije Universiteit Amsterdam randomly split students into no-AI, unguided-AI and trained-AI groups, and the no-AI group came last in both years.
Thibault Schrepel of Vrije Universiteit Amsterdam ran the experiment over two years in his own Law of AI course. The task was identical for all three groups: in teams of four or five, they had 20 minutes to improve a provision of the EU AI Act, graded on substance, clarity, proportionality and innovation. The first group could not use ChatGPT. The second received ChatGPT-generated revision suggestions embedded directly in the text and could keep using the tool with no guidance on how. The third got hands-on training in legal prompt engineering and in checking AI suggestions for consistency and accuracy. Over both years, the group without AI finished last. "I was wrong," Schrepel writes; he had assumed unguided AI use would do more harm than good.
Limitations: this is a single-course experiment with a 20-minute task, scored through a group exercise plus the same exams for everyone (including a multiple-choice test), so sample size and grading are bounded by the course setup.
Source: The Decoder
Regional and Early Signals
08/11
Zhipu closes about $5 billion for the next GLM and "fully self training"
The raise combines roughly $2 billion in share placement with roughly $3 billion in convertible bonds, earmarked for the next GLM base model, a self-training system and compute infrastructure.
Zhipu announced on September 13th that it had closed roughly $5 billion in financing (Chinese-language source).
| Component | Amount |
|---|---|
| Share placement | about $2 billion (about HK$15.68 billion) |
| Convertible bonds | about $3 billion (about RMB 20.14 billion) |
| Total | about $5 billion (about RMB 33.654 billion) |
The announcement says the money mainly goes to research on the next-generation GLM base model and a Fully Self Training system, plus deployment and upgrades of large-scale training, production inference, compute resources and related infrastructure. Zhipu describes fully self training as the next GLM being trained in environments built by the previous GLM, forming a recursive self-improvement loop, with spending on automated generation and filtering of training data, task-environment construction, longer-horizon reasoning, plus domestic chip adaptation, operator development and inference optimization.
Limitations: the announcement gives only the funding structure and its direction — no timeline, parameter count or benchmark target for the next GLM. Fully self training is at this point Zhipu's own description and plan, with no third-party results to check it against. Chinese-language source.
Source: ITHome
09/11
Figma's security agents cut complex alert handling time by about 70%
Figma's engineering team documented an agent system built on Panther SIEM, self-reporting roughly 70% less time on complex alerts and 20% fewer on-call pages.
Reported via a Chinese-language source.
- Investigation scope: checks audit logs in AWS, Okta, GitHub, GCP and osquery, queries more than 100 other data sources, and can open pull requests
- Triage agent: uses a Claude Opus-class model, fed the full Slack thread as context plus its own guided memory, with the toolset an on-call security engineer normally needs
- Retrieval stack: AWS Bedrock knowledge bases, Amazon Kendra, Tines and a Snowflake-based tool for pulling past alerts and analyzing Panther data
- Three kinds of memory: past alerts, behavioral guidelines and learned database schemas, each managed separately
- Numbers from a companion post: the agents found more than 100 previously undetected vulnerabilities, including two high-severity ones traditional tools missed; code review accuracy reached 80% within a month; a second review pass raised detection of known vulnerabilities by about 30%; certain coding mistakes dropped roughly 50% after automated guidance was added
Limitations: part of the ~70% speedup comes from downgrading the severity of some alerts. The authors explicitly say they cannot tell you what to do, since it depends on company size, risk and existing feedback loops, and advise raising precision before recall. Agent-generated PRs default to draft and human review stays in place. Chinese-language source.
Source: InfoQ China · AWS Bedrock knowledge bases docs
10/11
Liangyuan Xinchuang ships three Physical AI releases in just over a month
QbitAI reports one navigation model transferring zero-shot across four robot bodies, trained on more than 2,000 real scenes ported into simulation.
Liangyuan Xinchuang (亮源新创) released LightParkour, LightNav-0 and Light REACT over the past month or so, and frames the whole line as a three-stage paradigm: scalable pre-training, scalable alignment and scalable deployment. The report frames the problem space as changing obstacles, changing scenes, changing bodies, and whether a robot keeps working after impacts, falls or joint damage (Chinese-language source).
The figures given in the report: more than 2,000 real-world scenes moved into simulation and a single navigation model working zero-shot across four robot embodiments; adding Point CoT raised average task success across eight evaluations by 8.4 percentage points and SPL by 5.7 percentage points; obstacle-crossing height went from 45 cm to 75 cm, about 83% of the robot's own height, at a 50 Hz control frequency. The report says the company released the models, the code and a technical report.
Limitations: these figures come from the company's own release and QbitAI's account, with no third-party reproduction, and the zero-shot cross-embodiment result is the company's claim. Chinese-language source.
Source: QbitAI
11/11
A $399 open source robot duck passed $2.6M in orders in 24 hours
Microduck, the 25-centimeter open source biped from Hugging Face's Pollen Robotics, took 10,500 orders in four days and has pushed delivery out to four months.
Microduck sells for $399 (about RMB 2,800) and is developed and open-sourced by Pollen Robotics, part of the AI community Hugging Face. It stands 25 centimeters tall (Chinese-language source).
| Time since orders opened | Cumulative order value | Cumulative units |
|---|---|---|
| 6 hours | passed $1 million | not given |
| 24 hours | passed $2.6 million | not given |
| 4 days | not given | 10,500 |
At peak it sold one unit every four seconds, and new orders now quote four-month delivery. Price is not the only draw: Microduck ships a full open source reinforcement learning pipeline from simulated training to real-robot deployment, so a user can teach a virtual duck a new motion in simulation and then transfer the learned behavior onto the physical robot. Chinese resale listings have appeared at up to RMB 5,567, roughly twice list price.
Limitations: these order figures come from a column republished by TMTPost, not from a Pollen Robotics disclosure, and the four-month queue means delivery and real-world use are still unverified. Chinese-language source.
Source: TMTPost
Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.
Generate free

