Overview
11 stories in this issue. The first 3 are today's priorities.
Hot model updates
- Top · Muse opens up with a dedicated secure VM for every user
- Top · V4.1 Flash enters a time-boxed internal beta on a natively multimodal architecture
- Top · GPT-6 Astra is generally available on Amazon Bedrock
- Cloudera brings Mistral's frontier models into hybrid data environments
- A 2B model with streaming spatial memory beats GPT-5 on spatial benchmarks
Global AI news
- A Chrome developer-experience veteran joins Anthropic to work on Claude Code
- An Anthropic pretraining researcher resigns, saying neither lab is acting responsibly
- SageMaker Feature Store adds feature-level writes
Regional and early signals
- Ant's Ling team open-sources a finance-tuned model and a matching research benchmark
- JD Logistics shows its robot fleet, with five new machines for warehouses, cold chain and pharmacies
- A domestic AI-for-science platform shows a typhoon tracking system in Shanghai

Jiufeng graphic based on the sources cited in this issue.
Hot model updates
01/11
Muse opens up with a dedicated secure VM for every user
Meta's personal agent keeps running after you close the app, with its browser and credentials sealed inside a per-user cloud VM.
Meta introduced Muse, a personal AI agent built to take actions rather than answer questions. It can send emails, book travel, negotiate bills and pursue long-term goals, it keeps working after the app is closed, and it comes back only when it needs approval. Interaction is built around messaging: a user describes a task or a goal, and Muse plans and executes, opening its own browser, filling forms and negotiating on the person's behalf. Meta's examples include selling a car for more and lowering a bill.
- Availability: a consumer service rolling out now in the US on iOS, Android and muse.ai, with a free tier and paid plans
- Isolation: each user gets a dedicated cloud VM, Muse Secure VM, holding the agent, its browser and all credentials
- Model: Muse Spark 1.3 is available today through the Meta Model API and Muse Code
- Weights: an open-weights release sits on Meta's stated roadmap
Limitations: Muse is a consumer service and developers cannot self-host it; the open-weights release is a roadmap statement with no date attached; the rollout is US-only for now; Meta published no task success rates, latency figures or price points.
Source: MarkTechPost · Muse on X
02/11
V4.1 Flash enters a time-boxed internal beta on a natively multimodal architecture
DeepSeek exposes V4.1 Flash through an API ID that expires, with the official release set for September 10.
DeepSeek opened a short internal beta for V4.1 Flash. The API ID is spelled out as deepseek-v4.1-flash-expires-on-0910, it is priced like V4 Flash and capped at 20 concurrent requests, and the model runs on a new architecture with native multimodal support. In its announcement, DeepSeek said internal and external testing showed V4.1 Flash surpassing V4 Pro on performance, cost, speed and task completion time, and that after the official launch — and before V4.1 Pro ships — every request sent to the Pro model will be routed to V4.1 Flash and billed at Flash pricing.
Limitations: the beta channel expires on September 10 and allows only 20 concurrent requests, so it cannot be load tested; "surpassed V4 Pro across all key metrics" is the vendor's own wording, with no benchmark names or scores in the announcement; the text reached readers as a user-posted copy on Hacker News rather than through an official page.
Source: Pandaily · Hacker News
03/11
GPT-6 Astra is generally available on Amazon Bedrock
Call it through the Bedrock APIs, or point ChatGPT Work and Codex at Astra running on Bedrock.
AWS said GPT-6 Astra from OpenAI is now generally available on Amazon Bedrock. There are two paths: call the model directly through the Bedrock APIs, or configure ChatGPT Work and Codex to use GPT-6 Astra on Bedrock. Alongside the launch, OpenAI is introducing enterprise plugins for ChatGPT Work that extend Astra's browser-use capabilities across common business applications. On safety, AWS said OpenAI evaluated GPT-6 Astra through its Preparedness Framework, with model-level safeguards working alongside Bedrock's own security and governance controls.
Limitations: this is AWS's own launch post, and it carries no benchmark scores, throughput numbers or Bedrock pricing; "deeper reasoning and sharper judgment" is vendor language with no third-party evaluation; the post does not list supported regions.
Source: AWS Machine Learning Blog · OpenAI Preparedness Framework
04/11
Cloudera brings Mistral's frontier models into hybrid data environments
A nine-figure partnership puts Mistral's models next to enterprise data in cloud, on-premises and hybrid deployments.
Cloudera announced a nine-figure strategic partnership with French model maker Mistral AI at its EVOLVE26 event in São Paulo. Mistral will integrate its large language models directly with Cloudera's hybrid data and AI platform, so models can run wherever enterprise data lives. Cloudera framed the problem around highly regulated industries: those companies want to use AI, but feeding their most sensitive and proprietary information to an external model is hard to do safely.
Limitations: the report gives only "nine-figure" with no exact amount, names no specific Mistral models, and offers no availability date or pricing; it is currently a single-outlet report with no joint announcement to check against.
Source: SiliconANGLE
05/11
A 2B model with streaming spatial memory beats GPT-5 on spatial benchmarks
Tsinghua, Tencent Hunyuan and NTU apply test-time training to continuous video so a small model can hold 3D memory.
The paper "Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training", accepted at ECCV 2026, comes from Tsinghua University, Tencent Hunyuan and Nanyang Technological University. The report says the method gives a 2B model a streaming spatial memory that beats GPT-5 on spatial benchmarks. It targets two engineering limits: with standard Transformer self-attention, compute and KV-cache memory grow quadratically with sequence length, so video running tens of minutes or longer overwhelms constrained hardware; and uniform frame sampling or keyframe selection destroys the continuous parallax and geometric cues that 3D spatial relationships depend on, with the dropped information unrecoverable afterwards. Instead of stretching the attention window further, the team has the model integrate and update 3D geometric memory during inference.
The report gives two ablations, both measured on VSI-Bench — one replaces every self-attention layer with pure TTT, the other removes the spatial prediction mechanism and falls back to pointwise linear projection:
| Ablation (VSI-Bench) | Metric | Full method | Ablated |
|---|---|---|---|
| Self-attention → pure TTT | Overall score | 64.4% | 53.9% |
| No spatial prediction | Numerical spatial reasoning | 64.0% | 60.7% |
Limitations: the two ablations change different components and are scored on different metrics, so the rows are not directly comparable; the report mentions no plans to release weights or code; the "beats GPT-5" claim rests on a single Chinese-language report with no independent reproduction. Chinese-language source.
Source: Leiphone
Global AI news
06/11
A Chrome developer-experience veteran joins Anthropic to work on Claude Code
Addy Osmani moves to Anthropic to take on the usability side of coding agents.
Addy Osmani, who spent roughly 14 years on developer tools at Google, said on X that he has joined Anthropic to work on Claude Code and improve it for developers. At Google he led developer experience across Chrome and later became a director at Google Cloud AI, with work spanning Chrome DevTools, Lighthouse, PageSpeed Insights, Puppeteer, Chrome Headless and Core Web Vitals; he also helped create TodoMVC and co-founded Yeoman. Anthropic has said Claude Code passed $2.5B in run-rate revenue by May.
Limitations: Osmani has not disclosed a formal title or which part of Claude Code he will own; the $2.5B figure is Anthropic's own May-dated disclosure, not audited financials.
Source: RuntimeWire · Addy Osmani on X
07/11
An Anthropic pretraining researcher resigns, saying neither lab is acting responsibly
Jacob Coxon announced his exit in a seven-post thread, concluding private labs cannot safely coordinate a slowdown.
Jacob Coxon, a pretraining researcher who worked at both OpenAI and Anthropic, resigned from Anthropic on Tuesday after concluding that competition among frontier labs is pushing them toward systems they may be unable to control. "Neither company is acting responsibly," he wrote in a seven-post thread published at 00:04 UTC on September 9th, accusing OpenAI and Anthropic of racing toward self-improving superintelligence and "gambling with our lives." His record sits inside the work he is criticizing: OpenAI listed him among contributors to GPT-4o and as a core research contributor to GPT-4.5, and he co-authored OpenAI research on weight-sparse transformers. The Wall Street Journal reported that he left OpenAI earlier in 2026 to join Anthropic. The report notes the departure follows a series of public disclosures, including three incidents Anthropic reported on July 30th.
Limitations: this is one researcher's personal statement, and the report carries no response from OpenAI or Anthropic; the accusations come without specific technical evidence or internal material; his reason for leaving OpenAI is relayed second-hand from the Wall Street Journal.

Image source: anthropic; mirrored on Jiufeng R2.
Source: RuntimeWire · Jacob Coxon on X · Anthropic incident report
08/11
SageMaker Feature Store adds feature-level writes
The new UpdateRecord API changes one feature value without reading and rewriting the whole record.
AWS added feature-level writes to Amazon SageMaker Feature Store. The new UpdateRecord API updates one or more feature values in a single call without reading or rewriting the entire record, and it works on both online store tiers: Standard, backed by Amazon DynamoDB, and In-Memory, backed by Amazon ElastiCache. Until now, changing even a single feature value required a full read-modify-write cycle: read the complete record with GetRecord, merge the new value in the application, then write the whole record back with PutRecord.
Limitations: the post describes the API capability but publishes no before/after numbers for write latency, throughput or cost, and does not state region availability or whether the call is billed separately.
Source: AWS Machine Learning Blog
Regional and early signals
09/11
Ant's Ling team open-sources a finance-tuned model and a matching research benchmark
Ling-3.0-flash-Fin targets the investment-research workflow, and the FinFIRST agent benchmark ships with it.
Ahead of the Inclusion Conference in Shanghai, Ant Group's Ling team released Ling-3.0-flash-Fin, its first finance-enhanced open model, and open-sourced it, alongside FinFIRST, a benchmark for financial search agents. The model is built on Ling-3.0-flash with continued pretraining on financial corpora, domain post-training and tool-use optimization, and focuses on four abilities — information retrieval, research reasoning, valuation modeling and report writing — which together cover the chain from finding and verifying sources to computing, modeling and producing a reviewable research output.
- Parameters: 124B total, 5.1B active
- Context: 256K
- Availability: open-sourced
- Companion: the FinFIRST financial search agent benchmark released at the same time
Limitations: the report gives no scores on FinFIRST or any other financial benchmark and names no comparison models, and it does not state the license or where the weights live; capabilities such as valuation modeling rest on the vendor's own description with no third-party verification. Chinese-language source.
Source: QbitAI
10/11
JD Logistics shows its robot fleet, with five new machines for warehouses, cold chain and pharmacies
A "Super Brain" model handles scheduling while 11 "Wolf" robots cover storage, sorting and delivery.
At the JDD conference on September 9th, JD Logistics showed the full lineup of its "Super Brain + Wolf" robot fleet. The Super Brain model handles decisions and scheduling across complex supply-chain scenarios, while the Wolf family has 9 products and 11 robots deployed across storage, sorting and delivery, with 5 new products and technologies added at the event. JD Logistics also stated a five-year procurement plan of 3 million robots, 1 million driverless vehicles and 100,000 drones.
Self-reported figures for the three new machines:
- Cang Wolf (in-warehouse mobile picking): claimed two-week deployment with no warehouse retrofit; 99.9% picking accuracy, over 85% SKU coverage and 80 units per hour (UPPH)
- Low-temperature Zhi Wolf (cold-chain goods-to-person): built for -20°C cold stores with active defrosting and 100% contactless power; claimed 100% more storage capacity, 200% higher throughput and 10% lower cost per order
- Mu Wolf Health (unstaffed pharmacy): mixed-batch and mixed-SKU storage, running 24/7 unattended; claimed 100% compliant traceability, 99.9% dispensing accuracy and 90 orders packed per hour
Limitations: every figure above comes from JD's own presentation at its own conference, with no third-party verification; the 3 million robots is a five-year procurement plan, not delivered volume; the report does not say how many warehouses each machine is running in. Chinese-language source.
Source: InfoQ China
11/11
A domestic AI-for-science platform shows a typhoon tracking system in Shanghai
Taichu Yuanqi builds on its own heterogeneous many-core AI chip and supports Loongson, Sunway, Phytium and x86 hosts.
At the 2026 Inclusion Conference in Shanghai on September 9th, Taichu (Hangzhou) Integrated Circuit exhibited the domestic AI-for-science computing platform it launched in July, aimed at weather forecasting, biomedicine, quantum mechanics, chemical materials and fluid dynamics. The platform runs on the company's own heterogeneous many-core AI chip and spans hardware, base components, acceleration libraries, model frameworks and a DevKit toolkit; it is compatible with Loongson, Sunway, Phytium and x86 CPUs and adapted to PyTorch, JAX, PaddlePaddle and MindSpore. The update shown for the first time is TecoWeatherNext, a typhoon tracking system that chains data ingestion, AI forecasting and typhoon identification, works with domestic weather models, and overlays typhoon tracks with wind fields, sea-level pressure, temperature and precipitation layers. The company said it has launched an industry-academia initiative on domestic AI-for-science compute with nearly 20 universities, research institutes and companies.
Limitations: the piece is an exhibition write-up, with no chip model number, compute figures or performance and cost comparisons against mainstream GPUs; the typhoon system's forecast accuracy and any comparison with operational weather systems are not published. Chinese-language source.
Source: QbitAI
Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.
Generate free

