AI Highlights

Meta launches Muse with a dedicated secure VM per user

Key Takeaways

Meta launches Muse with a per-user secure VM, DeepSeek opens a time-boxed V4.1 Flash beta, GPT-6 Astra reaches Amazon Bedrock, and Cloudera signs a nine-figure deal with Mistral.

jiufeng
September 9, 2026
42 min read
In this article

Overview

11 stories in this issue. The first 3 are today's priorities.

Hot model updates

  1. Top · Muse opens up with a dedicated secure VM for every user
  2. Top · V4.1 Flash enters a time-boxed internal beta on a natively multimodal architecture
  3. Top · GPT-6 Astra is generally available on Amazon Bedrock
  4. Cloudera brings Mistral's frontier models into hybrid data environments
  5. A 2B model with streaming spatial memory beats GPT-5 on spatial benchmarks

Global AI news

  1. A Chrome developer-experience veteran joins Anthropic to work on Claude Code
  2. An Anthropic pretraining researcher resigns, saying neither lab is acting responsibly
  3. SageMaker Feature Store adds feature-level writes

Regional and early signals

  1. Ant's Ling team open-sources a finance-tuned model and a matching research benchmark
  2. JD Logistics shows its robot fleet, with five new machines for warehouses, cold chain and pharmacies
  3. A domestic AI-for-science platform shows a typhoon tracking system in Shanghai
AI signal map for 2026-09-09

Jiufeng graphic based on the sources cited in this issue.

Hot model updates

01/11

Muse opens up with a dedicated secure VM for every user

Meta's personal agent keeps running after you close the app, with its browser and credentials sealed inside a per-user cloud VM.

Meta introduced Muse, a personal AI agent built to take actions rather than answer questions. It can send emails, book travel, negotiate bills and pursue long-term goals, it keeps working after the app is closed, and it comes back only when it needs approval. Interaction is built around messaging: a user describes a task or a goal, and Muse plans and executes, opening its own browser, filling forms and negotiating on the person's behalf. Meta's examples include selling a car for more and lowering a bill.

  • Availability: a consumer service rolling out now in the US on iOS, Android and muse.ai, with a free tier and paid plans
  • Isolation: each user gets a dedicated cloud VM, Muse Secure VM, holding the agent, its browser and all credentials
  • Model: Muse Spark 1.3 is available today through the Meta Model API and Muse Code
  • Weights: an open-weights release sits on Meta's stated roadmap

Limitations: Muse is a consumer service and developers cannot self-host it; the open-weights release is a roadmap statement with no date attached; the rollout is US-only for now; Meta published no task success rates, latency figures or price points.

Source: MarkTechPost · Muse on X

02/11

V4.1 Flash enters a time-boxed internal beta on a natively multimodal architecture

DeepSeek exposes V4.1 Flash through an API ID that expires, with the official release set for September 10.

DeepSeek opened a short internal beta for V4.1 Flash. The API ID is spelled out as deepseek-v4.1-flash-expires-on-0910, it is priced like V4 Flash and capped at 20 concurrent requests, and the model runs on a new architecture with native multimodal support. In its announcement, DeepSeek said internal and external testing showed V4.1 Flash surpassing V4 Pro on performance, cost, speed and task completion time, and that after the official launch — and before V4.1 Pro ships — every request sent to the Pro model will be routed to V4.1 Flash and billed at Flash pricing.

Limitations: the beta channel expires on September 10 and allows only 20 concurrent requests, so it cannot be load tested; "surpassed V4 Pro across all key metrics" is the vendor's own wording, with no benchmark names or scores in the announcement; the text reached readers as a user-posted copy on Hacker News rather than through an official page.

Source: Pandaily · Hacker News

03/11

GPT-6 Astra is generally available on Amazon Bedrock

Call it through the Bedrock APIs, or point ChatGPT Work and Codex at Astra running on Bedrock.

AWS said GPT-6 Astra from OpenAI is now generally available on Amazon Bedrock. There are two paths: call the model directly through the Bedrock APIs, or configure ChatGPT Work and Codex to use GPT-6 Astra on Bedrock. Alongside the launch, OpenAI is introducing enterprise plugins for ChatGPT Work that extend Astra's browser-use capabilities across common business applications. On safety, AWS said OpenAI evaluated GPT-6 Astra through its Preparedness Framework, with model-level safeguards working alongside Bedrock's own security and governance controls.

Limitations: this is AWS's own launch post, and it carries no benchmark scores, throughput numbers or Bedrock pricing; "deeper reasoning and sharper judgment" is vendor language with no third-party evaluation; the post does not list supported regions.

Source: AWS Machine Learning Blog · OpenAI Preparedness Framework

04/11

Cloudera brings Mistral's frontier models into hybrid data environments

A nine-figure partnership puts Mistral's models next to enterprise data in cloud, on-premises and hybrid deployments.

Cloudera announced a nine-figure strategic partnership with French model maker Mistral AI at its EVOLVE26 event in São Paulo. Mistral will integrate its large language models directly with Cloudera's hybrid data and AI platform, so models can run wherever enterprise data lives. Cloudera framed the problem around highly regulated industries: those companies want to use AI, but feeding their most sensitive and proprietary information to an external model is hard to do safely.

Limitations: the report gives only "nine-figure" with no exact amount, names no specific Mistral models, and offers no availability date or pricing; it is currently a single-outlet report with no joint announcement to check against.

Source: SiliconANGLE

05/11

A 2B model with streaming spatial memory beats GPT-5 on spatial benchmarks

Tsinghua, Tencent Hunyuan and NTU apply test-time training to continuous video so a small model can hold 3D memory.

The paper "Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training", accepted at ECCV 2026, comes from Tsinghua University, Tencent Hunyuan and Nanyang Technological University. The report says the method gives a 2B model a streaming spatial memory that beats GPT-5 on spatial benchmarks. It targets two engineering limits: with standard Transformer self-attention, compute and KV-cache memory grow quadratically with sequence length, so video running tens of minutes or longer overwhelms constrained hardware; and uniform frame sampling or keyframe selection destroys the continuous parallax and geometric cues that 3D spatial relationships depend on, with the dropped information unrecoverable afterwards. Instead of stretching the attention window further, the team has the model integrate and update 3D geometric memory during inference.

The report gives two ablations, both measured on VSI-Bench — one replaces every self-attention layer with pure TTT, the other removes the spatial prediction mechanism and falls back to pointwise linear projection:

Ablation (VSI-Bench)MetricFull methodAblated
Self-attention → pure TTTOverall score64.4%53.9%
No spatial predictionNumerical spatial reasoning64.0%60.7%

Limitations: the two ablations change different components and are scored on different metrics, so the rows are not directly comparable; the report mentions no plans to release weights or code; the "beats GPT-5" claim rests on a single Chinese-language report with no independent reproduction. Chinese-language source.

Source: Leiphone

Global AI news

06/11

A Chrome developer-experience veteran joins Anthropic to work on Claude Code

Addy Osmani moves to Anthropic to take on the usability side of coding agents.

Addy Osmani, who spent roughly 14 years on developer tools at Google, said on X that he has joined Anthropic to work on Claude Code and improve it for developers. At Google he led developer experience across Chrome and later became a director at Google Cloud AI, with work spanning Chrome DevTools, Lighthouse, PageSpeed Insights, Puppeteer, Chrome Headless and Core Web Vitals; he also helped create TodoMVC and co-founded Yeoman. Anthropic has said Claude Code passed $2.5B in run-rate revenue by May.

Limitations: Osmani has not disclosed a formal title or which part of Claude Code he will own; the $2.5B figure is Anthropic's own May-dated disclosure, not audited financials.

Source: RuntimeWire · Addy Osmani on X

07/11

An Anthropic pretraining researcher resigns, saying neither lab is acting responsibly

Jacob Coxon announced his exit in a seven-post thread, concluding private labs cannot safely coordinate a slowdown.

Jacob Coxon, a pretraining researcher who worked at both OpenAI and Anthropic, resigned from Anthropic on Tuesday after concluding that competition among frontier labs is pushing them toward systems they may be unable to control. "Neither company is acting responsibly," he wrote in a seven-post thread published at 00:04 UTC on September 9th, accusing OpenAI and Anthropic of racing toward self-improving superintelligence and "gambling with our lives." His record sits inside the work he is criticizing: OpenAI listed him among contributors to GPT-4o and as a core research contributor to GPT-4.5, and he co-authored OpenAI research on weight-sparse transformers. The Wall Street Journal reported that he left OpenAI earlier in 2026 to join Anthropic. The report notes the departure follows a series of public disclosures, including three incidents Anthropic reported on July 30th.

Limitations: this is one researcher's personal statement, and the report carries no response from OpenAI or Anthropic; the accusations come without specific technical evidence or internal material; his reason for leaving OpenAI is relayed second-hand from the Wall Street Journal.

Investigating three real-world incidents in our cybersecurity evaluations

Image source: anthropic; mirrored on Jiufeng R2.

Source: RuntimeWire · Jacob Coxon on X · Anthropic incident report

08/11

SageMaker Feature Store adds feature-level writes

The new UpdateRecord API changes one feature value without reading and rewriting the whole record.

AWS added feature-level writes to Amazon SageMaker Feature Store. The new UpdateRecord API updates one or more feature values in a single call without reading or rewriting the entire record, and it works on both online store tiers: Standard, backed by Amazon DynamoDB, and In-Memory, backed by Amazon ElastiCache. Until now, changing even a single feature value required a full read-modify-write cycle: read the complete record with GetRecord, merge the new value in the application, then write the whole record back with PutRecord.

Limitations: the post describes the API capability but publishes no before/after numbers for write latency, throughput or cost, and does not state region availability or whether the call is billed separately.

Source: AWS Machine Learning Blog

Regional and early signals

09/11

Ant's Ling team open-sources a finance-tuned model and a matching research benchmark

Ling-3.0-flash-Fin targets the investment-research workflow, and the FinFIRST agent benchmark ships with it.

Ahead of the Inclusion Conference in Shanghai, Ant Group's Ling team released Ling-3.0-flash-Fin, its first finance-enhanced open model, and open-sourced it, alongside FinFIRST, a benchmark for financial search agents. The model is built on Ling-3.0-flash with continued pretraining on financial corpora, domain post-training and tool-use optimization, and focuses on four abilities — information retrieval, research reasoning, valuation modeling and report writing — which together cover the chain from finding and verifying sources to computing, modeling and producing a reviewable research output.

  • Parameters: 124B total, 5.1B active
  • Context: 256K
  • Availability: open-sourced
  • Companion: the FinFIRST financial search agent benchmark released at the same time

Limitations: the report gives no scores on FinFIRST or any other financial benchmark and names no comparison models, and it does not state the license or where the weights live; capabilities such as valuation modeling rest on the vendor's own description with no third-party verification. Chinese-language source.

Source: QbitAI

10/11

JD Logistics shows its robot fleet, with five new machines for warehouses, cold chain and pharmacies

A "Super Brain" model handles scheduling while 11 "Wolf" robots cover storage, sorting and delivery.

At the JDD conference on September 9th, JD Logistics showed the full lineup of its "Super Brain + Wolf" robot fleet. The Super Brain model handles decisions and scheduling across complex supply-chain scenarios, while the Wolf family has 9 products and 11 robots deployed across storage, sorting and delivery, with 5 new products and technologies added at the event. JD Logistics also stated a five-year procurement plan of 3 million robots, 1 million driverless vehicles and 100,000 drones.

Self-reported figures for the three new machines:

  • Cang Wolf (in-warehouse mobile picking): claimed two-week deployment with no warehouse retrofit; 99.9% picking accuracy, over 85% SKU coverage and 80 units per hour (UPPH)
  • Low-temperature Zhi Wolf (cold-chain goods-to-person): built for -20°C cold stores with active defrosting and 100% contactless power; claimed 100% more storage capacity, 200% higher throughput and 10% lower cost per order
  • Mu Wolf Health (unstaffed pharmacy): mixed-batch and mixed-SKU storage, running 24/7 unattended; claimed 100% compliant traceability, 99.9% dispensing accuracy and 90 orders packed per hour

Limitations: every figure above comes from JD's own presentation at its own conference, with no third-party verification; the 3 million robots is a five-year procurement plan, not delivered volume; the report does not say how many warehouses each machine is running in. Chinese-language source.

Source: InfoQ China

11/11

A domestic AI-for-science platform shows a typhoon tracking system in Shanghai

Taichu Yuanqi builds on its own heterogeneous many-core AI chip and supports Loongson, Sunway, Phytium and x86 hosts.

At the 2026 Inclusion Conference in Shanghai on September 9th, Taichu (Hangzhou) Integrated Circuit exhibited the domestic AI-for-science computing platform it launched in July, aimed at weather forecasting, biomedicine, quantum mechanics, chemical materials and fluid dynamics. The platform runs on the company's own heterogeneous many-core AI chip and spans hardware, base components, acceleration libraries, model frameworks and a DevKit toolkit; it is compatible with Loongson, Sunway, Phytium and x86 CPUs and adapted to PyTorch, JAX, PaddlePaddle and MindSpore. The update shown for the first time is TecoWeatherNext, a typhoon tracking system that chains data ingestion, AI forecasting and typhoon identification, works with domestic weather models, and overlays typhoon tracks with wind fields, sea-level pressure, temperature and precipitation layers. The company said it has launched an industry-academia initiative on domestic AI-for-science compute with nearly 20 universities, research institutes and companies.

Limitations: the piece is an exhibition write-up, with no chip model number, compute figures or performance and cost comparisons against mainstream GPUs; the typhoon system's forecast accuracy and any comparison with operational weather systems are not published. Chinese-language source.

Source: QbitAI

Generate one yourself with JIUFENG

Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.

Generate free