Overview
8 stories in this issue. The first 3 are today's priorities.
Hot Model Updates
- Top · Qwen-Audio-3.1-Realtime launches with an ~85% price cut
Global AI News
- Top · Google Research open-sources RRSI for self-improving agent harnesses
- Top · One vLLM-Omni container generates an image, then a video
- AWS uses Nova Act agents for synthetic monitoring
- Mistral opens a Munich hub with BMW and Siemens Energy projects
Regional & Early Signals
- AMD to acquire World Labs for $8.2 billion
- Google starts retiring Google Assistant on Android phones
- Doubao adds a travel and local-services entry to its home screen

Jiufeng graphic based on the sources cited in this issue.
Hot Model Updates
01/08
Qwen-Audio-3.1-Realtime launches with an ~85% price cut
Qwen's new five-model Qwen-Audio-3.1 stack is built around a full-duplex realtime model. Realtime pricing drops about 85%, and ASR pricing drops by up to 95%.
Alibaba's Qwen team released Qwen-Audio-3.1, which covers ASR, TTS and realtime interaction. The main model, Qwen-Audio-3.1-Realtime, is a full-duplex speech model for voice agents that call tools. It is trained to reason, act and decide when to speak. On an adapted τ-Voice benchmark, task success rises from 78.4% to 82.0%.
- Access:
qwen-audio-3.1-realtime-plusis live on QwenCloud over WebSocket - Context: 262K tokens, with 245K max input and 16K max output
- Default limits: 60 requests and 100K tokens per minute
- Features: function calling, web search, structured outputs, context cache, fine-tuning
- Training: GRPO. Scoring checks the terminal state first, then permitted writes, then behavioral assertions, so a fluent reply can't make up for a failed state check
| Billing item | USD per 1M tokens |
|---|---|
| Audio input | 6.4 |
| Text input | 0.8 |
| Text and audio output | 24 (output text not charged) |
The same release cuts prices by about 85% on Realtime, about 70% on TTS and up to 95% on ASR. A companion model, Qwen-Audio-3.1-ASR-Flash-Filetrans, handles offline file transcription.
Limitations: The model is only available as a managed API, and Qwen did not announce open weights. The τ-Voice result comes from an adapted version of the benchmark and from Qwen's own release materials. No third party has reproduced it yet.
Source: QwenCloud model page · MarkTechPost · The Decoder
Global AI News
02/08
Google Research open-sources RRSI for self-improving agent harnesses
RRSI lets an agent rewrite its own prompts, tools and memory while its model weights stay frozen. Constraints on the improvement loop are meant to keep gains from being limited to the set being optimized.
RRSI (Regularized Recursive Self-Improvement) comes from Google Cloud AI Research, working with UNC-Chapel Hill, Stanford and Washington University in St. Louis. An agent can rewrite its harness: prompts, tools, memory, control flow and sub-agents. The model weights never change. The paper names three failure modes of harness-evolution loops: benchmark-specific fitting, noise chasing and complexity accumulation. RRSI counters them with constraints that include a leakage critic and a noise floor.
- License: Apache 2.0 code
- Requirements: Python 3.10+, accepts any LiteLLM model string
- Default policy model: Claude Opus 4.8 on Vertex AI
| Benchmark | Change |
|---|---|
| SWE-bench Verified (not used for selection) | 82.0% → 83.8% |
| JobBench (out of distribution) | +4.7 pts |
| GDPval (out of distribution) | +3.5 pts |
| APEX-Agents (out of distribution) | +3.7 pts |
| Harvey LAB held-out split | +2.3 pts |
The authors report gains on all six held-out splits.
Limitations: RRSI is released as a research framework. The results come from the authors' own paper, and the defaults rely on Claude Opus 4.8 on Vertex AI. Gains with other policy models need to be verified separately.

Image source: GitHub; mirrored on Jiufeng R2.
Source: GitHub: google-research/rrsi · arXiv paper · MarkTechPost
03/08
One vLLM-Omni container generates an image, then a video
An AWS tutorial deploys two endpoints from one vLLM-Omni Deep Learning Container. FLUX.2-klein-4B generates images in real time, and Wan2.1-VACE-1.3B animates them asynchronously.
This is Part 2 of AWS's vLLM-Omni series. A text prompt goes to the real-time FLUX.2-klein-4B endpoint, which returns a still image. The image and a motion prompt then go to the asynchronous Wan2.1-VACE-1.3B endpoint. The resulting MP4 is retrieved from Amazon S3, and an optional Streamlit UI is included. vLLM-Omni extends vLLM to audio, image and video through OpenAI-compatible APIs, and the AWS container adds routing middleware for SageMaker AI.
Limitations: This is a deployment tutorial, not a new service launch. The sample uses fixed ml.g6.xlarge and ml.g6e.xlarge instance types, and the video endpoint runs asynchronously, with output retrieved from S3.
Source: AWS Machine Learning Blog
04/08
AWS uses Nova Act agents for synthetic monitoring
AWS published a synthetic-monitoring approach built on Amazon Nova Act and Amazon Bedrock AgentCore. Agents run critical user journeys on a schedule, and a sample repository includes the complete implementation.
Synthetic monitoring emulates real user journeys through automated transactions. It validates critical workflows such as logins, purchases and form submissions on a schedule, so teams catch degraded performance or broken UI interactions before customers do. The AWS post describes an agent-driven version that aims to replace brittle UI scripts with resilient, managed journey validation. It also covers the architecture and design patterns. The sample repository runs the monitor every 10 minutes via EventBridge Scheduler, and AWS says Nova Act exceeded 90% browser-workflow accuracy in early enterprise deployments.
Example use cases in the post include ecommerce product search, pricing and inventory display, cart operations and checkout confirmation. Others are account access and transaction flows in financial services, booking journeys in travel and hospitality, signup and subscription upgrades for SaaS, and appointment-scheduling portals in healthcare.
Limitations: This is an AWS architecture and sample-code post, not a new service launch. The 90%-plus accuracy figure is AWS's own claim about early enterprise cases, with no independent evaluation, so results need to be tested against your own site.
Source: AWS Machine Learning Blog · GitHub sample repository
05/08
Mistral opens a Munich hub with BMW and Siemens Energy projects
On September 28, Mistral opened a hub in Munich. Its research teams focus on Physics AI and Industrial AI, and its applied engineers work directly with enterprise partners.
Mistral says Germany is the EU's largest industrial economy. The hub targets automotive, energy, aerospace and advanced manufacturing, with sovereignty at the center of the move. The announcement names several partnerships:
- BMW: crash simulation and engineering AI in Munich
- Siemens Energy: industrial AI applications
- Technical University of Munich (TUM): research partnership using its wind tunnel for automotive aerodynamics digital twins
- Team: more than 30 researchers who joined through the Emmi AI acquisition
Limitations: This item cites only Mistral's own announcement. The partnership details and progress are Mistral's description, with no independent reporting yet.
Source: Mistral AI
Regional & Early Signals
06/08
AMD to acquire World Labs for $8.2 billion
According to QbitAI, AMD is buying the world-model company World Labs outright, and Fei-Fei Li will become AMD's EVP and chief scientist after closing.
QbitAI reports the deal is valued at $8.2 billion and that Fei-Fei Li will report directly to Lisa Su. After closing, co-founders Justin Johnson and Ben Mildenhall will bring the World Labs team into AMD to form a frontier AI research organization. World Labs raised a $1 billion Series C on February 18, 2026 at a $5.4 billion post-money valuation, and AMD Ventures had invested in earlier rounds. In early September, World Labs released Atlas, a multimodal world model pretrained from scratch that can reconstruct a scene from one to dozens of images and generate new viewpoints.
Limitations: Chinese-language source. The deal has not closed, and the terms and roles come from a Chinese-language media report; this item does not link AMD's original announcement. The description of Atlas comes from World Labs, with no independent evaluation.
Source: QbitAI (Chinese-language source)
07/08
Google starts retiring Google Assistant on Android phones
The switch Google announced earlier is now underway. Many Android users report they can no longer use Google Assistant, and the "Switch to Google Assistant" option is gone from the Gemini app menu.
IT Home, citing Android Authority, reports that Google began emailing users last month to say mobile Google Assistant would be discontinued. The change affects phones and phone accessories such as smartwatches, Android Auto and earbuds. Google has said smart speakers and smart displays will still support Google Assistant.
Limitations: Chinese-language source. The rollout status comes from user reports and media relays. The report notes that some users say Gemini still hallucinates and lacks some offline voice commands.
Source: IT Home (Chinese-language source)
08/08
Doubao adds a travel and local-services entry to its home screen
The Doubao app, which has more than 168 million daily active users, added a "出行用豆包" (travel with Doubao) entry above its input box. It sits alongside "Chat" and "Help me write."
The entry has four modules: maps, transport, hotels, and food and entertainment. According to TMTPost, maps come from Baidu Maps and ride-hailing from Caocao Mobility, whose AI ride service with Doubao launched this month, first in Beijing, Hangzhou and Suzhou. Hotels link to Douyin's travel service, and dining links to Douyin local services. Doubao has responded that merchants cannot pay to influence recommendations or rankings.
| Channel | Hotel order take rate |
|---|---|
| Doubao entry | ~12% (11.4% fee + 0.6% payment) |
| Douyin main app | 8% |
| Ctrip, Meituan | ~15% |
Limitations: Chinese-language source. The fee and partner details come from a TMTPost analysis piece. The article itself notes that Meituan, Didi and Ctrip have spent years building fulfillment and supply-chain moats.
Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.
Generate free

