Overview
10 stories in this issue. The first 3 are today's priorities.
Hot model updates
- Top · Qwen3.8-LiveTranslate cuts average interpretation lag to 2.3 seconds
- Top · StepFun's Step 5 Preview enters the open-weight top three at one-eighth of Opus 5's task cost
- Top · Open-weight models reach 78.4% of Vercel AI Gateway token volume
- Opus 5.2 and a likely Gemini 4 Pro are already live with no announcement
Global AI news
- Runway wants AI video generation to become a stream you steer live
- Microsoft open-sources TauGrid, collapsing GPU cluster plumbing into one Helm install
- Microsoft and the University of Illinois train AI tutors on simulated students that make mistakes
- Google ships Agent Development Kit for Kotlin 1.0 at parity with Python
Regional and early signals
- APUS open-sources a Jev reproduction that runs browser decisions on a local 9B model
- Spirit AI's Moz1 wheeled humanoids are running on CATL and JD.com lines

Jiufeng graphic based on the sources cited in this issue.
Hot model updates
01/10
Qwen3.8-LiveTranslate cuts average interpretation lag to 2.3 seconds
Alibaba's Qwen team shipped a real-time interpretation model that returns translated text and speech while the speaker is still talking, available now as a hosted API.
Qwen3.8-LiveTranslate listens to live speech, optionally with video frames, and translates as it listens. The core change is a new Interleave architecture; average lagging (LAAL) drops from 2.8 seconds to 2.3 seconds. The release also adds real-time speaker diarization, synchronized bilingual display, and long-context disambiguation.
- Latency: LAAL down from 2.8s to 2.3s, where LAAL measures how far the translation trails the speaker on a length-adaptive basis
- Languages: 60
- Input: live speech, with optional video frames
- Availability: hosted API on Alibaba Cloud Model Studio and QwenCloud as qwen3.8-livetranslate-flash-realtime over WebSocket
Limitations: the only shipping form is a hosted API — the report states deployability as a hosted API and mentions no open weights or self-hosting option; the claimed gains in faithfulness, fluency and conciseness come without comparison scores; LAAL is a latency metric, so lower lag does not by itself mean better translation.
Source: MarkTechPost · QwenCloud model page · LAAL paper
02/10
StepFun's Step 5 Preview enters the open-weight top three at one-eighth of Opus 5's task cost
A 600B sparse MoE with 27B active parameters, with weights promised for October 15.
StepFun released Step 5 Preview on September 20, aimed at AI coding, software engineering, professional knowledge work and finance. It scores 44 on the Artificial Analysis Intelligence Index, placing in the top three among open models globally, and sits on the Pareto frontier between intelligence and per-task cost.
- Architecture: sparse MoE, 600B total parameters, 27B active
- Context: 1M tokens, native text and vision input
- Cost: per-task cost is 1/8 that of Claude Opus 5
- Open weights: the company says the model will be open sourced on October 15
Limitations: what shipped is a preview, with weights held until October 15; Leiphone's capability claims rest on "public and internal evaluations," and the internal part cannot be checked externally; ifanr's hands-on is a subjective read of Three.js generation cases (a gomoku game, a speedboat game, a hot-air-balloon scene, Niagara Falls, a web desktop) with no benchmark numbers.
Source: Leiphone · ifanr (Chinese-language source)
03/10
Open-weight models reach 78.4% of Vercel AI Gateway token volume
In a single-day snapshot from Vercel's founder, open models carried more than three times the tokens closed models did.
Vercel founder and CEO Guillermo Rauch published the gateway's token split in a September 19 daily snapshot; Aligned News highlighted the figure on September 20.
| Model type | Share of gateway tokens (%) |
|---|---|
| Open-weight | 78.4 |
| Closed | 21.6 |
Limitations: this is one gateway's single-day snapshot, published by Rauch himself on X with no third-party verification; the report itself notes Vercel benefits as models become easier to swap, and that Rauch is repositioning the gateway as the control plane between AI applications and models.
Source: RuntimeWire · Aligned News snapshot
04/10
Opus 5.2 and a likely Gemini 4 Pro are already live with no announcement
Next-generation frontier models are being routed into existing products and blind-test arenas before they are named.
On the evening of September 17, Anthropic cut its grey-rollout channel and accounts that had been routed to a new model dropped back, prompting concentrated discussion among developers; a day later the channel returned, expanding from Claude Code to Chat and Cowork. Around the same time a model labeled gemini-3.8-flash appeared on the third-party blind-test platform Arena. Users who drew it reported a pelican drawn with a full pedaling motion, a PS5 vector illustration that took ten minutes and thousands of lines of output, and a voxel pagoda that rotated in real time in the browser. The community concluded it was Google's next flagship Gemini 4 Pro, internally codenamed Argon; the label was then changed to gemini-3.7-flash. Three days earlier, developers inspecting API requests found that Claude Code showed Opus 5 in the interface while the backend model identifier already pointed to 5.2. Opus 5 shipped on July 24, with no 5.1 in between.
For contrast, launching a generation publicly is expensive: GPT-6 Astra was trained on more than 100,000 GPUs at the Stargate site in Texas and is priced at $10 per million input tokens and $50 per million output tokens, 2.5x the previous Sol generation.
Limitations: these rollouts come with no model card, no API identifier and no price list, and Anthropic has said nothing; the identity of gemini-3.8-flash is inferred by the community from output behavior and unconfirmed by Google; the evidence is developer request captures and blind-test observation, and vendors can roll any of it back.
Source: TMTPost (Chinese-language source)
Global AI news
05/10
Runway wants AI video generation to become a stream you steer live
Instead of prompting and waiting for a finished clip, users would describe a scene and watch it stream.
Runway shared its research into real-time video generation. Today's models work in separate steps: enter a prompt, wait seconds or minutes, get a finished video, start over if it's wrong. Runway says users repeatedly report losing the most time generating and revising, and it wants to minimize time to first frame and then stream video as users prompt. The base is GWM-1, the company's first General World Model, introduced in December 2025; it builds on Gen-4.5, generates video frame by frame, and accepts camera movements, robot commands or audio as controls.
The one latency figure in the report is a design target: a real-time model built with Nvidia and running on the Vera Rubin platform aims to deliver a first frame within 100 milliseconds. Runway first discussed the approach in March with Runway Characters and showed a research preview on X; weeks ago it demonstrated Solaris, which uses Gen-4.5 to generate user interfaces frame by frame in response to clicks or voice.
Limitations: this remains a research preview with no launch date, pricing or availability; the 100-millisecond figure is a design target, not a measured result.
Source: The Decoder · Runway research preview
06/10
Microsoft open-sources TauGrid, collapsing GPU cluster plumbing into one Helm install
Researchers submit workloads without learning Kubernetes, and platform teams stop maintaining the glue.
TauGrid is a cloud-native platform for managing, scheduling and monitoring AI workloads on GPU-equipped Kubernetes clusters, covering data preparation, distributed training, fine-tuning and inference. It bundles Kueue (queueing and quota management), KubeRay (orchestration), GPU node health monitoring and observability, replacing the submission scripts, queue wrappers, health checks and result retrieval that platform teams otherwise assemble themselves — all through a single Helm install with clear ownership boundaries. Workloads are defined in tau.yaml and submitted with tau run, which validates the config and creates a Kubernetes Job or KubeRay RayJob, then lets Kueue queue it by remaining quota and priority. Failed jobs can resume from checkpoints, and status, logs and experiment evidence are retained for reproduction and diagnosis. The codebase is mainly Go.
Limitations: TauGrid is still under development, and its roadmap lists multi-tenant workspaces, RBAC and quotas, PyTorch DDP/FSDP and DeepSpeed, LoRA/QLoRA workflows, dataset lifecycle management and multi-cluster/multi-cloud execution as planned rather than shipped; it requires Kubernetes 1.30+, kubectl and Helm 3.0 or later; comparable platforms already exist, including Kubeflow and Nvidia Run:AI.

Image source: GitHub; mirrored on Jiufeng R2.
Source: TauGrid roadmap · InfoQ China
07/10
Microsoft and the University of Illinois train AI tutors on simulated students that make mistakes
Digital replicas of individual learners give tutoring systems fast feedback without waiting on real students.
StudentSim builds a separate digital replica for each student even when very few records of that person's work exist, and uses those replicas in place of real learners for rapid feedback. Training an AI tutor with a large, diverse group of students is "prohibitively expensive and time-consuming," the authors write in their paper, which is why improvements to tutors have lagged advances in AI models. The method was tested across 60 students on subjects including chess, and the code is on GitHub at microsoft/StudentSim.
The paper's abstract reports these scores on chess:
| Method | Behavioral fit F | Responsiveness R |
|---|---|---|
| StudentSim | 0.51 | 0.91 |
| GPT-5.4 | 0.23 | 0.72 |
| Maia2 | 0.45 | 0.27 |
Limitations: those scores cover chess only; the sources give no figure for how much the tutors improved as a result; this is a paper plus code, not a deployable product.
Source: The Decoder · Paper · GitHub
08/10
Google ships Agent Development Kit for Kotlin 1.0 at parity with Python
ADK for Kotlin closes the gap with the Python version and supports on-device models.
Google released Agent Development Kit (ADK) for Kotlin 1.0, a production-ready framework for building AI agents across Kotlin, Android and JVM/server applications. The release brings Kotlin to feature parity with Python and adds support for on-device AI.
Limitations: the report covers positioning and scope only, with no performance data, supported-model list or licensing details; "feature parity" is the vendor's framing and the report provides no item-by-item comparison.
Source: InfoQ
Regional and early signals
09/10
APUS open-sources a Jev reproduction that runs browser decisions on a local 9B model
A third-party reimplementation derived from public docs, packaged as an offline browser-automation skill under MIT.
On September 19, the AI lab of APUS (Kylin Hesheng Network Technology) published an independent open-source reproduction of Jev, packaged as a ready-to-run Agent Skill called fast-browser-use. It supports macOS, Linux and Windows, runs on GPU servers as well as Macs and PCs without a GPU, and is released entirely under MIT. It reproduces single-token logits decisions that skip autoregressive decoding, plus KV-cache broadcasting and concurrent batch evaluation. For browser automation it collects the page's genuinely visible, interactive elements into a numbered candidate action set, and a locally running Qwen3.5-9B decides where to click and what to select in one forward pass, which structurally removes the possibility of the model hallucinating a selector or a malformed output.
APUS reports that on an Apple M2 Pro consumer laptop, a fully offline real Wikipedia retrieval task took a median of about 18 seconds, while form filling and in-site navigation took about 3 seconds, with 4 scoring calls per task and no cloud calls or API fees.
Limitations: this is a third-party reproduction inferred from Jev's public documentation, not an official implementation, and TypeSafe has not disclosed Jev's architecture; the timings are APUS's own measurements with no independent verification, and APUS itself notes Jev's prior performance numbers all came from TypeSafe's own testing; coverage is a single Chinese-language report.
Source: QbitAI (Chinese-language source)
10/10
Spirit AI's Moz1 wheeled humanoids are running on CATL and JD.com lines
The founder puts a GPT-3-style moment for natural-language robot brains in mid-2027.
Pandaily reports that Spirit AI's Moz1 wheeled humanoid robots are deployed on production lines at CATL and JD.com, and that founder Gao Yang expects a GPT-3-style breakthrough for natural-language robot brains around the middle of 2027.
Limitations: this is summary-level post-launch coverage with no unit counts, station types, cycle times or acceptance metrics; the mid-2027 date is a founder's forecast rather than a roadmap commitment; there is a single report and no primary announcement to check against.
Source: Pandaily
Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.
Generate free

