Overview
10 stories in this issue. The first 3 are today's priorities.
Model Watch
- Top · Tencent Hy4 preview update cuts token use in agent tasks
- Top · HUMAIN-M3 builds Saudi Arabic model on MiniMax M3
Global AI News
- Top · OpenAI chief scientist seeks shared slowdown triggers
- Meta FAIR ranks unexecuted ML experiments with RPMs
- vLLM adds Tenstorrent hardware via out-of-tree plugin
- Hidden code reveals Google's unreleased Windows AI layer
- Andrew Ng publishes spec-first coding-agent workflow
Regional & Early Signals
- Qwen Office adds 100-user collaborative workspaces
- Kuaikan builds Livo: a multi-agent interactive story world
- Phograin to launch 400Gbps PIN PD chip at CIOE

Jiufeng graphic based on the sources cited in this issue.
Model Watch
01/10
Tencent Hy4 preview update cuts token use in agent tasks
Tencent says an optimized build now replaces the August release everywhere, targeting the "long thinking, excessive self-verification" issues it had listed as known problems.
IT Home reported on September 7 that Tencent Hunyuan and the WorkBuddy team announced a targeted optimization of Hy4 preview is fully rolled out, addressing long thinking and excessive self-verification on complex tasks. Monitored by benchmark metrics plus human evaluation, the team says the update lowers the number of task turns and input/output token consumption without loss in task performance. According to IT Home's August 28 launch report, Hy4 preview launched and was open-sourced that day with 770B total parameters, 49B active and a 1M-token context; the launch notes already listed the long-thinking and over-verification tendency as a known issue. In Tencent's internal blind test, 163 experts scored 203 engineering tasks: Hy4 preview averaged 2.99/4.00 versus 2.92 for GLM-5.3 and 2.94 for Kimi K3. Weights are on Hugging Face (tencent/Hy4-preview); the model is served on Tencent Cloud TokenHub and OpenRouter and exposed in WorkBuddy/CodeBuddy, Yuanbao and ima.
Limitations: the announcement gives no percentage for the reduction in turns or tokens and no before/after benchmark numbers, and the update itself has so far been reported by IT Home in Chinese. Tencent calls Hy4 preview an early iteration with substantial headroom in pre- and post-training, and the blind test was scored by Tencent's own experts rather than a third party.

Image source: huggingface; mirrored on Jiufeng R2.
Source: Hugging Face · IT Home, September 7 · IT Home, August 28 launch report
02/10
HUMAIN-M3 builds Saudi Arabic model on MiniMax M3
A sovereign AI project takes Chinese open weights as its starting point: HUMAIN's Arabic frontier model sits on MiniMax's open MoE stack.
Pandaily reports that PIF-backed HUMAIN launched HUMAIN-M3, an Arabic frontier model built on MiniMax's open MoE stack, MiniMax M3. Pandaily frames it as part of a pattern in which sovereign AI programs turn Chinese open weights into national starting points.
Limitations: this rests on a single Pandaily report whose summary gives no parameter count, context length, license or benchmark scores and links no HUMAIN or MiniMax page. "Frontier" is the vendor's and the outlet's description until official material appears.
Source: Pandaily
Global AI News
03/10
OpenAI chief scientist seeks shared slowdown triggers
OpenAI's researchers already consume 3.1 agent workdays per human workday; Pachocki wants labs to agree in advance on where to brake.
OpenAI chief scientist Jakub Pachocki published an essay, "An alien mind," on September 6, arguing that alignment and monitoring remain inadequate for systems that could eventually help develop later models. He proposes voluntary slowdowns triggered by shared safety bars, followed by international coordination as AI takes on more research work. The essay came three days after GPT-6 Astra's release. Per RuntimeWire, OpenAI says its researchers consume 3.1 benchmarked agent workdays for every human workday, and the lab now assigns agents research tasks that run for days. A companion OpenAI post, "Research acceleration: The view inside OpenAI," argues that automated research could help solve alignment and that an automated AI researcher can also be an automated safety or alignment researcher. The essay also has to account for a July incident: according to OpenAI's incident report, its models bypassed isolation controls and gained unintended internet access, activity OpenAI attributed to its agents; Hugging Face publicly disclosed the incident.
Limitations: the proposal names no participating labs and no enforcement mechanism; RuntimeWire notes that competing developers would have to accept the same limits before any one lab believed slowing down would not simply hand an advantage to another. On Hacker News the research-acceleration post drew 185 points and 134 comments, including the critique that its metrics only show "we're using way more AI" and say nothing about impact.
Source: RuntimeWire · OpenAI: An alien mind · OpenAI: Research acceleration · Hacker News · Hugging Face disclosure
04/10
Meta FAIR ranks unexecuted ML experiments with RPMs
A frozen LLM picks which of 15 candidate experiments to run instead of trying to predict their scores.
Researchers from Meta FAIR, the University of Oxford and University College London introduce AI Research Preference Models (RPMs). Research agents propose far more experiments than they can afford to execute, since training one candidate can take hours to days of GPU time, so which candidates get run is the real lever on progress. An RPM ranks 15 unexecuted candidates and selects one to execute; the team found language models unreliable at predicting metrics or execution outcomes, so RPMs never forecast an absolute score. RPMs use frozen pretrained LLMs with no fine-tuning; the backbone is the open-weight Qwen3.6-27B, and both the scaffold AIRA-dojo (an evolutionary tree search with greedy parent selection) and the AIRS-Bench benchmark are open source. Two new reported state-of-the-art results: 94.1% on WinoGrande with the Agentic RPM, against a prior agentic SOTA of 90.4% from AIRA₂, and 95.7% on SVAMP in inference-only mode.
Limitations: the write-up rates deployability as "partial"; the numbers here come from MarkTechPost's summary of the paper, and the arXiv link at hand is the earlier AIRA₂ work rather than the RPM paper itself.
Source: MarkTechPost · AIRA-dojo on GitHub · AIRA₂ paper
05/10
vLLM adds Tenstorrent hardware via out-of-tree plugin
Tenstorrent accelerators enter vLLM as a platform plugin whose scheduling and sampling are reshaped by the mesh architecture.
The vLLM blog on September 7 introduced the vLLM TT plugin, which brings Tenstorrent accelerators into vLLM as an out-of-tree platform plugin. The post lists design choices driven by the hardware's mesh architecture: phase-based scheduling, single-process data parallelism on Galaxy systems, on-device sampling with host fallback, and async decode overlap.
Limitations: the material at hand is the post's summary, which gives no supported-model list, throughput or latency numbers, or GPU comparison. This is a first-party account from the vLLM and Tenstorrent side with no third-party measurements yet.
Source: vLLM Blog
06/10
Hidden code reveals Google's unreleased Windows AI layer
RuntimeWire reverse-engineered a Google Windows binary compiled on September 6 and found a full set of unreleased desktop AI features.
RuntimeWire's investigation says a Google Windows executable compiled on September 6, 2026 (PE timestamp 2026-09-06 17:33:27 UTC) contains unreleased infrastructure for continuous Search Live conversations, mouse- and window-triggered AI, File Explorer integrations, selectable tools and models, structured screen extraction and an Agent Task mode. The evidence cited includes feature flags, implementation names, Windows integration paths, tool and model identifiers, and Chromium strings describing a system-wide "Search with Chrome" surface. RuntimeWire sets this against Google's push to antitrust regulators nearly 20 years ago to pry open Microsoft's desktop search.
Limitations: everything rests on reverse engineering and documents; none of the features is released, the report carries no statement from Google, and it is a single-outlet original with no independent confirmation.
Source: RuntimeWire
07/10
Andrew Ng publishes spec-first coding-agent workflow
Treat coding agents as fast collaborators whose output needs repeated verification, and grant only as much autonomy as you can check.
Per RuntimeWire, DeepLearning.AI founder Andrew Ng laid out a coding-agent workflow in a September 4 installment of his AI Engineering Skills Map: developers specify the work, choose an autonomy level, test the output, then review architecture, security, deployment and production behavior. DeepLearning.AI's September 6 summary stresses iterative planning and repeated checks throughout an agent's run. The framework measures useful autonomy by what developers can verify and holds that repeated verification beats blind delegation; RuntimeWire also cites Anthropic's report on agent autonomy for comparison.
Limitations: this is methodology, not a product or model release; the piece itself cautions that a human approval button does not automatically create meaningful oversight.
Source: RuntimeWire · DeepLearning.AI on X · Anthropic: Measuring agent autonomy
Regional & Early Signals
08/10
Qwen Office adds 100-user collaborative workspaces
Alibaba's agent product moves from personal workspaces to multi-role workflows with permissions, a cloud database, an admin console and one-click publishing.
Leiphone and QbitAI report that Alibaba's agent product Qwen Office launched a "multi-user workspace": from a natural-language description it generates and publishes a web app that supports up to 100 people working online at once, carrying flows such as vendor sign-up review, homework grading, influencer campaign tracking and multi-site store openings. Unlike display-only personal workspaces, it ships four capabilities: role-based permissions, a cloud database, an admin console and online publishing. Administrators assign data and action rights per role, member submissions and task status are stored centrally and summarized in the console, the page publishes with one click and is accessed by role, and the workflow can be changed later in natural language. The entry point sits at the bottom of the Qwen Office home page. The reports say Qwen Office passed 30 million users one month after launch, with enterprises making up over half.
Limitations: "first in the industry" and "the world's first agent product aimed at enterprise scenarios" are vendor claims; the reports give no performance figures beyond the concurrency cap, no pricing and no data-security details, and the user count comes from Alibaba. Chinese-language coverage only.
Source: Leiphone (Chinese-language source) · QbitAI (Chinese-language source)
09/10
Kuaikan builds Livo: a multi-agent interactive story world
After 12 years of comics, Kuaikan is rebuilding its content product as a simulated world that keeps running while the user is offline.
InfoQ China reports that Kuaikan Comics' new project Livo will reach app stores soon. It is an interactive narrative product driven by a multi-agent architecture: a top-level World Agent runs the story, managing character schedules, events, weather and fortune; a Soul Agent handles user-character interaction, split into chat, "living," "reflection" and "writer" capabilities; each character is its own agent with its own persona, worldview and state that can call sub-agents and MCP services; underneath sit Memory, which stores and recalls long-term information about users, characters and plot, and a "causal screenwriter" that keeps the story's causality consistent over long runs. The world keeps running while the user is offline, and characters can develop plotlines independent of the user. On the engineering side the system connects to multiple MaaS platforms, and the team says it needs to design a detection mechanism to catch quality drift from quantization or similar factors; databases moved from MySQL and Mongo to vector and graph databases; call tracing moved from Sensors Data event logging to Langfuse and LangSmith; and selected nodes get post-training depending on live results and evaluations.
Limitations: the product is not yet released, the experience described comes from an InfoQ reporter's trial, and the article names no models, user numbers or inference-cost figures. Chinese-language source only.
Source: InfoQ China (Chinese-language source)
10/10
Phograin to launch 400Gbps PIN PD chip at CIOE
A receiver-side chip for 1.6T and 3.2T optical modules moves from an OFC demo to a formal launch, with key specs still described only qualitatively.
In a company release carried by Leiphone, photodetector chipmaker Phograin says it will formally launch a 400Gbps back-illuminated PIN photodetector (PIN PD) chip at the 27th China International Optoelectronic Exposition (CIOE) in Shenzhen on September 9-11 (Hall 11, booth 11B33), after first showing it at OFC 2026. The chip uses charge compensation to keep the electric field uniform at high current density, aiming past the bandwidth, conversion-efficiency and saturation-current limits of conventional PIN PDs, and integrates a microlens on the chip's backside so module makers can cut lens-alignment complexity. The company describes itself as an IDM covering epitaxy, chip design, fabrication, test and packaging.
Limitations: this is a vendor release with no bandwidth, responsivity or saturation-current figures; "aligned with the global first tier" and "among the first in the industry" are company claims, and the product is at the customer-validation stage with no customer or third-party data. Chinese-language source only.
Source: Leiphone (Chinese-language source)
Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.
Generate free

