Overview
9 stories in this issue. The first 3 are today's priorities.
Trending model updates
- Top · Qwen3.8-Max leads open-weight models on commerce agent bench
- Top · MiniMax H3 hits generation faster than playback on vLLM-Omni
- Top · A Bedrock error hints Anthropic is staging Claude Fable 5.1
Global AI news 4. Gradium's new default TTS: 81% hard-case pass rate, 216 ms first audio 5. DataAgent raises $10M to fix production faults inside Kubernetes 6. Tencent's Marvis opens custom models: plug in Kimi, Zhipu GLM, DeepSeek
Regional & early signals 7. Tsinghua AIR's self-evolving Zeva lifts success rate from 26% to 73% 8. Leiphone dissects 11,000 DeepSeek Harness plugins: almost no governance 9. Lovart's update centers on the design workflow: a clipper and 100+ skills

Jiufeng graphic based on the sources cited in this issue.
Trending model updates
Qwen3.8-Max leads open-weight models on commerce agent bench
Alibaba's Accio team says Qwen3.8-Max is the strongest open-weight model on its stateful commerce benchmark — but the scores are self-reported.
Yukun Lian, Sicong Xie and ten-plus colleagues on Alibaba International's Accio team built Commerce Agent Bench to test whether agents can finish commercial work inside stateful replicas of business software, scoring completed business state rather than transcripts. Qwen3.8-Max (part of the Qwen3.8 family) completed 56 of 107 tasks through Accio, 47 through OpenClaw and 53 through Pi; Alibaba's official Qwen account called it the strongest overall result among open-weight models. The repository's three result tables put Qwen3.8-Max two aggregate passes ahead of DeepSeek V4 Pro, while closed model Claude Opus 5 finished 35 passes ahead of Qwen.
Limitations: the figures come from Alibaba's own self-run benchmark, so this is a published result, not an independent audit. RuntimeWire notes harness variance and the absence of immutable, public task-level bundles, which weakens auditability.

Image source: GitHub; mirrored on Jiufeng R2.
Source: RuntimeWire · Qwen on X · Commerce Agent Bench
MiniMax H3 hits generation faster than playback on vLLM-Omni
vLLM's official blog says system-wide optimization plus FastVideo's four-step FastH3 lets MiniMax H3 generate faster than playback.
vLLM-Omni optimizes and scales the full MiniMax H3 stack, then integrates FastVideo's four-step FastH3, so the MiniMax H3 stack generates faster than playback for real-time serving. The write-up is a first-party engineering post on the vLLM blog.
Limitations: the post centers on the serving stack and optimization path; the public summary gives no concrete throughput (tok/s), latency or hardware numbers, and "faster than playback" is a qualitative claim to be read against the full post.
Source: vLLM Blog
A Bedrock error hints Anthropic is staging Claude Fable 5.1
An unlisted Fable 5.1 slug returns a 404 on Bedrock instead of an invalid-ID error — an early staging clue, not proof of imminent GA.
Per an API response surfaced on August 31, researcher Leo (@synthwavedd) queried Amazon Bedrock with a Fable 5.1 model slug and got a 404 "Model not found," while an Opus 5.1 slug and a deliberately nonsensical Anthropic identifier both returned a 400 "provided model identifier is invalid." AWS says an incorrect Bedrock model identifier normally yields that validation error; a 404 can indicate the service recognizes the resource shape but cannot serve the model to that account, endpoint or region. Bedrock's public catalog still points to Fable 5.
Limitations: this is a single error-code difference. Leo says this "usually happens in the days before release," while RuntimeWire is explicit that the clue does not prove Fable 5.1 is ready for general availability; Anthropic has not confirmed anything.
Source: RuntimeWire · @synthwavedd on X · Bedrock model card
Global AI news
Gradium's new default TTS: 81% hard-case pass rate, 216 ms first audio
Gradium made a new TTS its API and Studio default, reporting an 81.0% hard-case pass rate and 216 ms time-to-first-audio, and open-sourced the eval set.
Gradium AI switched a new speech model on as the default across its API and Studio on August 31; existing voices, including custom clones, keep working with no migration. On a 500-sentence hard-case set spanning five languages (EN, DE, FR, ES, PT), it reports an 81.0% human-rated pass rate, ahead of Cartesia Sonic 3.6 at 75.1% and ElevenLabs v3 Conversational at 65.4%; time to first audio is 216 ms at P50 on Coval, 170 ms faster than the model it replaces. The evaluation set is open-sourced on Hugging Face under CC BY 4.0, with 100 items across 10 criteria in five languages covering spelling, acronyms, alphanumeric tokens, dates, numbers and email.
Limitations: pass rate and latency are Gradium's own numbers, and competitors' scores were measured by Gradium on the same set; the eval set is open, but the leaderboard is still a vendor self-report.
Source: MarkTechPost · Hugging Face dataset
DataAgent raises $10M to fix production faults inside Kubernetes
Israeli startup DataAgent launched with $10M in pre-seed funding and agents that repair production faults inside a customer's own Kubernetes clusters.
DataAgent formally launched with $10 million in pre-seed funding for a platform that repairs production faults inside a customer's own Kubernetes clusters, spending the money to accelerate North American adoption. CEO Ishay Yaari and CTO Nati Shalom founded the company in January; both previously worked at open-source cloud-orchestration company Cloudify, which Dell acquired in 2023 for a reported $100 million. Unlike conventional observability that detects a fault, hands it to an engineer and ships a second copy of logs and metrics to a vendor cloud, DataAgent's agents read live system state and topology and act in place.
Limitations: the positioning and cost claims are the company's own. The story cites a Grafana Labs 2025 survey that such observability spend averages 17% of total compute-infrastructure spending (most commonly 10%) as industry context, not DataAgent's measured results.
Source: SiliconANGLE
Tencent's Marvis opens custom models: plug in Kimi, Zhipu GLM, DeepSeek
Tencent's OS-level assistant Marvis added custom models on September 1, letting users connect third-party models or run a local one.
Tencent's operating-system-level AI assistant Marvis rolled out a custom-model feature on September 1, letting users connect third-party models — including Kimi, Zhipu GLM, DeepSeek and MiniMax — or run a local model, shifting Marvis from a fixed backend to a user-chosen one.
Limitations: only Pandaily has reported this so far; the summary gives no detail on integration method, pricing, availability or local-model hardware requirements, so treat it as an early product signal.
Source: Pandaily
Regional & early signals
Tsinghua AIR's self-evolving Zeva lifts success rate from 26% to 73%
With weights frozen, Zeva learns from its own interactions via "in-context causal learning," lifting cumulative success across embodied benchmarks from 26% to 73%.
Tsinghua AIR and Yuanbianhuan released Zeva, described as the first embodied WAM to achieve in-context causal learning (ICCL): learning physical causality from the consequences of its own actions without updating weights. The paper reports lifting a frozen model's cumulative success rate from 26% to 73% across several standard embodied benchmarks, with a real chemistry-lab robotic arm growing steadier over repeated attempts. It self-evolves at deployment via Causal Interaction Extraction, a dual-timescale causal memory and in-context policy injection.
Limitations: results come from the team's paper as reported by QbitAI, with no third-party reproduction yet; beyond cumulative success rate, the summary gives little cross-task or cross-embodiment quantification. Chinese-language source.
Source: QbitAI
Leiphone dissects 11,000 DeepSeek Harness plugins: almost no governance
After DeepSeek Harness went open source, 11,000+ GitHub repos appeared under the dsh-plugin tag; Leiphone's hands-on test found almost no official plugin governance.
After DeepSeek Harness was open-sourced, its "everything is a plugin" slogan spread fast. Leiphone says GitHub now hosts more than 11,000 repositories under the dsh-plugin tag; it counted through them, installed a probe plugin to test them, and ran more than twenty of the most-downloaded plugins overnight, concluding that DeepSeek has established almost no plugin-governance mechanism.
Limitations: this is a single-source hands-on investigation by Leiphone; the summary discloses no specific vulnerability samples, affected-plugin list or quantified risk data — only the governance gap itself. Chinese-language source.
Source: Leiphone
Lovart's update centers on the design workflow: a clipper and 100+ skills
Lovart quietly shipped a big update, shifting focus from generation quality to the design workflow with a browser clipper and a 100+-skill marketplace.
In a hands-on by ifanr, Lovart's update reorients around the work surrounding design: a new "inspiration clipper" browser extension one-click saves and auto-categorizes assets from Pinterest, Xiaohongshu, Taobao, Amazon and Instagram into the user's account; during generation the agent auto-searches for reference material, with Pinterest and Google Images added as sources. It also adds a skills marketplace with 100+ official and community skills across five categories: marketing assets, brand visuals, social content, product design and creative styles.
Limitations: this is a single hands-on review by ifanr, with no model version, quantified quality evaluation or pricing disclosed. Chinese-language source.
Source: ifanr
Get the latest AI model insights and tutorials from Jiufeng.
Explore more


