Overview
10 stories in this issue. The first 3 are today's priorities.
Hot Model Watch
- Top · Cognition adds Grok 4.6 to Devin's coding platform
- Top · Alibaba open-weights Qwen3.8-2.4T-A95B, its first Max-tier weights
- Top · DeepSeek V4 Pro goes live with a 1M-token tier
- Gemini loses market share as Anthropic climbs to 14.9%
Global AI 5. Google DeepMind ships sign-language-to-text model SL2T 6. Liquid AI releases edge vision model LFM2.5-VL-3B 7. Researchers reconstruct LLM prompts from output text 8. Twitch adds an opt-out from training Amazon's AI
Regional & Early Signals 9. Robot sorts 1,816 packages/hour in a livestream (Chinese-language source) 10. Tencent Cloud open-sources agent sandbox Cube Sandbox (Chinese-language source)

Jiufeng graphic based on the sources cited in this issue.
Hot Model Watch
Cognition adds Grok 4.6 to Devin's coding platform
Cognition folds xAI's Grok 4.6 into Devin, reinforcing its pitch that orchestration—not any single foundation model—is the durable product.
On August 12, Cognition added xAI's Grok 4.6 to Devin Desktop and Devin CLI, giving developers another foundation model for agent-driven work inside local codebases. The company announced the integration in a thread on X and published a short technical note describing Grok 4.6 as particularly strong at exploring repositories and diagnosing root causes before editing code, and credited it with following repository conventions and testing its changes. Led by Scott Wu, Cognition says its founding team collectively holds 10 International Olympiad in Informatics gold medals.
Cognition positions Devin as a model-neutral control layer; the capability description comes from Cognition itself, with no third-party benchmark yet, and no pricing or quota details were disclosed at launch.
Source: RuntimeWire · Cognition on X
Alibaba open-weights Qwen3.8-2.4T-A95B, its first Max-tier weights
A Qwen-Max-class model gets open weights: 2.4T MoE, 95B active per token, native 256K context extendable to ~1M.
Late on August 12, Alibaba's Qwen team released the weights of Qwen3.8-2.4T-A95B on the ModelScope community: 2.4T total parameters, 95B activated per token, native 262,144 (256K) token context, extendable to 1,010,000 tokens. Alibaba says this is the first time a Qwen-Max-class model has been open-weighted.
The model page states parameter and context specs; license terms, benchmark scores and inference throughput should be confirmed on the official ModelScope page, and no independent third-party evaluation is available yet.
Source: ModelScope
DeepSeek V4 Pro goes live with a 1M-token tier
DeepSeek-V4-Pro-0813 lands on the API with a one-million-token context and agent tooling, priced 3x the Flash tier.
DeepSeek rolled out DeepSeek-V4-Pro-0813 on its API platform, pairing a one-million-token context window with explicit agent-developer tooling at a 3x price premium over the Flash tier. Pandaily reports the release "closes in on the frontier tier."
"Closes in on frontier" is Pandaily's framing; the report cites no head-to-head benchmark scores against named frontier models, and this remains a single-outlet report without independent verification.
Source: Pandaily
Gemini loses market share as Anthropic climbs to 14.9%
Three datasets agree: Pangram shows Gemini falling from 12% to 1.9%, OpenAI holding above 50%, and Anthropic rising from 4.3% to 14.9%.
Per The Decoder's roundup of three data points: text-detection platform Pangram, measuring submitted texts, finds OpenAI above 50% every month since tracking began, Anthropic up from 4.3% to 14.9% (mainly in technical and scientific fields), and Gemini down from 12% to 1.9% in July 2026—what Pangram calls a "collapse," echoed by OpenRouter. Similarweb shows Google's website share dipping from 27% to 26.8% last month, up from 9.4% a year ago; Google separately claims 1 billion monthly Gemini App users.
The report spells out each source's blind spots: Pangram mostly measures writing tasks where Gemini has always trailed, OpenRouter skews toward open-weight models, and MAU is a "very low-quality signal." The datasets diverge in method and agree only that Gemini's share is shrinking.
Source: The Decoder · Similarweb
Global AI
Google DeepMind ships sign-language-to-text model SL2T
Sign language AI reaches consumer products for the first time, starting with American Sign Language to English.
Google DeepMind introduced a massively multilingual sign-language-to-text (SL2T) translation model, calling it a breakthrough in quality and generality and the first time sign language AI has shipped in consumer products: SL2T powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with ASL to English. DeepMind notes there are more than 200 sign languages and an estimated 70 million Deaf and hard-of-hearing people whom prior speech AI had not reached.
The launch covers only ASL-to-English; DeepMind says more devices and languages are "coming soon" but gives no timeline or accuracy figures for additional languages.

Image source: deepmind; mirrored on Jiufeng R2.
Source: Google DeepMind
Liquid AI releases edge vision model LFM2.5-VL-3B
A 3B vision-language model for on-device use, focused on screen/UI understanding, grounding and function calling.
Liquid AI released LFM2.5-VL-3B, a vision-language model that runs on your own hardware and pairs a SigLIP2 encoder. It lists four improvements over prior releases: screen/UI understanding across devices, grounding and object detection from natural-language queries, multi-image reasoning, and stronger function calling in text-only and vision-text settings; the model answers directly rather than reasoning, to keep responses fast on-device.
Liquid AI gives no numeric score for the function-calling gains; specific benchmarks and latency should be read from the Hugging Face model page, and a 3B edge model has different capability limits than larger models.
Source: Hugging Face
Researchers reconstruct LLM prompts from output text
IIT Bombay and Adobe Research's PTP inverse model rebuilds the original prompt from output alone.
Researchers at IIT Bombay and Adobe Research proposed "Previous-Token Prediction" (PTP): they train an inverse language model to predict the previous token, reconstructing the original prompt from an LLM's output text with near-perfect accuracy. The inverse model is trained from scratch on synthetic data generated by the target LLM, needs no access to model weights, and even applies to third-party models.
Because many prompts can yield similar outputs, such reconstruction had been considered impractical; the "near-perfect accuracy" is the paper's result under its own setup, and it remains an arXiv preprint awaiting peer replication.
Source: The Decoder · arXiv
Twitch adds an opt-out from training Amazon's AI
Users can now stop their streams, VODs, clips and chats from training Amazon's generative AI models.
Twitch added a toggle letting users opt out of having their content train Amazon's generative AI models. Per The Verge, opting out means "your streams, VODs, clips, stream chats, and pictures and text on your channel" won't be used in "future training" of an Amazon AI model meant to generate text, audio, images or video; other "AI-supported" features like captions and safety tools still work.
The opt-out only covers "future" training—per The Verge, opting out means your content won't be used in future training, so it applies going forward; the toggle itself reads "Allow your channel content to train generative AI content models at Amazon…," and turning it off is how you opt out. And if you chat on someone else's stream, their opt-out preference governs whether that chat can be used.
Source: The Verge · Twitch on X
Regional & Early Signals
Robot sorts 1,816 packages/hour in a livestream (Chinese-language source)
Vendor livestream claims a dual-arm-plus-gripper rig beats Figure AI's public record by 45%.
Per QbitAI, X Square Robot (自变量机器人) ran a public livestream in which a dual-arm, gripper-based rig sorted random irregular packages for a continuous hour at 1,816 items/hour, claiming to beat Figure AI's previously published 1,248/hour average by about 45%, at roughly 70% lower hardware cost than Figure's humanoid-plus-five-finger-hand setup. Packages included cardboard boxes, soft mailers, cylinders and foam-packed fresh goods in random poses, with its in-house WALL-B "unified world model" identifying items and matching grasp strategies in real time.
These are claims from the vendor's own livestream; the efficiency and cost comparisons come solely from X Square Robot, with no independent third-party verification, and this is a single Chinese-language source.
Source: QbitAI
Tencent Cloud open-sources agent sandbox Cube Sandbox (Chinese-language source)
A sub-100ms-startup agent execution environment, merged into a third-party provider stack in July.
Per InfoQ, Tencent Cloud open-sourced Cube Sandbox in April 2026 as an execution-environment substrate for AI agents, emphasizing sub-100-millisecond startup; its foundation evolved from a Serverless production system dating to around 2023, later moving through code execution, data analysis, agent RL and agent runtime. In July, OpenClaw creator Peter Steinberger submitted and merged a PR integrating Cube into his provider system, alongside sandboxes like E2B and Modal. The context is that terminal agents granted filesystem, browser and account permissions need an isolated, recoverable execution environment.
InfoQ frames Cube as a sample for observing agent-infrastructure evolution and gives no quantified performance comparison with E2B or Modal; this is a Chinese-language source, with the GitHub repository provided for verification.
Get the latest AI model insights and tutorials from Jiufeng.
Explore more


