Overview
10 stories in this issue. The first 3 are today's priorities.
Hot Model Watch
- Top · SpaceXAI ships Grok 4.6, tying GPT-5.6 on the AA Index
- Top · DeepSeek raises V4 API prices on August 16 with a peak/off-peak rate card
- Top · Claude's text watermark goes live — and a developer already ships a Mac remover
- WeChat gray-tests native assistant "Xiaowei," with WeLM reportedly grown to 617B parameters
Global AI News 5. OpenAI starts rolling out giftable ChatGPT credits 6. 25 frontier-lab researchers warned on automated AI research — several milestones have already fallen 7. AllenAI opens OlmoEarth embedding exports, with code and weights released 8. OneAdvanced self-hosts Llama 4 on UK-sovereign AWS, deploying over 50 agents
Regional & Early Signals 9. BAAI's FlagOS completes Day0 adaptation of Qwen3.8-2.4T across nine chips (Chinese-language source) 10. Brex describes a "runtime-independent" AI workflow pattern

Jiufeng graphic based on the sources cited in this issue.
Hot Model Watch
SpaceXAI ships Grok 4.6, tying GPT-5.6 on the AA Index
A post-training upgrade pushes Grok 4.6 into the frontier tier — 61 points, level with GPT-5.6 Sol, at more than 60% lower price.
SpaceXAI (xAI) released Grok 4.6 on August 12. Per The Decoder, it scores 61 on the Artificial Analysis Intelligence Index, up five points from Grok 4.5, tying OpenAI's GPT-5.6 Sol and trailing only Anthropic's Claude Opus 5 (63) and Claude Fable 5 (62). On the GDPval-AA v2 benchmark for real-world knowledge work it ranks second with an Elo of 1,753, finishing complex tasks in about 53 steps versus Opus 5's roughly 103. Pricing is $2/$6 per million input/output tokens — more than 60% cheaper than Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30) — and the first week doubles usage quotas in Grok Build and Cursor. MarkTechPost notes it is a post-training upgrade on an unchanged Grok 4.5 base — regenerated supervised fine-tuning trajectories plus reinforcement learning in agentic environments — and adds a new "xhigh" reasoning-effort level.
The model is generally available via the xAI API (grok-4.6), is the default in Grok Build, ships in Cursor on all plans, and is routable through OpenRouter, Vercel, and Cloudflare. The limits are explicit: there is no open-weights release and no self-hosting path, so air-gapped deployments are out, and this is a post-training iteration rather than a larger new base model.

Image source: spacexai; mirrored on Jiufeng R2.
Source: The Decoder · MarkTechPost · xAI docs
DeepSeek raises V4 API prices on August 16 with a peak/off-peak rate card
The new schedule splits pricing by UTC hour — off-peak is half of peak, yet every off-peak rate still sits above today's price.
Per RuntimeWire, Liang Wenfeng's DeepSeek will raise V4-model API prices from 16:00 UTC on August 16, introducing a two-tier rate card that nudges flexible compute jobs out of its busiest hours. On X, DeepSeek defines peak hours as 01:00–04:00 UTC and 06:00–10:00 UTC — seven hours a day — with the other 17 hours off-peak, and says off-peak sits 50% below peak. But that gap is only between the two new tiers: every off-peak rate in the table is higher than DeepSeek charges today. For V4-Flash, current prices of $0.0028 (cache-hit input), $0.14 (cache-miss input) and $0.28 (output) per million tokens rise off-peak to $0.007, $0.22 and $0.66; off-peak output rates climb 136% for V4-Flash and 128% for V4-Pro.
In other words, even the cheapest off-peak window costs more than the current price — the "50% lower" framing is relative to the new peak tier, not a cut versus today. RuntimeWire advises scheduled workloads to revisit their unit economics before August 16.
Source: RuntimeWire · DeepSeek on X
Claude's text watermark goes live — and a developer already ships a Mac remover
Anthropic is applying EU-driven machine-readable marks to Claude output worldwide; a local rewriting tool arrived at once, but the detector needed to verify it is unpublished.
Ars Technica reports Anthropic is embedding machine-readable watermarks into Claude-generated text — invisible for now, and flagging anything Claude processed, even human writing it only edited. RuntimeWire adds that Brian Roemmele (@BrianRoemmele), founder of ReadMultiplex, published a Mac-only tool and local rewriting workflow on August 12 called "AI Watermarks Cleaner," paired with LM Studio for running models locally. By his description the workflow relies on local rewriting rather than simply deleting metadata or invisible Unicode characters, separating visible formatting before rewriting the body text.
The limitation: Anthropic has not published the detector needed to check for its marks, so the cleaner's effectiveness cannot be independently verified — and Ars stresses the watermark is invisible "for now," leaving room for that to change.
Source: Ars Technica · RuntimeWire · Brian Roemmele on X
WeChat gray-tests native assistant "Xiaowei," with WeLM reportedly grown to 617B parameters
Pandaily says the WeChat AI team's WeLM has moved to a sparse MoE architecture and scaled to 617B parameters — still in gray testing.
Per Pandaily, WeChat has begun gray-testing "Xiaowei," a native AI assistant embedded inside the app, powered by WeLM, the WeChat AI team's long-developed large language model. The report says WeLM has shifted from a dense design to a sparse MoE architecture and quietly grown to 617 billion parameters.
A caveat: this is a thin-material item. The 617B figure comes from Pandaily's own decoding of the model rather than an official spec — WeChat has not published WeLM's parameters — and Xiaowei is only in gray testing, not general availability.
Source: Pandaily
Global AI News
OpenAI starts rolling out giftable ChatGPT credits
One-time claim links let you buy prepaid usage for someone else; entry points are already shipped in the Windows app.
A RuntimeWire investigation finds OpenAI is gradually rolling out a way for personal ChatGPT users to buy credits for another person via one-time claim links. OpenAI has published a live Help Center article covering purchase, delivery, redemption, expiration and refund rules; separately, RuntimeWire examined the Electron archive in OpenAI's Windows app and found a dedicated gift-credit module, a direct purchase route at https://chatgpt.com/gifts/credits, remote controls named purchase_flow_enabled and desktop_beacon_enabled, and UI entry points on the home screen, profile and Usage settings. Product researcher Tibor Blaho flagged the live page on X.
Limitations: the feature is still rolling out — it was not enabled on RuntimeWire's own account, which had to force local display conditions to render the "Gift credits" banner; once redeemed, credits are added to the recipient's personal ChatGPT balance.
Source: RuntimeWire · OpenAI Help Center · Tibor Blaho on X
25 frontier-lab researchers warned on automated AI research — several milestones have already fallen
An interview study finds 20 of 25 respondents rank "automating AI research" among the most severe, urgent risks, and some milestones they named have already been hit.
The Decoder reports that IAPS fellow Severin Field interviewed 25 researchers from OpenAI, Anthropic, Google DeepMind, Meta and US universities about recursive self-improvement (RSI) — a system capable enough at AI development to build a stronger version of itself, which then does the same. Twenty of the 25 rated automating AI research as one of the most severe and urgent AI risks. Respondents repeatedly cited the nonprofit METR's Task Horizon benchmark as their go-to progress measure: the length of tasks AI agents can complete on their own has been doubling roughly every six months. Summing up his late-summer-2025 study on his blog, Field writes that several of the named milestones have already been passed and RSI can no longer be dismissed as marketing hype.
Limitations: this is a qualitative study of 25 interviews plus a retrospective blog post; the definition and urgency of RSI remain contested in the field.
Source: The Decoder · Interview study (arXiv)
AllenAI opens OlmoEarth embedding exports, with code and weights released
The OlmoEarth Earth-observation foundation model can now compute and export embedding vectors from Studio, with source code, weights and paper all public.
On the Hugging Face blog, AllenAI (Ai2) introduced OlmoEarth embeddings: its Earth-observation modeling platform OlmoEarth Studio can now compute and export embedding vectors — compact numerical representations of Earth-observation data produced by the open-source OlmoEarth foundation models. The embeddings are a low-cost entry point to OlmoEarth, supporting downstream tasks from similarity search to segmentation to unsupervised exploration; locations with similar surface characteristics end up with similar vectors, while different ones land far apart. The team says the source code and model weights are public alongside the research paper, so the community can inspect exactly how the embeddings are generated.
Limitations: the performance claims come from AllenAI's own benchmarking plus one independent evaluation (paper), and still need broader third-party replication.
Source: Hugging Face blog · Source code (GitHub) · Independent evaluation (arXiv)
OneAdvanced self-hosts Llama 4 on UK-sovereign AWS, deploying over 50 agents
To keep data inside the UK, OneAdvanced self-hosts Llama 4 Maverick and Llama Guard 4 and runs more than 50 specialized agents.
An AWS Machine Learning Blog case study describes how OneAdvanced, a UK enterprise software provider serving over 10,000 customers, self-hosts the open-weight models Llama 4 Maverick and Llama Guard 4 on Amazon SageMaker AI over infrastructure it fully controls — because at the time those models weren't yet available through managed services in the UK region. The solution pairs a RAG pipeline on Amazon Aurora PostgreSQL with the pgvector extension, more than 50 specialized agents powered by the Strands Agents SDK, and a tool layer running on Amazon ECS, all to ensure no data leaves the UK.
Limitations: self-hosting was a forced choice (managed services were unavailable in the UK region at the time), and the account comes from AWS's own vendor blog.
Source: AWS Machine Learning Blog · pgvector (GitHub) · Llama (Hugging Face)
Regional & Early Signals
BAAI's FlagOS completes Day0 adaptation of Qwen3.8-2.4T across nine chips (Chinese-language source)
For the freshly open-sourced 2.4T MoE, the FlagOS community completed Day0 adaptation on nine domestic and foreign AI chips and added unified INT8 quantization.
Per Leiphone, BAAI's FlagOS community completed Day0 multi-chip adaptation of Alibaba's freshly open-sourced ultra-large MoE model Qwen3.8-2.4T-A95B, covering nine chips — T-Head, NVIDIA, Moore Threads, Huawei Ascend, MetaX, Kunlunxin, Hygon, Tsingmicro and Enflame — with BF16/FP8/INT8 precision options. To handle a parameter count about six times larger than the prior generation, the FlagOS stack added unified multi-chip INT8 quantization, using FlagOS-Compressor to open two compression paths from native FP8/BF16 weights to INT8. The community says it has now cumulatively completed cross-chip Day0 adaptation for 7 leading model teams, 12 mainstream open models and 10 chips.
Limitations: this is a single Chinese-language source; precision-alignment claims (that error versus NVIDIA's CUDA FP8 build stays within an aligned range) are self-reported by the community and not independently verified. The Qwen3.8-2.4T model release itself was reported separately; this item is about the chip adaptation, not model capability.
Source: Leiphone (Chinese-language source)
Brex describes a "runtime-independent" AI workflow pattern
Treating the durable runtime as a plugin behind an interface keeps production reliability and evaluation-iteration speed from dragging on each other.
InfoQ published Brex's engineering practice for AI workflows: the distinctive difficulty is that an LLM step's output quality drifts with every prompt tweak or model change and must be validated by offline evaluation — yet production durability needs a heavyweight, distributed, persistent runtime, while eval iteration needs a lightweight, in-process, ephemeral loop that can re-run hundreds of times in seconds. Those two demands conflict. Brex's answer is to abstract the runtime into a plugin behind an interface: the platform is written in TypeScript and maintained by five engineers, its worker processes run on Kubernetes and connect to Temporal Cloud to execute long-running agents, and agents reach LLMs through the Vercel AI SDK to a shared LLM gateway.
Limitations: this is a single company's (Brex) engineering pattern, and the authors note it only pays off when you have multiple workflows, real production-grade reliability requirements and a rigorous evaluation process.
Source: InfoQ (Chinese) · InfoQ (English original)
Get the latest AI model insights and tutorials from Jiufeng.
Explore more


