Overview
10 stories in this issue. The first 3 are today's priorities.
Hot model updates
- Top · WeatherNext 3 adds hourly refresh and ships into Search and Gemini
- Top · Cloudflare wires GPT-5.6 Cyber into vulnerability discovery
- Top · Claude Fable 5.1 cracks a 1653 number cipher in 44 minutes
- Gemini voice controls arrive in Gmail, Docs and Keep
- Four major AI model services go down at almost the same time
Global AI news 6. Nvidia agrees to buy Hugging Face for $12.93 billion 7. Nvidia's PAIR pools idle home PCs for local inference 8. NeoMME releases 260M and 800M multimodal multilingual encoders
Regional and early signals 9. GLM-5.3-Flash overnight quota discount runs through September 20 10. Tencent's WorkBuddy platform opens with 30-plus hardware brands

Jiufeng graphic based on the sources cited in this issue.
Hot model updates
WeatherNext 3 adds hourly refresh and ships into Search and Gemini
Google DeepMind's flagship forecasting model now feeds Search, Gemini, Maps and Cloud.
Google DeepMind and Google Research released WeatherNext 3 on September 3. The announcement says the model adds real-time satellite data, hourly refreshes, higher resolution, more precise precipitation forecasting and clean-energy variables, and that it is already integrated across Search, Gemini, Maps, the Google Maps Platform and Google Cloud. Ferran Alet, weather science lead at Google DeepMind, worked on the project with researchers from both teams.
For comparison, the previous WeatherNext 2 — documented in the public repository — runs on a 0.25-degree global grid, roughly 30 kilometers at the equator, initializes every six hours and produces 64 ensemble members.
The evidence behind the claims differs by claim: the official blog cites Brightband's independent live evaluation when calling WeatherNext 3 the most advanced and accurate global weather AI model to date, and RuntimeWire notes that Brightband's Operational WeatherBench now includes WeatherNext 3, while the headline "improves CRPS by up to 60%" is a bounded figure — an upper bound measured against NASA IMERG precipitation observations.

Image source: Google; mirrored on Jiufeng R2.
Source: Google Blog · RuntimeWire · TechCrunch · GitHub
Cloudflare wires GPT-5.6 Cyber into vulnerability discovery
Edge traffic context plus an OpenAI model: find the flaw, check whether it is reachable, then propose both a patch and a temporary firewall rule.
Cloudflare opened invitation-only early access to Vulnerability Discovery and Remediation on September 3. The service sits inside Cloudflare Managed Defense and uses OpenAI's GPT-5.6 Cyber for reconnaissance, vulnerability hunting and validation. Cloudflare then adds network context — whether the affected code is deployed and which routes can reach it — and proposes software patches alongside temporary firewall rules. GPT-5.6 Cyber was introduced through OpenAI's restricted Daybreak program.
Availability is narrow: access is invitation-only, scanning is limited to customer-authorized code, humans approve every change, and Managed Defense staff help set the boundaries before the service reaches a broader market. The report is explicit that the approach only pays off if those checks hold up in production.
Source: RuntimeWire · OpenAI
Claude Fable 5.1 cracks a 1653 number cipher in 44 minutes
A 17th-century puzzle treated as unsolved fell — through brute persistence, the evaluator says, not better cryptanalysis.
The Decoder reports that Anthropic's Claude Fable 5.1 solved the "Cyphral Distich" from Sir Thomas Urquhart's 1653 publication: two lines of 32 numbers each. According to Vals AI, the model produced the answer in 44 minutes with no human help. Each number points to a word in one of the 32 sections of the book, and the first letters of those words spell out "O God uphold King Charls the Second and make him the supreme ruler of this land."
The author had spent months testing other frontier models on unsolved puzzles without a verifiable solution. This time the model was asked to find a solvable puzzle itself; it sifted through candidates and flagged the distich as promising.
Vals AI adds its own caveats: the success came from systematic trial and error and persistence rather than stronger cryptanalysis, and in hindsight the solution is simple enough that humans could have found it. Only one English outlet has covered this, and the evidence chain rests on Vals AI's account.
Source: The Decoder
Gemini voice controls arrive in Gmail, Docs and Keep
Google turns Gemini into an action layer for mail, documents and notes, reserving the strongest workflow for pricier plans.
Google began rolling out Gemini-powered voice interfaces for Gmail, Docs and Keep on September 3. Gmail Live lets users ask questions about information buried in their inboxes and keep the conversation going, Keep Live organizes spoken notes, and Docs Live turns a spoken stream of ideas into a structured document. Sundar Pichai announced the rollout on X and singled out Docs Live; the features were first previewed at I/O earlier this year.
The tiering is clear: Gmail Live and Keep Live reach Plus, Pro and Ultra subscribers, while Docs Live requires Pro or Ultra. Business access is only a promise for later, with no date attached.
Source: RuntimeWire · Google Workspace blog
Four major AI model services go down at almost the same time
ChatGPT, Claude, Grok and Gemini broke practically simultaneously.
Ars Technica reports that service interruptions hit ChatGPT, Claude, Grok and Gemini practically simultaneously — a rare case of four major model services failing in the same window.
What is verifiable stops there: only the headline and standfirst of that single English report are available here, so the start and end times, the scope of the impact and the cause have no citable account for now.
Source: Ars Technica
Global AI news
Nvidia agrees to buy Hugging Face for $12.93 billion
The main distribution hub for open models changes hands, but the deal cannot close before the first half of 2027.
Jensen Huang announced on September 3, in a post on Nvidia's blog, that the company has agreed to acquire Hugging Face for $12,930,300,000. The Decoder breaks that down as roughly $11.9 billion in purchase price plus a stock program of up to $1 billion to retain staff, and notes that per the SEC filing the deal won't close until the first half of 2027 and still needs regulatory approval. More than 18 million developers and 200,000 companies use the hub.
Nvidia promises Hugging Face stays open: developers pick their own models, frameworks, clouds, inference providers and compute platforms, and Nvidia compute will not be required to build or deploy through the hub. Those are company commitments for now, on a deal that has not been approved yet; The Decoder's reading is that Nvidia has bought the front door to open AI along with a new sales channel for compute.
Source: NVIDIA · The Decoder
Nvidia's PAIR pools idle home PCs for local inference
A free tool that syncs compatible home computers to share local inference work, alongside Ollama and LM Studio.
Nvidia announced the Personal AI Router (PAIR), a free tool that syncs up home computers for local AI inference tasks with tools like Ollama and LM Studio; the open-source software is published in the NVIDIA/Personal-AI-Router repository on GitHub. The Verge gets the obvious confusion out of the way first: despite what the name implies, PAIR is not a hardware router but software. Compatible devices are mostly Nvidia GeForce GPUs from the RTX 20-series onward.
That is also where the constraints sit: the pooling only helps if several compatible machines are already sitting idle on the same home network, and the RTX 20-series floor leaves older hardware out.
Source: The Verge · PAIR page · GitHub
NeoMME releases 260M and 800M multimodal multilingual encoders
One bidirectional Transformer takes text tokens and raw image patches directly, cutting index storage to 6 kB per page.
H company published NeoMME on Hugging Face, a family of 260M and 800M multilingual multimodal encoders. It uses no separate pretrained vision tower and no causal language model; a single bidirectional Transformer processes text tokens and raw image patches, and the whole model is trained from scratch with a masked discrete-diffusion objective. The team fine-tuned NeoMME-Retriever for visual document retrieval using ColPali's page-image approach, returning dense and late-interaction embeddings in one forward pass.
The published numbers: both sizes sit on the ViDoRe v3 Pareto frontier for nDCG@10 versus model size; at a matched 2048×2048 image input on an NVIDIA L40S, the 260M model encodes about 51 pages per second, roughly twice ColModernVBERT's throughput; and hierarchical token pooling plus asymmetric quantization cut late-interaction index storage from about 1.5 MB to 6 kB per page, a 255× reduction.
All of these are author-reported results. The Pareto claim is scoped to two axes on ViDoRe v3, the throughput comparison is tied to one GPU and one input size, and there is no third-party reproduction yet.
Source: Hugging Face · arXiv
Regional and early signals
GLM-5.3-Flash overnight quota discount runs through September 20
Zhipu opens a nightly window for paying GLM Coding Plan users, 23:00 to 09:00, with terms that depend on the client. (Chinese-language source)
Per IT Home, Zhipu AI announced a "Flash × ZCode" overnight campaign for the GLM Coding Plan: from now through September 20, every night between 23:00 and 09:00, all paying subscribers get it automatically with no toggle to flip. The terms for GLM-5.3-Flash in that window depend on the client: used through Zhipu's official coding tool ZCode, quota consumption is zero; used through the other agents the plan supports, available quota is doubled (×2).
Boundaries: this is a time-limited promotion that ends on September 20 and applies only to GLM-5.3-Flash. Only Chinese-language media coverage is available for this item; no English-language primary announcement was found.
Source: IT Home (Chinese-language source)
Tencent's WorkBuddy platform opens with 30-plus hardware brands
An agent platform reaching outward to hardware and vertical software; the launch numbers are partner counts, not usage. (Chinese-language source)
InfoQ China reports that Tencent held its WorkBuddy ecosystem launch in Shenzhen on September 2, opening with nine co-branded hardware devices spanning smart glasses, recording cards, lavalier microphones, an AI desk companion and a keyboard microphone. The partner numbers announced were 30-plus hardware brands, 30-plus Buddy applications and 100-plus open-platform co-creation partners, including Anker, Rokid, GF Securities, Beisen and FanRuan. WorkBuddy launched in March this year and has shipped more than 50 versions in six months.
Liu Yi, Tencent Cloud vice president and head of CodeBuddy and WorkBuddy, said AI productivity gains depend on how much data and capability a tool can connect to and how many real scenarios it can enter. Open-platform lead Yu Hongxiao used financial reconciliation to illustrate the bottleneck: bank statements sit in online banking, invoices in the tax system and vouchers in the ERP, so the model cannot reach the data and does not hold the domain know-how.
All of these figures come from the launch event. The report gives no user numbers, no ship dates for the co-branded hardware and no terms for joining the open platform, and there is no independent verification; Chinese-language sourcing only.
Source: InfoQ China (Chinese-language source)
Get the latest AI model insights and tutorials from Jiufeng.
Explore more


