AI Highlights

ChatGPT's web traffic share rebounds to 55.5 percent

Key Takeaways
  • •Similarweb puts ChatGPT back at 55.5 percent of chatbot web traffic
  • •vLLM runs GLM 5.3 at 1M context on one 8-GPU node
  • •Anthropic names three labs over distillation.
jiufeng
September 8, 2026
33 min read
In this article

Overview

9 stories in this issue. The first 3 are today's priorities.

Hot model news

  1. Top · ChatGPT's web traffic share rebounds to 55.5 percent
  2. Top · GLM 5.3 runs a full 1M context on a single 8-GPU node
  3. Top · Anthropic names three labs behind industrial-scale distillation of Claude

Global AI news

  1. Attackers used an agent framework to steal credentials in under six hours
  2. Mistral raises €3B with Samsung leading the round
  3. Motional opens 20,000 driving edge cases as a reasoning benchmark
  4. Meta tests Muse, a personal agent that can spend your money

Regional and early signals

  1. Perplexity's local agent now runs on Qwen3.8-27B
  2. DeepCtrls draws strategic investment from CATL and Aramco Ventures
AI signal map for 2026-09-08

Jiufeng graphic based on the sources cited in this issue.

Hot model news

01/09

ChatGPT's web traffic share rebounds to 55.5 percent

Similarweb has ChatGPT back up from 52.7 to 55.5 percent in three months, while Claude climbed from 1.9 to 9.3 percent year over year.

Similarweb data on AI chatbot website traffic puts ChatGPT at 55.5 percent, up from 52.7 percent three months ago. Google Gemini is slipping again after a strong comeback, from 27.8 to 25.6 percent. The year-over-year picture looks different: ChatGPT is down from 73.3 percent, Gemini doubled its share, and Anthropic's Claude grew from 1.9 to 9.3 percent. DeepSeek (3.4 percent), Grok (2.4 percent), Copilot (1.6 percent) and Perplexity (0.9 percent) remain bit players.

These figures cover website traffic only. The Decoder notes that Google likely drives many mobile Gemini interactions through its Android ecosystem that never appear here, sending push notifications that open the Gemini app with a summary instead of a web link, and that OpenAI is pushing its own mobile app and new desktop client. The ranking describes web entry points, not total usage.

Source: The Decoder · Similarweb via X

02/09

GLM 5.3 runs a full 1M context on a single 8-GPU node

vLLM wires HiSparse in as a pressure-driven memory tier, so requests keep decoding when their KV cache no longer fits in GPU memory.

vLLM published the first batch of optimizations it built for GLM 5.3. HiSparse is integrated as a pressure-driven memory tier that composes with the Hybrid Memory Allocator and KV offloading: when a request's KV cache runs out of room on the GPU, blocks move to host memory and the request keeps decoding. Until now there were two options. Preemption drops a request's KV cache and re-prefills it later, so the request pays its full time-to-first-token again on every eviction. Offloading moves blocks to host memory, but dense attention requires every token to be resident on the GPU, so concurrency stays bounded by GPU memory. For a sparse-MLA KV cache the indexer selects top-K tokens and attends only to those, and HiSparse exploits that to offload the rest. The result is an aggregated deployment on a single 8x H200 node running GLM 5.3 at the full 1 million context length, which was previously impossible on this hardware.

This is part one of a two-part series; part two is not out yet. HiSparse itself is described in a separate arXiv paper, and the 1 million context figure applies to that specific aggregated 8x H200 single-node setup.

Source: vLLM Blog · HiSparse paper

03/09

Anthropic names three labs behind industrial-scale distillation of Claude

The company says DeepSeek, Moonshot and MiniMax pulled more than 16 million exchanges out of Claude through roughly 24,000 fraudulent accounts.

In a policy announcement, Anthropic says it identified industrial-scale campaigns by three AI laboratories — DeepSeek, Moonshot and MiniMax — to illicitly extract Claude's capabilities in order to improve their own models. The labs generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts, in violation of Anthropic's terms of service and regional access restrictions. The technique is distillation: training a less capable model on the outputs of a stronger one.

Anthropic itself stresses that distillation is a widely used and legitimate training method, and that frontier labs routinely distill their own models into smaller, cheaper versions. Its argument is that competitors can use it to acquire powerful capabilities at a fraction of the time and cost. The post says these campaigns are growing in intensity and sophistication and that the window to act is narrow, requiring fast coordination across institutions.

Detecting and preventing distillation attacks

Image source: anthropic; mirrored on Jiufeng R2.

Source: Anthropic

Global AI news

04/09

Attackers used an agent framework to steal credentials in under six hours

Google's threat intelligence group says troubleshooting and IP rotation ran with no operator in the loop.

Google Threat Intelligence Group released "From Prompting to Autonomy: The Evolution of Adversarial AI," covering activity it tracked over the second quarter. Mandiant investigators traced one campaign to a suspected financially motivated actor that first broke into an organization's cloud infrastructure, then assembled an autonomous framework out of an AI coding chatbot, a prompt and a set of agent instructions, with preconfigured markdown playbooks driving the scanning and credential harvesting that followed. Thousands of credentials were compromised in under six hours. Troubleshooting and IP rotation ran without an operator, and traffic left the victim's own addresses, so it looked legitimate on the way out.

GTIG says that what has changed since its May edition — which documented the first confirmed case of criminals using AI to build a working zero-day exploit — is how little human involvement is left. The report as covered does not name which vendor's model or product the attacker used.

Source: SiliconANGLE

05/09

Mistral raises €3B with Samsung leading the round

Post-money valuation tops €21 billion, and the money is aimed at owning more compute.

Mistral announced a €3 billion Series D at a post-money valuation of more than €21 billion, which the company calls the largest equity fundraising round ever completed by a European technology company, three years after launch. Samsung Electronics led the round, joined by co-leads Scaleup Europe Fund, managed by EQT, and existing investor PSG Equity. Mistral says the money will expand its frontier research, scale compute capacity for training, and accelerate commercial growth and international footprint. The company operates across 20 countries and supports more than 125 global enterprises including Airbus, ASML and HSBC. CNBC reported the financing as roughly $3.5 billion at a valuation equivalent to about $24 billion.

RuntimeWire notes the sovereignty pitch only works commercially if enterprise revenue catches up with the infrastructure bill, and that the "open-weight" label does not describe every Mistral model under identical terms — licensing varies model by model. CEO Arthur Mensch told CNBC that Mistral plans to focus on manufacturing work with Samsung, similar to ASML, which led the €1.7 billion Series C in September 2025 before integrating Mistral's AI into manufacturing processes.

Source: Mistral AI · RuntimeWire · Mistral licensing guidance

06/09

Motional opens 20,000 driving edge cases as a reasoning benchmark

nuReasoning asks models to explain why one maneuver is safer, not to imitate the recorded trajectory.

Motional released the nuReasoning dataset, packaging 20,000 scenarios and more than 105 hours of difficult driving events with 247,000 human-verified annotations covering spatial relationships, driving decisions and counterfactual reasoning. It targets the unusual road situations where perception alone is not enough: models are asked to explain their decisions and justify why one maneuver is safer than another, rather than simply imitate the recorded vehicle trajectory. The material comes from Motional's own fleet archive; autonomous-driving programs usually keep such edge cases private.

The coverage reports no baseline scores for any model on the dataset. The September 8 announcement follows several months of public research, with the accompanying paper on arXiv.

Source: RuntimeWire · nuReasoning paper

07/09

Meta tests Muse, a personal agent that can spend your money

An invite-only iPhone app wired to email, calendars, health and financial accounts, asking for approval before it acts.

Meta is testing Muse, a standalone personal AI agent designed to manage schedules, make purchases, track spending and keep working after users close the app. A Meta-authored App Store listing gives the clearest account: an iPhone agent with access to email, calendars, health information, financial accounts and Meta's social apps. Examples in the listing include negotiating a bill, selling a car, planning and booking a trip, auditing subscriptions and buying products after receiving approval. Muse can also run in the background, monitor unfinished work and notify the user when a decision is required.

The primary evidence is a September 7 TestingCatalog post plus that public App Store listing; Meta has published no formal announcement. Access is restricted through a waitlist and invite codes. RuntimeWire argues the approval system is what will determine whether users trust Muse to act on their behalf.

Source: RuntimeWire · TestingCatalog

Regional and early signals

08/09

Perplexity's local agent now runs on Qwen3.8-27B

Chinese-language source: the PPLX 27B series is built on Qwen3.8-27B and tuned for NVIDIA DGX Spark.

Leiphone reported on September 8, citing US outlet SiliconANGLE, that Perplexity's new local agent product Portable Computer runs on Alibaba's open-source Qwen3.8-27B, with the PPLX 27B series optimized for NVIDIA DGX Spark hardware to handle file processing, data analysis and coding. The report puts Perplexity's latest valuation above $30 billion and says its model choices previously centered on Claude and GPT. The same piece lists other Qwen adopters: Hugging Face's summer report calls Alibaba's Qwen a foundation base for global open-source AI, Airbnb's business is said to rely heavily on Qwen, Pinterest uses it for a shopping assistant and chatbot, and Reuters has built its own model on Qwen.

This is a Chinese-language retelling of English reporting, with no Perplexity announcement attached. The "local Opus4.6" nickname is a community label, not a benchmark result, and none of the adopters' usage volumes or deployment scopes are quantified.

Source: Leiphone (Chinese-language source)

09/09

DeepCtrls draws strategic investment from CATL and Aramco Ventures

Chinese-language source: a third round in two months, aimed at physical AI control for compute and energy infrastructure.

Physical AI company DeepCtrls closed a B+ round of several hundred million yuan, led by CATL, with Aramco Ventures, Taiping Innovation Investment, GF Xinde and Fosun Capital participating and existing backers Source Code Capital and Guangyuan Capital adding on. The company says it has completed three rounds in the past two months. Technically, DeepCtrls has worked since 2018 on combining AI with physical mechanisms, centered on its PhyAI engine for real-time optimization and control; its products are moving from industrial settings into compute infrastructure, including liquid-cooling control and compute-power coordination. CATL is also described as a long-standing enterprise customer.

The report follows company-supplied framing and discloses no exact amount, valuation or revenue. Claims such as being "among the first in the world to deploy closed-loop physical AI control" and serving "several hundred leading enterprises globally" come from the company itself, without third-party verification.

Source: QbitAI (Chinese-language source)

Generate one yourself with JIUFENG

Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.

Generate free