AI Highlights

Qwen opens 2.4T model to self-hosting with 95B active params

Key Takeaways

Qwen opens a 2.4T-parameter model for self-hosting, DeepSeek reroutes V4 Pro traffic to the cheaper V4.1 Flash, and X's revised terms make users liable for its AI agents.

jiufeng
September 10, 2026
33 min read
In this article

Overview

9 stories in this issue. The first 3 are today's priorities.

Hot model updates

  1. Top · Qwen opens a 2.4T model to self-hosting with 95B active params
  2. Top · DeepSeek reroutes V4 Pro API traffic to V4.1 Flash
  3. Top · Mistral agents migrate 40,000 lines of Fortran 77 to C++
  4. Suno rebuilds its training data and ships v6 in three tiers

Global AI news

  1. AWS publishes a HyperPod deployment recipe for the 2.4T Qwen model
  2. X's new terms make users liable for its AI agents
  3. IBM ships a new Granite time-series foundation model

Regional and early signals

  1. Tencent Cloud turns the database schema into an agent app's backend contract
  2. AI coding hackathon sets a Guinness record with 15,577 apps
AI signal map for 2026-09-10

Jiufeng graphic based on the sources cited in this issue.

Hot model updates

01/09

Qwen opens a 2.4T model to self-hosting with 95B active params

Alibaba Cloud's Qwen team put a Max-class mixture-of-experts model into downloadable form, giving teams willing to run it a fully self-hosted option.

Qwen3.8-2.4T-A95B was released on August 12th with downloadable weights and configuration files rather than access through a hosted provider only. The model card lists:

  • Scale: 2.4 trillion total parameters, 95 billion active during inference; a 92-layer architecture with 512 experts per mixture-of-experts block, of which 10 routed experts and one shared expert are active
  • Context: 262,144 tokens natively, extensible to 1,010,000
  • Serving: released artifacts support Hugging Face Transformers, with serving engines including vLLM, SGLang and TokenSpeed
  • Positioning: adjustable reasoning depth, described by RuntimeWire as a self-hosted alternative to OpenAI's GPT-5.5 Pro, a higher-compute reasoning model aimed at demanding coding, research and knowledge work

Limitations: these figures come from the model card and official materials; the report gives no hardware footprint, throughput or cost for running the model yourself. The 1,010,000-token figure is an extended ceiling, not the native length, and Alibaba still offers managed access through Qwen Cloud, so self-hosting is one path rather than the only one.

GitHub - QwenLM/Qwen3.8 at runtimewire

Image source: GitHub; mirrored on Jiufeng R2.

Source: RuntimeWire · Qwen team GitHub · Primary source Gist

02/09

DeepSeek reroutes V4 Pro API traffic to V4.1 Flash

Until V4.1 Pro arrives, every V4 Pro API request goes to the cheaper V4.1 Flash — an early retirement for the current premium endpoint.

DeepSeek plans to release V4.1 Flash around September 10th in Beijing and temporarily route all V4 Pro API requests to it, cutting token prices for some Pro workloads by more than 70%. The company had already moved an updated Flash API into public beta on July 31st and released an experimental model, DeepSeek-V4-Flash-Vision-Exp, whose weights page is on Hugging Face.

An early look from Open Design offers one narrow comparison on design tasks:

ModelOpen Design arena score (out of 100)
GPT-6 Astra82.7
DeepSeek V4.1 Flash81.2
GPT-5.6 Sol77.6

In the same run, V4.1 Flash scored 28.4 out of 30 for requirement fulfillment, 52.8 out of 70 for design quality, a 57.7% delivery rate and a 5.3-minute completion time.

Limitations: DeepSeek says V4.1 Flash beat V4 Pro on performance, inference cost, generation speed and total task completion time, but those remain company claims from internal and external testing until independent evaluations cover coding, reasoning, agent and multimodal workloads. The Open Design arena measures design tasks only. Applications tuned around V4 Pro's output, tool use or reasoning behavior need retesting even if their API configuration is unchanged.

Source: RuntimeWire · Hugging Face model page

03/09

Mistral agents migrate 40,000 lines of Fortran 77 to C++

A European energy operator handed a reservoir simulator with no test suite and no centralized documentation to Mistral's agents for language migration.

The physics-intensive simulator carried 40,000 lines of Fortran 77, and once the original authors left, the knowledge embedded in that code became hard to recover. Mistral's own read of the project: translating syntax between languages is largely a solved task — hand a recent model a snippet in a non-obscure language and it converges in a few iterations — while migrating a full procedural system into object-oriented C++ forces architectural work.

Limitations: this is a vendor-published customer story with a single point of view; the customer is not named and no third party verified the result. The original system had neither a test suite nor centralized documentation, a difficulty Mistral itself flags, which makes post-migration correctness the hardest part to deliver.

Source: Mistral AI

04/09

Suno rebuilds its training data and ships v6 in three tiers

v6 is Suno's first model made with record-industry support, and it arrives as three models — v6, v6-wild and v6-mini — with mini free for everyone.

Suno's Jack Brody told The Verge that v6 was "trained from the ground up, with a new set of data that does not include the same data that our previous models were trained on." That data includes content licensed from Warner Music Group, BMG and Believe, plus "user data." Of the three tiers, v6-mini is the one available free to all users, focused on fast, resource-light creation, with simpler results that more often carry artifacts betraying their origin.

Limitations: The Verge notes it is unclear whether the new training data is completely free of dubiously obtained content, and reports that v6 still cannot capture the "natural imperfections" of real human music.

Source: The Verge

Global AI news

05/09

AWS publishes a HyperPod deployment recipe for the 2.4T Qwen model

The official walkthrough covers running Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM, from cluster provisioning to an OpenAI-compatible endpoint.

The target is the same 2.4-trillion-parameter open-weight model whose weights just went out, and the walkthrough covers cluster provisioning, NVFP4 quantization and exposing an OpenAI-compatible endpoint — removing the last interface difference between self-hosting and a managed API.

Limitations: this is a deployment guide rather than an evaluation, and the abstract carries no throughput, latency or cost numbers. NVFP4 quantization trades precision for speed, so compare outputs before and after on your own benchmark, and bring your own cluster capacity and GPU quota.

Source: AWS Machine Learning Blog

06/09

X's new terms make users liable for its AI agents

Published September 9th and effective October 9th, the revised terms assign users responsibility for prompts, outputs, information and actions produced through autonomous features.

Users get 30 days before the change takes effect; continued use after October 9th counts as acceptance. They become responsible for what autonomous features generate and do, including compliance with laws and X policies, while X retains a broad liability cap. Dispute provisions now extend across a corporate group that includes SpaceXAI, Cursor and SpaceX: the class-action and jury waivers apply to all users to the extent permitted by law and cover those affiliates, and the terms also shorten filing periods and limit available remedies. What splits by region is the forum and governing law — users outside the EU, EFTA states and the UK agree to Texas law and a Texas forum, with individual arbitration when that forum is unavailable, while users in the EU, EFTA states and the UK fall under Irish law and keep local protections that cannot be waived.

Limitations: as RuntimeWire puts it, the agent exists in the contract before it exists in the product — X has mapped the liability without explaining what these autonomous features will actually do or when they ship.

Source: X Privacy Center · RuntimeWire · X Terms of Service

07/09

IBM ships a new Granite time-series foundation model

PatchTST-FM-r2 arrives under a commercial-friendly license at roughly 385M parameters, aimed at zero-shot forecasting.

  • Scale: about 385 million parameters, part of the Granite TSFM family and a new version of PatchTST-FM-r1
  • Versus the predecessor: updated architecture, larger pretraining corpus, plus probabilistic forecasting and imputation of missing values
  • Ranking claim: IBM says that as of September 8, 2026 it is the top performing zero-shot model released under a commercial-friendly license
  • Where to get it: model on Hugging Face, code in ibm-granite/granite-tsfm

The pitch behind time-series foundation models is a change in how forecasting systems get built: instead of training and maintaining a separate model per dataset, users run a pretrained model zero-shot.

Limitations: the ranking comes from IBM's own post, so readers should check the evaluation setup themselves. Zero-shot accuracy shifts with data distribution, so validate on your own series before relying on it.

Source: Hugging Face Blog · GitHub

Regional and early signals

08/09

Tencent Cloud turns the database schema into an agent app's backend contract

CloudBase for Supabase maps REST and RPC interfaces straight out of the schema, and pooled resources return a database in under three seconds.

At the DBTalk database AI salon in Beijing on September 5th, Tencent Cloud's PostgreSQL team argued that AI shortens code generation time, not the production engineering checklist — authentication, permission isolation, elastic scaling, cost control and data security all still apply. Their approach treats the schema as the backend contract: the platform maps REST APIs and RPC interfaces from it, then pairs JWT, table-level authorization and row-level security policies to push backend capabilities down into the database layer.

Delivery changes too: provisioning a traditional cloud database still takes minutes, while pooled resources can return a database in under three seconds, billed in CUs where 1 core / 2 GB per hour counts as 1 CU. For the three failure modes named for multi-agent work — unreliable long-term memory, unclear isolation boundaries and hard-to-roll-back intermediate state — the team's answers were factual memory in relational tables, pgvector for semantic retrieval and Apache AGE for relational memory; Cube Sandbox for microVM-level isolation; and PG 18 branching to try a change on a branch before rolling back to a point, with a demo figure of a 3.8 GB database clone dropping from just over 1,200 milliseconds to 56.8 milliseconds. Two outside figures were cited on stage: Neon's estimate that AI-driven database requests could reach 80% of the total, and Gartner's projection of 33% enterprise agent adoption by 2027.

Limitations: Chinese-language source, with no global English reporting; these are vendor statements made at the vendor's own salon, and the three-second, 80%, 33% and clone-latency figures come with no public benchmark or third-party verification. Every remedy offered sits inside Tencent Cloud's own product stack.

Source: InfoQ China (Chinese-language source)

09/09

AI coding hackathon sets a Guinness record with 15,577 apps

The Bund Hackathon AI Coding contest was certified by Guinness World Records, with nearly 70% of its participants under 18.

On September 9th the event took the record for "most applications developed in an online generative AI programming marathon" with 15,577 submissions from more than ten thousand participants. Close to 70% were under 18, the youngest was 6, and a 9-year-old took first place. At the awards ceremony, 11 leading AI coding platforms jointly published a technical specification for open AI coding challenges, described by the organizers as the first such specification written jointly by multiple platforms in China.

Limitations: Chinese-language source. Certification counts original applications that meet an MVP bar and pass official review — a measure of volume, not of application complexity, retention or continued usability — and the specification's actual clauses are not detailed in the report.

Source: Leiphone (Chinese-language source)

Generate one yourself with JIUFENG

Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.

Generate free