AI Highlights

DeepSeek V4.1 Flash hits GA with a 4x smaller KV cache

Key Takeaways

DeepSeek-V4.1-Flash hits GA with a 1M context and an 890-byte/token KV cache under MIT, while vLLM triples MiniMax M3 throughput on AMD MI355X.

jiufeng
September 10, 2026
35 min read
In this article

Overview

9 stories in this issue. The first 3 are today's priorities.

Hot model updates

  1. Top · DeepSeek-V4.1-Flash reaches GA with an 890-byte-per-token KV cache
  2. Top · vLLM triples MiniMax M3 output throughput on AMD MI355X
  3. Top · OpenAI turns a 250-person security sprint into a Defense Factory

Global AI news

  1. Anthropic builds a threat intelligence operation that treats activism as a security risk
  2. Hugging Face rebuilds AUTOMATIC1111 as a Gradio Workflow: 73 nodes, 11 pipelines
  3. TorchServe is no longer maintained, and AWS offers Ray Serve Deep Learning Containers
  4. Six Chinese AI firms accused of aggressively copying US frontier models

Regional and early signals

  1. Zhipu says API revenue reached 86.5% of the total, with token calls up more than 40x since January
  2. Heurist buys premium financial data per query through AgentCore payments
AI signal map for 2026-09-10

Jiufeng graphic based on the sources cited in this issue.

Hot model updates

01/09

DeepSeek-V4.1-Flash reaches GA with an 890-byte-per-token KV cache

DeepSeek shipped the Flash line to general availability with open MIT weights and a global KV cache of 890 bytes per token.

DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with a 552B backbone, 196B additional Engram parameters and a 1M-token context window. It activates 8B parameters per token during prefill and 16B during decode. The number MarkTechPost builds its breakdown around is the global KV cache footprint:

ModelGlobal KV cache (relative to V4.1-Flash)
DeepSeek-V4.1-Flash1x (890 bytes/token)
DeepSeek-V4-Flash~4x
DeepSeek-V1~437x

The 40-layer backbone splits into a 20-layer causal encoder and a 20-layer decoder. Inspired by YOCO, the decoder does not compute its own global KV, which halves prefill. The main KV cache is quantized to E2M1 with one E4M3 scale per 16 channels, following NVFP4 without its global scale.

  • Parameters: 552B backbone plus 196B Engram; 8B active at prefill, 16B at decode
  • Context: 1M tokens
  • Licence and deployment: open weights under MIT, with vLLM, SGLang and Transformers paths on Hugging Face
  • API: low, high and max reasoning tiers
  • Benchmarks: MarkTechPost's Training and Results section compares V4.1-Flash with V4-Flash, Opus 5 and GPT-5.6 Sol on Terminal-Bench 2.1, DeepSWE v1.1, Terminal-Bench 4.0, Automation-Bench and GPQA Diamond

Limitations: the parameter, cache and licence details here all come from these two reports, with no third-party reproduction, and the 890-byte figure is the number MarkTechPost's breakdown is built around. Pandaily separately reports that the new Flash pricing took effect on September 10th.

deepseek-ai (DeepSeek)

Image source: huggingface; mirrored on Jiufeng R2.

Source: MarkTechPost · Pandaily · DeepSeek on Hugging Face · YOCO

02/09

vLLM triples MiniMax M3 output throughput on AMD MI355X

Public InferenceX numbers put MXFP8 serving at 3.14x the day-0 result at concurrency 32.

The baseline is the day-0 MiniMax M3 implementation, which already covered MiniMax Sparse Attention, multimodal inputs, reasoning and tool outputs, MXFP8 weights and EAGLE3. The figures come from the public SemiAnalysis InferenceX benchmark:

Configuration (MI355X)Day-0 (output tok/s/GPU)Optimized (output tok/s/GPU)Reported gain
MXFP8 / concurrency 32109.1342.43.14x
MXFP8 / concurrency 128297.8623.72.09x
MXFP4 / concurrency 128212.1716.8—

Latency moved with it: at concurrency 32, median TTFT fell from 1.46 to 0.67 seconds and mean TPOT from 69.1 to 22.1 milliseconds; at concurrency 128, median TTFT fell from 3.53 to 1.54 seconds and mean TPOT from 100.7 to 48.8 milliseconds. All three rows run on the TP4/EP1 four-GPU path; a later TP2/EP1 run reached 943.5 output tok/s/GPU, 31.6% above that TP4 checkpoint. One of the changes was vLLM #45725, which split the kernel launcher so prefill's many tokens and decode's few rows are handled separately.

Limitations: vLLM stresses that the useful result is not a single higher throughput number but a way to decide what to optimize next as the bottleneck keeps moving; the 943.5 figure uses a different parallel topology and is not directly comparable with the TP4/EP1 rows.

Source: vLLM Blog · vLLM #45725

03/09

OpenAI turns a 250-person security sprint into a Defense Factory

OpenAI published a reference architecture for continuously finding, validating and fixing vulnerabilities, saying agents closed 53 urgent or high-priority issues on day one.

It started as an internal "code red" that brought the Security, Applied and Research groups together to examine hundreds of systems across more than 100 service areas. The resulting Defense Factory connects AI agents to source-control platforms, security scanners, issue trackers and isolated development environments: agents can inventory systems, scan code, test whether suspected vulnerabilities are exploitable, route confirmed findings to an owner, prepare patches and verify deployed fixes. OpenAI published the blueprint in early September and promoted it again on September 9th, saying its latest cyber models uncovered vulnerabilities its security staff might otherwise have missed. In practice the blueprint uses OpenAI's Codex products alongside two controlled cyber access levels, and humans retain review authority over consequential changes.

Limitations: the 53 figure is OpenAI's own, covers day one only, and carries no external verification in the reporting. RuntimeWire notes the bet is that defenders can use frontier models faster than attackers adopt comparable tools — a claim that has not been tested.

Source: RuntimeWire · OpenAI · Daybreak overview

Global AI news

04/09

Anthropic builds a threat intelligence operation that treats activism as a security risk

Reporting says Anthropic monitors protests near executives and facilities, investigates people deemed potential threats, and exchanges information with law enforcement.

A report published Wednesday by The American Prospect ties together Anthropic's job listings, its statements to other publications, and an interview in which security officials described using an AI-assisted risk platform to track a protest. New roles call for OSINT investigations and ties to police.

The contrast is with February 2026, when Anthropic refused to lift restrictions on using Claude for mass domestic surveillance and fully autonomous weapons. Amodei argued at the time that AI could combine location, browsing and association records into comprehensive personal profiles at unprecedented scale.

Limitations: the reporting itself notes the internal program differs from the government surveillance Amodei opposed — it is a corporate security operation. The evidence base is job listings, statements to other outlets and a single interview, not a formal disclosure from Anthropic.

Source: RuntimeWire · Anthropic policy for government data requests

05/09

Hugging Face rebuilds AUTOMATIC1111 as a Gradio Workflow: 73 nodes, 11 pipelines

Workflow1111 recreates most of stable-diffusion-webui's feature set on a single workflow canvas.

Hugging Face published Workflow1111, a gr.Workflow graph of eleven media pipelines built from seventy-three nodes. It covers text-to-image, hi-resolution fix, image-to-image, prompt-matrix grids, VLM interrogate, detection-to-inpaint masks, ControlNet-style annotators, background removal, PNG Info storing and image-to-video. Signing in with a Hugging Face account or supplying an access token runs the pipelines, with model calls billed against your own quota; the Space can also be duplicated and rewired.

Limitations: the post says it rebuilt "most" of AUTOMATIC1111's feature set without listing what is missing, and gives no performance comparison against the original. Running it requires sign-in and consumes the user's own quota.

Source: Hugging Face Blog · GitHub

06/09

TorchServe is no longer maintained, and AWS offers Ray Serve Deep Learning Containers

The official project notice says there are no planned updates, bug fixes or security patches, and AWS is pitching Ray Serve DLCs to take over those inference workloads.

AWS spells out the situation: TorchServe is no longer actively maintained, the official project notice states there are no planned updates, bug fixes, new features or security patches, and vulnerabilities might not be addressed. For teams running inference on TorchServe today, that means security patches stop and compatibility updates with newer PyTorch and CUDA versions stop, leaving engineers owning the whole dependency chain — picking compatible versions across the GPU stack, patching every layer, and debugging subtle failures when a component drifts. The answer extends the Deep Learning Container approach from training to inference: the Ray Serve DLC is a pre-built, performance-optimized Docker image bundling the framework, its dependencies and the GPU stack into a tested, patched combination.

Limitations: this is a migration path rather than a drop-in compatibility layer, and the post gives no migration effort estimate, performance comparison or version timeline.

Source: AWS Machine Learning Blog

07/09

Six Chinese AI firms accused of aggressively copying US frontier models

Ars Technica reports that six Chinese AI companies stand accused of aggressively copying US frontier models, alongside a push for AI firms to identify Chinese users and switch them to weaker models.

Ars Technica's report carries two threads: six Chinese AI firms accused of aggressively copying US frontier models, and a US push for AI companies to identify Chinese users and then secretly switch them to less-capable models.

Limitations: this is single-outlet reporting. The six named companies, the evidence behind the accusation, the body making the recommendation and any enforcement mechanism are not set out in the report summary, so it should not be treated as established fact before the underlying documents are available.

Source: Ars Technica

Regional and early signals

08/09

Zhipu says API revenue reached 86.5% of the total, with token calls up more than 40x since January

Zhipu's first interim report since listing shows revenue shifting from on-premises deployment to its open platform, with ARR at $1.6B at the end of August. (Chinese-language source)

Zhipu reported RMB 954M in revenue for the first half of 2026, up 399.7% year on year and already above its full-year 2025 total. The revenue mix is where the report moved most:

Revenue lineAmount (RMB M)Share of totalYear-ago share
Open platform and API82586.5%15.2%
On-premises deployment12913.5%84.8%
  • ARR: $1.6B at the end of August, up 60% from $1.0B in early July; $133M in August alone
  • Users: over 7.4M registered MaaS accounts, up 144% since the start of the year; paying DAU up 603%
  • Volume: token calls up more than 40x since January, with Coding Plan calls up more than 23x
  • Pricing: average API price up about 101%; API gross margin moved from -0.4% a year earlier to 24.6%

Limitations: these are self-reported interim figures with no independent verification. The same report notes that losses widened alongside the revenue growth and that results still fell short of market expectations. Chinese-language source, single outlet.

Source: TMTPost

09/09

Heurist buys premium financial data per query through AgentCore payments

A conversational investment workbench for retail investors puts pay-per-query data purchases inside the agent's execution chain.

Heurist Finance folds several institutional-style workflows into one chat experience: gathering market data, reading filings and news, running deep research, building and stress-testing portfolios, and monitoring positions, with each answer reflecting the user's own holdings and preferences. It is built on Amazon Bedrock AgentCore, using AgentCore payments to buy premium data per query and combining Identity, Memory, Code Interpreter and observability so that paid data access, sandboxed analysis, identity and memory come together in an auditable response. Heurist's stated motivation: the premium market, macroeconomic, fundamental and alternative data retail investors need sits behind paywalls and bespoke APIs, with no single vendor covering it.

Limitations: this is an AWS-published customer story with no accuracy, cost or user-scale figures and no third-party evaluation; the actual per-query prices and the list of data vendors are not disclosed.

Source: AWS Machine Learning Blog

Generate one yourself with JIUFENG

Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.

Generate free