AI Highlights

Google ships Gemini 3.8 Live with background tool calls

Key Takeaways

Google DeepMind ships Gemini 3.8 Live for voice agents that reason and call tools mid-conversation, plus a shift in OpenRouter spend, enterprise pushback on Claude log retention, and Salesforce inside Claude.

jiufeng
September 16, 2026
35 min read
In this article

Overview

9 stories in this issue. The first 3 are today's priorities.

Hot model updates

  1. Top · Gemini 3.8 Live ships with background tool calls
  2. Top · OpenAI overtakes Anthropic in weekly OpenRouter spend
  3. Top · Claude's enterprise week: Salesforce skills in, Nvidia use narrowed
  4. Paying for frontier models buys about four months

Global AI news

  1. WhatsApp Business setup handed to coding agents
  2. Qwen3-8B tuned into a dedicated product tagger
  3. Acing a task once is not enough for agents

Regional and early signals

  1. DeepSeek hires a CFO from Hillhouse's venture arm
  2. A car-parts maker pays 800M yuan for a compute integrator
AI signal map for 2026-09-16

Jiufeng graphic based on the sources cited in this issue.

Hot model updates

01/09

Gemini 3.8 Live ships with background tool calls

Google DeepMind released two audio-to-audio models, and the Extended Thinking variant keeps talking while it reasons and runs tools in the background.

On September 15th, Google DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two models for developers building voice agents.

  • Gemini 3.8 Live: tuned for low-latency conversation and direct tasks, placed second on Speech Agent Arena according to Google, and available through the Gemini API, the Gemini app, AI Studio, Google Cloud and Search Live;
  • Gemini 3.8 Live Extended Thinking: the more deliberate option, aimed at multi-step work such as diagnosing a technical problem, with reasoning and tool calls running in the background, spoken updates covering the wait, reasoning configurable at low/medium/high, and availability in the Gemini app, the Gemini API, AI Studio, Google Cloud and parts of Workspace (Gmail, Docs, Keep);
  • Shared capabilities: audio-to-audio, real-time visual input, and external tool calls during a conversation.

Benchmark results the official post attributes to Extended Thinking:

BenchmarkExtended Thinking score
Speech-to-speech quality index82.6 (Artificial Analysis, 1st)
τ-Voice68.6%
τ-Voice-banking35.1% (Sierra)
Big Bench Audio97.7%

Limitations: Google publishes no latency figures and describes pricing only as "highly competitive," without listing rates. On integration, Extended Thinking requires non-blocking tool declarations — synchronous function calls raise an error — and apps must poll interaction_status themselves; the model card also notes it can be slower or time out.

Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking

Image source: Google; mirrored on Jiufeng R2.

Source: Google DeepMind · RuntimeWire

02/09

OpenAI overtakes Anthropic in weekly OpenRouter spend

GPT-6 Astra took the largest share of dollars while the cut-price GPT-5.6 Luna took the largest share of tokens.

Data posted on X by Peter Walker, OpenRouter's head of insights, shows OpenAI models drew more spending than Anthropic models on OpenRouter during the week of September 7th to 13th. Walker said the last time that happened was February 26th, 2024.

ModelShare of dollars spent, both families (%)
GPT-6 Astra19
Claude Opus 516

On the token side, Luna led by a wide margin. Released on July 9th, Luna had its API rates cut by 80% on July 30th, and current model documentation lists $1.20 per million output tokens. OpenAI positions it as the GPT-5.6 option for cost-sensitive, high-volume jobs.

Limitations: This is one platform and one week, counting only these two companies' models. The report itself notes that more time is needed to separate durable production traffic from launch demand.

Source: RuntimeWire · Peter Walker on X · OpenAI pricing note · Model documentation

03/09

Claude's enterprise week: Salesforce skills in, Nvidia use narrowed

Anthropic put 37 Salesforce skills into a Claude beta, while its 30-day log retention has Nvidia keeping Fable to less sensitive work.

Salesforce in Claude: Anthropic launched Salesforce in Claude on September 15th as a beta, packaging 37 skills for account research, call preparation, deal reviews, forecasting and CRM updates. Claude can collect context from Salesforce, Slack, email and call transcripts, then propose changes such as moving a close date, adding a contact, logging a call or creating a follow-up task. It works under existing Salesforce permissions, Salesforce stays the system of record, and writes go back only after approval by default.

Log retention narrows large accounts: According to The Information, as relayed by The Decoder, Anthropic said in June that it would retain Fable usage logs for 30 days to defend against "complex and novel attacks." Nvidia now uses Fable only for less sensitive tasks such as open-source projects, running its own Nemotron models for internal work like AI-powered supply chain monitoring. Justin Boitano, Nvidia's VP of Enterprise AI, said the company believes zero data retention should be on by default. Palantir and defense contractor Booz Allen Hamilton are named as also limiting their use: Booz Allen CTO Bill Vass said he worries Fable "may be learning our code," while Palantir CEO Alex Karp told a customer event that enterprises are tired of being "exploited" by AI labs. The report also notes that both labs collect metadata and technical usage data from corporate customers; OpenAI calls such data "de-identified," meaning stripped of information traceable to individual customers, and says classification results are metadata about business data and do not contain any business data itself, as stated on its enterprise privacy page.

Limitations: The Salesforce release is a beta with no announced regions, pricing or supported Salesforce editions, and nothing is said about capabilities beyond the 37 skills or whether the approval step can be turned off. The retention details come from The Information's report and The Decoder's summary, and no contract sizes or usage volumes are given.

Source: RuntimeWire · Claude on X · The Decoder · OpenAI enterprise privacy

04/09

Paying for frontier models buys about four months

Mozilla's report puts the premium tier's advantage at roughly a four-month head start, at five times the cost.

Ars Technica previewed an as-yet unpublished Mozilla report whose headline finding is that paying for frontier AI models buys about a four-month capability head start at five times the cost. The report itself is about how cheap open-weight models caught up on capability with Silicon Valley's frontier systems.

Limitations: Ars saw only a preview and Mozilla's report is not public yet, so for now the four-month and five-times figures rest on Ars's account.

Source: Ars Technica

Global AI news

05/09

WhatsApp Business setup handed to coding agents

Meta shipped a WhatsApp Business MCP server so developers can do the onboarding by chatting with a coding agent.

On September 15th, Meta said developers can now let an AI agent of their choosing set up and manage WhatsApp Business messaging, via a new WhatsApp Business MCP server that covers setup, messaging templates, testing and troubleshooting; the reporting names Claude, Cursor, Codex and ChatGPT among the agents that can drive it. Meta says the process previously required moving between the Developer Console, Business Manager, the API reference and an editor; now developers can describe what needs to be done in conversation. The news came alongside Meta's new AI-focused subscription plans.

Limitations: The report gives no regions, pricing or supported API versions, and does not say whether agent-made changes require human confirmation.

Source: TechCrunch

06/09

Qwen3-8B tuned into a dedicated product tagger

AWS published a serverless recipe: supervised fine-tuning first, then RL with verifiable rewards, so an 8B model returns attributes in a fixed schema.

The AWS Machine Learning Blog published a walkthrough that customizes Qwen3-8B with supervised fine-tuning (SFT) and then reinforcement learning with verifiable rewards (RLVR), using Amazon SageMaker serverless model customization. The premise is that manually tagging thousands of catalog products is slow and inconsistent; the goal is a small model that returns the right attributes in the right schema consistently, rather than prompting a general frontier model each time. The RLVR stage optimizes against deterministic, programmatically scored rewards. On the operational side, when no compute configuration is supplied, SageMaker selects and releases the training capacity for the customization job itself; official notebooks live in the SageMaker Python SDK v3 serverless model customization examples.

Limitations: This is AWS's own how-to post, tied to SageMaker serverless model customization and its official example repository; regional availability and billing need the documentation.

Source: AWS ML Blog · SageMaker Python SDK examples

07/09

Acing a task once is not enough for agents

IBM Research added consistency guidelines to altk-evolve, raising repeat success on the same task by 16 percentage points.

IBM Research introduced consistency guidelines, a new guideline type in the open-source altk-evolve toolkit, in a post on Hugging Face. The premise is that most benchmarks hide variability behind an average: the same request can succeed once and then take a different path and fail, which is a showstopper for work like reconciling a financial transaction or checking a contract for an obligation. The team reports same-task Pass⁵ up 16.0 percentage points and similar-task Pass⁵ up 13.0 percentage points, without costing anything in average accuracy. Full methodology and evaluations are in a technical report on arXiv, with code on GitHub.

Limitations: The write-up is IBM Research's own and the numbers are self-reported with no third-party replication; the post does not cover how results vary across base models or longer task chains.

Source: Hugging Face Blog · altk-evolve · Technical report

Regional and early signals

08/09

DeepSeek hires a CFO from Hillhouse's venture arm

Chinese reporting says DeepSeek's new CFO, Yan Wentao, was a partner at Hillhouse's venture arm who backed Zhipu and MiniMax.

Chinese outlet QbitAI reports that DeepSeek's new CFO is Yan Wentao, born in 1991 and previously a partner at Hillhouse's venture arm, where he invested in Zhipu and MiniMax. DeepSeek had already posted a CFO opening back in February 2025. QbitAI also relays a Reuters report saying DeepSeek has engaged CITIC Securities to prepare a STAR Market listing, with filing as early as the end of this year and a target listing in 2027.

Limitations: Chinese-language source only, and the listing portion is secondhand relay; DeepSeek has not publicly confirmed the appointment details or the timeline. Much of the piece is biographical background unrelated to the listing process. This is a capital-markets signal, not a model or product change.

Source: QbitAI (Chinese-language source)

09/09

A car-parts maker pays 800M yuan for a compute integrator

Xiangshan wants an AI compute business, and the target booked just 2.94M yuan of revenue in 2024.

On the evening of September 14th, Xiangshan Co. disclosed a restructuring plan to buy 100% of Zhejiang Wuluo Smart City Technology for a provisional 800 million yuan, 51% in cash and 49% in shares issued at 35 yuan each. The target only entered the compute business in 2024 and was long a trading distributor: 2.94 million yuan of revenue and a loss that year, 36.66 million yuan of revenue and 19.07 million yuan of net profit in 2025, then 305 million yuan of revenue and 28.1 million yuan of net profit in the first half of 2026. It shifted from trader to hardware integrator after its Ningbo plant came online in December 2025, and had been the fifth-largest new customer of domestic GPU maker MetaX. The performance undertaking skips 2026 entirely, committing only to at least 400 million yuan of cumulative non-GAAP net profit across 2027-2029, about 133 million yuan a year. The stock hit its down limit at 47.70 yuan when trading resumed on September 15th, after rising 35.69% over the 20 sessions before the suspension.

Limitations: Chinese-language source only (TMTPost), and the deal is still at the plan stage without approval. The target's results depend heavily on MetaX supply and chip prices and on downstream AI data center capex, and TMTPost itself flags the difficulty of meeting the three-year 400 million yuan commitment.

Source: TMTPost (Chinese-language source)

Generate one yourself with JIUFENG

Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.

Generate free