AI Highlights

GPT-6 Astra beats Pokemon in 18 hours, undone by one Creeper

Key Takeaways

GPT-6 Astra finishes Pokemon FireRed in 18 hours and completes Factorio and Fallout 3, then spends hours farming potatoes after a single Creeper wipes out its chest.

jiufeng
September 18, 2026
44 min read
In this article

Overview

11 stories in this issue. The first 3 are today's priorities.

Model watch

  1. Top · GPT-6 Astra clears four games, then farms potatoes in Minecraft
  2. Top · Anthropic opens a life sciences lane with looser biology guardrails
  3. Top · Claude Code relaunches Projects to run several agents at once
  4. Cooley puts ChatGPT Work into its IPO process

Global AI news

  1. Google opens its experimental CC agent to six people per household
  2. UN statistics move into Data Commons as one searchable platform
  3. vLLM adds GPU hardware video decoding to unblock captioning
  4. Crusoe raises $3.9B and adds small modular "AI factories"

Regional and early signals

  1. Tesla ships a Doubao-model voice assistant to more cars
  2. Jiushi discloses an L4 cluster of over 10,000 accelerators
  3. A pharma company signs a 1.72 billion yuan compute services contract
AI signal map for 2026-09-18

Jiufeng graphic based on the sources cited in this issue.

Model watch

01/11

GPT-6 Astra clears four games, then farms potatoes in Minecraft

One model set records in Pokemon, Factorio and Fallout 3, then lost hours to potato farming after a single Creeper.

According to runs compiled by The Decoder, GPT-6 Astra finished Pokemon FireRed in 18 hours, down from 96 previously, launched a rocket in Factorio, rolled the credits in Fallout 3 and completed Portal — games that had stumped every model before it. The runs come from different projects: the Pokemon figures from GPT Plays Pokemon, the Factorio run from a community setup that hooks the game up through a custom Lua mod and an MCP interface, the Fallout 3 credits from YouTube user imjustnewatai handing a save file to Codex, and the Minecraft run from evaluator Vals AI.

GameAstra's resultRun by
Pokemon FireRedFinished in 18 hours (96 previously)GPT Plays Pokemon
FactorioRocket launch completedCommunity setup
Fallout 3Credits rolledYouTube user imjustnewatai
PortalCompleted—
MinecraftGathered eye-of-ender materials, then lost everythingVals AI

Vals AI says no AI system had ever gotten as far in Minecraft. Astra built a semi-automatic blaze farm in the Nether, killed more than half a dozen Endermen in a warped forest, and gathered the ingredients for eyes of ender — six blaze rods and three ender pearls — leaving only the portal to the final battle to find. Then it put its loot in a chest. A Creeper blew up the chest and the bed, everything was gone, and by the time Astra noticed it had started to rain. It spent hours farming potatoes afterwards, and wrote itself a note partly in caps: "ALWAYS CARRY CRITICAL ITEMS with keepInventory; don't store in unguarded chest."

The report does give same-environment comparisons: in the identical Factorio setup, GPT-5.6 Luna and Fable 5.1 never finished power supply or oil prospecting, and on ARC-AGI-3 Astra scored 62.7% against 7.78% for GPT-5.6 Sol.

Limitations: the results come from separate projects using different harnesses rather than one official benchmark measured on the same terms, and the report gives no repeat counts or failure rates. The potato-farming stretch is itself the report's illustration of a model failing to return to its goal after one unexpected loss.

Source: The Decoder · GPT Plays Pokemon

02/11

Anthropic opens a life sciences lane with looser biology guardrails

Verified life science teams get Mythos, Opus and Sonnet for work the generally available models block.

Anthropic launched the Life Sciences Verification Program (LSVP) on September 17, giving life science professionals access to its Mythos, Opus and Sonnet models under a refined set of safeguards that is more permissive for biology-related work. The company says dozens of organizations were already onboarded through an early-access program, and applications are now open to the broader community, from academic labs to startups and pharma companies. The target is work currently blocked in the generally available Fable models: drug discovery, research biology, clinical development and manufacturing.

Access comes in two tiers, Standard Use and High-risk Use. The latter removes all safeguards that block life science requests, is granted per individual research project and must be renewed every six months; Anthropic states that other protections, such as its cybersecurity classifiers, remain fully in place under LSVP. Applicants go through a verification process that reviews research credentials, security standards and ethics.

Limitations: the program is in beta and initially limited to teams and institutions, with individual Pro and Max plans only promised over time. The announcement gives no acceptance rate and no review turnaround, and the only assessment of the residual risk is Anthropic's own.

Introducing the Life Sciences Verification Program

Image source: anthropic; mirrored on Jiufeng R2.

Source: Anthropic

03/11

Claude Code relaunches Projects to run several agents at once

Each thread is its own cloud session on its own branch; overlaps come back as ordinary merge conflicts.

Claude Code has relaunched Projects, letting several agents run under one roof with shared memory, goals and a shared library of files and artifacts. A project splits into "threads" running different tasks in parallel under a "coordinator"; under the hood each thread is a Claude Code cloud session working on its own branch and copy of the repo. A thread can split its own assigned work further using subagents, loops and workflows. The Verge compares the setup to Grok Bot and other tools that manage groups of AI agents.

The rollout starts today in beta for some Claude Pro and Max subscribers, then extends to all Pro, Max, Team and Enterprise users as well as Cowork and regular Claude chat. Threads currently run in the cloud, with local support said to be coming soon.

Limitations: the coordinator only keeps work organized — when two threads touch the same code, the overlap surfaces as a merge conflict to be resolved like any other PR. The report gives no pricing and no cap on how many threads can run in parallel.

Source: The Verge

04/11

Cooley puts ChatGPT Work into its IPO process

An OpenAI customer story: GO Public, built on ChatGPT Work.

OpenAI published a customer story describing how Cooley built GO Public with ChatGPT Work to bring intelligence to the IPO process. The stated effect is that lawyers surface issues earlier and concentrate their judgment where it matters most.

Limitations: this is OpenAI's own customer story. It gives no usage volume, document counts, time-saved comparison or accuracy figures, and no independent assessment from the client side.

Source: OpenAI

Global AI news

05/11

Google opens its experimental CC agent to six people per household

Up to six family members share one agent, and each decides what it is allowed to see.

Google opened its experimental AI agent CC to families on September 17, with as many as six people in a household sharing one. Each member picks what the agent may see, and CC sorts that material into a shared daily brief, calendar entries and a running task list; the brief goes out each morning as "Your Day Ahead." Google's example: school notices, practice schedules and vet reminders normally sit in one parent's inbox, and pooling them lets the household see who has to be where and what CC already dealt with the day before.

CC also has a verified Google account of its own. Google Labs says that gives it a distinct identity when it appears in a member's inbox or a chat thread, and a clearer place to hold what each person agreed to share. Senders can be added to an auto-share list so that everything from a school or an airline routes to CC, and each member gets a private weekly rundown of new senders to approve or ignore. The family version is live on web and mobile in the US for users 18 and over with a personal Google account; existing users get an upgrade email and everyone else joins a waitlist.

Limitations: CC is still experimental, and the report gives no waitlist size or pace of admission. Sharing means household members' mail flows into one agent; Google's answer is per-member visibility control, but no data on how well that holds up is provided.

Source: SiliconANGLE

06/11

UN statistics move into Data Commons as one searchable platform

Global statistics scattered across UN agencies in conflicting formats now sit in a single open resource.

Google and the UN system launched UN System Data Commons on September 17, an open, AI-ready platform that pulls the statistics compiled across UN entities into one searchable resource. Prem Ramaswami, Head of Data Commons, writes that these agencies hold some of the highest-integrity data in the world, but that the statistics needed to address big global challenges have lived in separate silos in conflicting formats — across, and even within, individual UN organizations — so that connecting the dots often meant months of manual work for analysts before any real analysis could start.

The platform is built on open standards including the Model Context Protocol (MCP), so AI agents can pull data autonomously and turn it into charts and draft reports. Google says every dataset is validated by statisticians and technical experts within the UN system, and the stated goal is to cover 80% of UN statistical datasets by 2027.

Limitations: the announcement does not say how many agencies, datasets or indicators are covered at launch. What the platform can answer depends on what each agency submits, and the 80% figure is a target with no staged timeline attached.

Source: Google

07/11

vLLM adds GPU hardware video decoding to unblock captioning

Moving decode onto the GPU's built-in decoder more than doubles throughput on an 8×H100 node.

vLLM announced support for the hardware video decoders built into NVIDIA GPUs; until now such workloads had to go through the CPU-based OpenCV+FFMPEG backend. The bottleneck is described precisely: video captioning outputs are relatively short, 100-200 tokens, so decoding takes up a much larger share of total time, and on a multi-GPU node running one vLLM server per GPU the CPU cores max out on decode with as few as 2 or 4 GPUs. The fix is integrating PyNvVideoCodec, a Python interface to NVIDIA's hardware decoder. vLLM's own measurement: on 8×H100, GPU decoding more than doubles throughput versus CPU decoding and scales well up to 8 GPUs. The post also notes that CUDA MPS is essential for good performance on this kind of multi-process, high-concurrency bulk VLM inference.

Limitations: the post gives no memory footprint and no list of supported codecs. The gain concentrates in video-in, short-output captioning and labeling; for longer-output tasks decoding was never as large a share to begin with.

Source: vLLM · PyNvVideoCodec · CUDA MPS

08/11

Crusoe raises $3.9B and adds small modular "AI factories"

The data center developer closed a Series F at a $30.9 billion valuation, with Nvidia among the participants.

Crusoe said on September 17 that it raised $3.9 billion in a Series F that pushes its valuation to $30.9 billion, co-led by Atreides Management, Mubadala Capital and Valor Equity Partners, with Founders Fund, GIC, Nvidia, the Qatar Investment Authority, Radical Ventures and TPG also participating. The company named three new board members: Cloudflare CFO Thomas Seifert; Bill Stein, partner and CIO at Primary Digital Infrastructure; and Redwood Materials founder and CEO JB Straubel, who also sits on Tesla's board. Straubel already had ties to Crusoe — he invested personally in 2021, and Crusoe later became the first customer of Redwood's energy storage business.

Limitations: the report covers the financing and board changes only. There is no disclosure of capacity under construction, delivery schedule or customer mix, and no specifications or timeline for the small modular "AI factories" named in the headline.

Source: TechCrunch

Regional and early signals

09/11

Tesla ships a Doubao-model voice assistant to more cars

The Doubao model reaches Tesla's cabin, but only behind the paid premium connectivity tier.

Per IT Home, Tesla released software version 2026.26.200.11 and began a staged rollout that brings a voice assistant built on the Doubao large model to more vehicle models. The assistant requires the premium in-car entertainment subscription, provides real-time information and natural conversation, and offers voice options such as "Weiwei" and "Yunzhou" plus personas including a general know-it-all, a music enthusiast and a storyteller. The same release adds pet mode, blind-spot ambient lighting, destination suggestions in navigation and a "familiar routes" preference.

Limitations: the feature is tied to a paid subscription and pushed in batches; the report lists no specific vehicle models, regions or rollout share, and no model version or response latency. This item rests on a single Chinese-language source, with no official release notes or second outlet confirming it.

Source: IT Home — Chinese-language source

10/11

Jiushi discloses an L4 cluster of over 10,000 accelerators

An urban-logistics driverless vehicle company says it has stacked nearly 15,000 accelerator cards for training.

Jiushi CTO Zhuang Li disclosed on September 17 that the company has built an L4-grade cluster of more than 10,000 accelerators, with total cluster scale approaching 15,000 cards, used to train its APEX multimodal foundation model. The company says that model is moving from tens of billions of parameters toward the hundred-billion range, trained on a mix of real L4 driving data, internet language and video corpora, and human decisions from live operations. A week earlier, on September 10, CEO Kong Qi announced a strategy shift to "city-scale physical AI" in Guangzhou.

Operating figures given by the company:

  • Fleet: more than 30,000 vehicles worldwide, with the first sold in 2023
  • Coverage: over 20 countries and more than 300 cities
  • Mileage: over 270 million kilometers of real operation
  • Compute: total cluster scale near 15,000 cards

Limitations: the card count, parameter range and the "industry-first L4 10,000-card cluster" claim are all the company's own. The report gives no chip models, utilization figures, model benchmark scores or third-party verification, and comes from a single Chinese-language source with no independent English coverage.

Source: Leiphone — Chinese-language source

11/11

A pharma company signs a 1.72 billion yuan compute services contract

A traditional medicine maker moves into compute leasing, borrowing to buy 1.141 billion yuan of servers first.

Kanghui Pharmaceutical disclosed after market close on September 16 that its wholly owned subsidiary Beijing Kanghui Zhichuang signed a five-year compute services contract with a customer identified only as Company A, worth about 1.72 billion yuan including tax. It simultaneously disclosed a server purchase contract with supplier Company G worth about 1.141 billion yuan, with each batch payable in full within 50 days of delivery. Compute is to be delivered in batches from the end of Q3 2026 through the end of Q1 2027. The company warns that the purchase is financed mainly through financial institutions and that its debt-to-asset ratio will rise from 69% at the end of June 2026 to roughly 78%. TMTPost notes that Guangdong Wanniansphere's controlled subsidiary Wanhong Zhisuan signed a similar framework order worth over one billion yuan on August 27.

ItemAmount
Five-year compute services contract (incl. tax)~1.72 billion yuan
Server purchase contract~1.141 billion yuan
Static gap (before costs)under 600 million yuan
Expected added 2026 revenue~30 million yuan

Limitations: TMTPost states explicitly that the static gap excludes financing interest, equipment depreciation, power and operations costs and does not represent achievable profit; the company says the effect on 2026 net profit cannot yet be determined. Neither Company A nor Company G is named, and this rests on a single Chinese-language financial outlet with no third-party verification of delivery or performance.

Source: TMTPost — Chinese-language source

Generate one yourself with JIUFENG

Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.

Generate free