AI Highlights

MiniMax H3 API renders a 5s 480p clip in about two seconds

Key Takeaways

Pruna AI wraps MiniMax H3 into a paid video endpoint, the WSJ says researchers used Claude to reach OpenAI's private code, and Microsoft patched 950+ flaws.

jiufeng
September 18, 2026
32 min read
In this article

Overview

9 stories in this issue. The first 3 are today's priorities.

Model Watch

  1. Top · MiniMax H3 API renders a 5s 480p clip in about two seconds
  2. Top · Researchers used Claude to reach OpenAI's private code
  3. Top · A 4B embodied model trained on zero real-robot data runs an hour-long livestream

Global AI News

  1. Microsoft patched 950+ flaws in one month, two zero-days exploited
  2. Anthropic and OpenAI look at 20-30 MW data centers
  3. Huawei's OceanStor M900 pools KV cache at petabyte scale
  4. PrismML raises a $22.25M seed round betting on small models

Regional and Early Signals

  1. Zhipu's ZCode admits its indexing feature could upload repository data
  2. Helix 2.5 enters 30 unfamiliar homes zero-shot at a 56% success rate
AI signal map for 2026-09-18

Jiufeng graphic based on the sources cited in this issue.

Model Watch

01/09

MiniMax H3 API renders a 5s 480p clip in about two seconds

Pruna AI packages the open-weight MiniMax H3 into a commercial endpoint billed per output second across four tiers.

Pruna AI launched P-Video-2-Pro in a thread on X on September 17th, giving developers a faster endpoint for generating short videos with audio from text or reference images. The company's stated timing: a five-second clip at 480p in Speed mode takes about two seconds, while the same length at 768p takes roughly 4.3 seconds.

  • Output: 5- to 15-second clips with audio, from text or reference images
  • Rate limit: 250 API requests per minute
  • Base model: MiniMax H3, released in late July and opened on August 3rd
  • Team: Rayan Nait Mazi (CEO), Bertrand Charpentier (president and chief scientist), John Rachwan (CTO) and Stephan Gunnemann (chief strategy officer), working between Munich and Paris on model efficiency rather than training foundation models from scratch

Pricing splits into four tiers by resolution and mode:

TierPer output second (USD)5-second clip (USD)
480p Speed0.020.10
480p Quality0.040.20
768p Speed0.0350.175
768p Quality0.0750.375

Less than seven weeks separate H3's opening from this commercial endpoint, which RuntimeWire reads as the commercial opening that open-weight models create for independent inference teams.

Limitations: the two-second figure applies only to the 480p Speed tier, with the same five-second clip taking about 4.3 seconds at 768p, and the source gives no comparable quality measurement between the Speed and Quality tiers. All the latency and pricing numbers come from Pruna's own launch thread, with no third-party test.

MiniMax H3 API renders a 5s 480p clip in about two seconds: Per output second (USD) and 5-second clip (USD)

Jiufeng graphic based on the sources cited in this issue.

Source: RuntimeWire · Pruna AI on X

02/09

Researchers used Claude to reach OpenAI's private code

RuntimeWire, citing The Wall Street Journal, says researchers got into an OpenAI employee's ChatGPT account and could suggest code changes.

RuntimeWire, citing a Wall Street Journal report, says independent security researchers used Anthropic's Claude to access an OpenAI employee's ChatGPT account, giving them a way to read and suggest changes to OpenAI's private software cache.

The operational question RuntimeWire draws from it: once an agent acquires an employee identity, can account permissions, repository gates and network controls contain it? The piece also carries one piece of background: on September 12th, Dario Amodei announced on X that Anthropic will unilaterally commit to giving external evaluators permanent, employee-level access to its systems, so they can verify safety measures, report incidents and assess model alignment during training.

Limitations: the report does not establish whether the access was authorized or whether any suggested change reached production. It also does not establish whether the researchers accessed customer data or caused operational damage, and it omits the researchers' identities and the incident date. Neither Anthropic nor OpenAI is quoted with a formal statement in this source.

Source: RuntimeWire · Dario Amodei on X

03/09

A 4B embodied model trained on zero real-robot data runs an hour-long livestream

Lexiang Technology's Aether, trained on about 200 hours of human video, drove two humanoids through a public stream in Shanghai.

Pandaily reports that Lexiang Technology's Aether model ran an outdoor barbecue service livestream in Shanghai lasting more than an hour, executed by two humanoid robots.

  • Size: roughly 4B parameters
  • Training data: about 200 hours of human video, zero real-robot data
  • Demo: outdoor barbecue service in Shanghai, dual humanoids, over an hour on stream
  • Framing: a public cross-embodiment demonstration

Limitations: Pandaily is the only report here, and its summary gives model size, training data and format but no success rate, task list or count of human interventions. "Zero real-robot data" is the vendor's own framing, with no independent reproduction, and weights and licensing are not described.

Source: Pandaily

Global AI News

04/09

Microsoft patched 950+ flaws in one month, two zero-days exploited

Roughly 2,750 fixes so far this year, more than double the 2020 record, with the pace attributed to AI-assisted security research.

Microsoft's September 2026 patch release fixed more than 950 vulnerabilities. That brings the year to roughly 2,750 fixes, more than double the annual record of about 1,250 set in 2020.

CategoryFixed this month
All vulnerabilities950+
Rated critical113
Remote code execution258
Elevation of privilege438

Two of them are zero-days under active exploitation: CVE-2026-81963 and CVE-2026-85880, both allowing privilege escalation on Windows. CVE-2026-85880 was found by researchers at Volexity and Proofpoint; CVE-2026-81963 was independently reported by Airbus Helicopters and the Microsoft Threat Intelligence Center. Ars Technica's Dan Goodin notes that the industry is shipping patches at unprecedented speed after OpenAI, Anthropic, AWS, Google and Microsoft published an open letter warning that AI-driven cyberattacks will become more common and more sophisticated.

Limitations: the security community's concern is that enterprises cannot keep up with the volume. Action1's Jack Bicer says the challenge at this scale is knowing what to prioritize, not working the list; Qualaix founder Marva Bailer says discovery is only step one, since organizations still have to assess exposure, test patches and deploy across thousands of devices, and AI puts more pressure on the window between discovery and deployment; Fortra's Tyler Reguly says that as long as Microsoft is catching up on patches, the raw vulnerability count has lost its meaning.

Source: InfoQ China · InfoQ · OpenAI

05/09

Anthropic and OpenAI look at 20-30 MW data centers

CNBC reports both labs are exploring smaller powered sites to bring inference capacity online sooner.

CNBC reported on September 18th that Anthropic and OpenAI are exploring data center deployments of roughly 20-30 megawatts to get capacity online faster. Four people familiar with the discussions say Anthropic has sounded out arrangements in the U.K. and Nordic countries; two sources say OpenAI explored opportunities in the Nordics, and one described conversations involving both labs about U.S. deployments at the same scale.

Anthropic's leadership page says compute is a co-founder-level function there: co-founder and chief compute officer Tom Brown runs the technical organization responsible for securing, scaling and using compute resources.

Limitations: these are discussions about potential capacity rather than completed transactions, and the source gives no sites, power providers, investment figures or timelines. The account rests on anonymous sources.

Source: RuntimeWire · Anthropic

06/09

Huawei's OceanStor M900 pools KV cache at petabyte scale

AI memory storage aimed at hyperscale inference, with up to 64 PB of KV cache per cluster.

  • KV cache capacity: up to 64 PB per cluster, pooled over Lingqu
  • Latency: about 60 microseconds from NPU to SSD
  • Aggregate bandwidth: about 40 TB/s
  • Scheduling: KV-aware scheduling

Limitations: these are vendor specifications. Pandaily's summary gives no measured throughput, no supported models or cluster sizes, no pricing or availability date, and no independent benchmark.

Source: Pandaily

07/09

PrismML raises a $22.25M seed round betting on small models

TechCrunch says the lab's wager is that capable reasoning LLMs do not have to be large.

TechCrunch reports that AI lab PrismML has so far raised only a $22.25 million seed round, and argues the reasons to watch it are the technical people involved and the direction of the work: PrismML is betting that capable, high-performing, reasoning large language models do not, in fact, have to be large.

Limitations: the accessible report names no model, parameter count, context length, benchmark score or release date. For now this is a directional signal, not a shipping product.

Source: TechCrunch

Regional and Early Signals

08/09

Zhipu's ZCode admits its indexing feature could upload repository data

The company apologized, says the issue is fixed, and will open-source the codebase and bring in third-party review.

On September 18th, ZCode, the coding product from Zhipu, apologized to affected users through its official group and said it had completed an internal review and fixed the issue. According to the response, the problem came from the "codebase indexing" feature, which builds a repository index locally to support session checkpoint restore, version rollback and Repo Wiki; generating a Wiki page in the cloud could trigger an upload of repository data.

Zhipu says uploaded data is destroyed immediately after the page is generated in the cloud and is not retained, and that users were affected because the feature was on by default at launch. It also announced it will open-source the ZCode codebase, invite third-party evaluators to review how the system operates, and publish the review's progress.

Limitations: the company has not disclosed how many users were affected, what data was uploaded or over what time window. "Destroyed immediately, not retained" is the vendor's own account, and the third-party review has not started. Chinese-language source.

Source: ITHome (Chinese-language source)

09/09

Helix 2.5 enters 30 unfamiliar homes zero-shot at a 56% success rate

No on-site data collection and no task-specific fine-tuning: 237 successes in 420 trials.

Figure AI released Helix 2.5 on September 17th and ran a zero-shot test in 30 real Bay Area homes: unfamiliar objects, unfamiliar room layouts, no prior data collection and no targeted fine-tuning, on chores such as folding towels, making beds and picking things up. The evaluation was blind, with 237 successes across 420 trials.

PolicyZero-shot success rate
Helix 2.5 (with Index pretraining)56%
Same hardware and task data, no Index pretraining9%

The source adds two figures: overall success improved 522%, and no single evaluation task accounted for more than 1.90% of the Index pretraining data. Against the previous generation, Helix 02 had to collect data on-site before training, while Helix 2.5 matched Helix 02's single-site success rate across 30 unseen homes using half the adaptation data.

Limitations: 56% also means nearly half the trials failed. The test is confined to Bay Area homes and tidying-type tasks, and the source gives no breakdown of failures or hardware details.

Source: ifanr (Chinese-language source)

Generate one yourself with JIUFENG

Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.

Generate free