AI Highlights

Claude Opus 5.5 wins a five-model 3D-printed bridge test

Key Takeaways

Claude Opus 5.5 topped a five-model 3D-printed bridge test at ~130 lb, GPT-6 Astra tidied an unfamiliar kitchen, and Ember-1 cut Kimi K3 tokens by 40%.

jiufeng
September 28, 2026
35 min read
In this article

Overview

10 stories in this issue. The first 3 are today's priorities.

Model watch

  1. Top · Claude Opus 5.5 tops a five-model 3D-printed bridge test
  2. Top · GPT-6 Astra drives a robot through an unfamiliar kitchen
  3. Top · Ember-1 trims Kimi K3 reasoning tokens by about 40%
  4. Atria Dawn Preview log: agents propose, humans decide

Global AI news

  1. Jev picks Minecraft moves in 24 milliseconds
  2. TinyAIArena puts four models in an 8x8 grid fight
  3. Engram turns local-model hallucinations into sound

Regional and early signals

  1. Anonymous model Space Bunny tops OpenRouter's daily call chart
  2. FermiQLLM 1.0 adds quantum methods to a Qwen base
  3. Doubao phone agent users get kicked out of Honor of Kings
AI signal map for 2026-09-28

Jiufeng graphic based on the sources cited in this issue.

Model watch

01/10

Claude Opus 5.5 tops a five-model 3D-printed bridge test

Anthropic's Claude Opus 5.5 came first in a five-model 3D-printed bridge challenge, with its bridge holding an estimated 130 lb.

Roberto Nickson (@rpnickson) reported the results in a September 27th post on X. The lineup was Anthropic's Claude Opus 5.5, Moonshot AI's Kimi K3, Meta's Muse Spark 1.3, OpenAI's GPT-6 Astra and SpaceXAI's Grok 4.7, all listed at "High" reasoning effort. Each bridge had to span two feet, use no more than 500 grams of filament and print in under 18 hours; Nickson specified the type of weight and where it would sit on the bridge.

ModelReported load
Claude Opus 5.5~130 lb
Muse Spark 1.326.5 lb
GPT-6 Astra~17.5 lb
Kimi K3did not complete
Grok 4.7did not complete

RuntimeWire writes that in Nickson's comparison Claude led on both load capacity and printing efficiency.

Limitations: this is one person's test under conditions he set, and the report says the results describe these designs under his test conditions rather than a universal ranking of the models' engineering ability. The 130 lb and 17.5 lb figures are explicitly labelled estimates, the results are Nickson's own account, and his X post carries an approximately 15-minute video.

Claude Opus

Image source: anthropic; mirrored on Jiufeng R2.

Source: RuntimeWire · Roberto Nickson on X · Claude Opus 5.5

02/10

GPT-6 Astra drives a robot through an unfamiliar kitchen

Stanford and Caltech researchers wired GPT-6 Astra straight into a Unitree G1, dropping the trained control layer in between.

The system is called HomeBody. It replaces the usual trained control layer between language model and robot with a swappable vision-language model — here GPT Astra — that calls directly into an extensible skill library for grasping, navigating and opening drawers. The robot explores the room first, builds a digital twin in Nvidia's Isaac Sim, and logs objects and locations in spatial memory, so it can still find items after they leave its field of view. For a task like "clean up the kitchen," the language model plans each step and self-corrects on errors.

Limitations: the researchers list Astra's latency, overheating finger servos and high compute costs as constraints. This is a research system, and the reported behaviour rests on the team's own account, with no independent reproduction.

Source: The Decoder · HomeBody project repository · Nvidia Isaac Sim docs

03/10

Ember-1 trims Kimi K3 reasoning tokens by about 40%

Fireworks says its Kimi K3-derived model gives comparable answers while generating roughly 40% fewer tokens.

Fireworks AI co-founder Dmytro Dzhulgakov (@dzhulgakov) promoted Ember-1 in a September 27th post on X. The company introduced the model on September 23rd, positioning it as a way to cut the cost of reasoning-heavy coding and agent workloads, and its launch materials back the token-reduction claim with benchmark and customer-test results. Dzhulgakov called Ember-1 "hot on HN" and said it was 40% faster and cheaper. He is one of seven Fireworks co-founders, and the company's team page identifies him as a former PyTorch core maintainer at Meta.

Limitations: RuntimeWire notes that the launch materials do not establish a general 40% improvement in response speed — fewer generated tokens can cut per-task cost without cutting latency across workloads. The 40% figure is vendor-reported and workload-specific.

Source: RuntimeWire · Dmytro Dzhulgakov on X

04/10

Atria Dawn Preview log: agents propose, humans decide

A team including Fudan University researchers audited its own model-building logs: agents supplied up to 55% of method proposals, but humans made more than 85% of final decisions.

  • Sample: 769 task logs from 56 participants, plus logs from the agents they used
  • Model built: Atria Dawn Preview, a mixture-of-experts agentic language model with 744 billion parameters, aimed at research and engineering tasks
  • Training: every task is tied to a real execution environment — it calls tools, produces intermediate results and is checked against external signals such as tests, metrics or source evidence
  • Self-reported scores: the team says it leads on five of 16 benchmarks, including web search and cybersecurity

Limitations: this is a team studying its own project, not an independent observation; it does not claim the top spot on the other 11 benchmarks, and all of the figures come from the team's own account.

Source: The Decoder · Atria Dawn Preview

Global AI news

05/10

Jev picks Minecraft moves in 24 milliseconds

Hao AI Lab posted a Minecraft combat demo in which TypeSafe's Jev is claimed to choose a move about every 24 milliseconds.

The demo went up on September 27th. The lab says the system makes up to 40 decisions a second while running DJev on an Nvidia B200 GPU, and describes that as faster than human reaction time. The design premise is to hand software a set of predefined choices and let a model score them quickly, rather than ask a general-purpose chatbot to generate a response. Implementation credit goes to Matt Mastracci (@mmastrac), whose DJev repository describes a Jev-style structured-decision server. TypeSafe closed a $40 million seed round led by DCVC on September 15th.

Limitations: the post publishes no match results, controlled comparison or details of the opponents, so the 24-millisecond figure and the claim that Jev is hard to beat are the lab's assertions rather than independently measured benchmarks. Fast choices among predefined moves do not establish reliable performance in production software.

Source: RuntimeWire · DJev repository

06/10

TinyAIArena puts four models in an 8x8 grid fight

A Show HN project stages life-or-death matches between four models on an 8x8 grid, with every match open to spectate.

hp6 posted TinyAIArena to Hacker News, framing it against the disappointment of clicking an "AI Arena" and getting a benchmark table instead of a fight. Any match on the page can be opened and watched, the code is published in the ai-arena repository, and the README spells out the rules:

  • Winning: the last model standing wins, and the server referees illegal moves
  • Actions: turn order is randomized each round; moving, attacking (15–24 damage) and waiting each cost 1 action point
  • Map and pickups: four obstacle squares are placed at random on the 8x8 board; a coin grants +1 AP, and a kill grants +1 AP plus 50 HP
  • Ranking: the leaderboard is computed by Elo

The post drew 93 points and 40 comments within about eight hours.

Limitations: outcomes and rankings come from the project's own server referee and Elo leaderboard, not from independent evaluation. The evidence here is the author's own description plus the repository; the Hacker News thread is a lead, not a source.

Source: Show HN thread · ai-arena repository

07/10

Engram turns local-model hallucinations into sound

Thoughtful Things opened a Kickstarter for Engram, a sampler that uses AI to mangle incoming audio and hallucinate entirely new sounds.

Engram is a sampler and groovebox that the company calls a "field recorder for latent space" and describes as circuit-bending tiny AI models. It deliberately is not a push-button-get-song device: the target is experimental, uncanny sound, pushing AI audio models past their limits. The unit is not connected to the internet and runs a "tiny AI" locally; per the Kickstarter listing, the model is designed in-house and custom-trained.

Limitations: this is crowdfunding-stage hardware, so the finished behaviour can only be judged after delivery, and the sound depends on that in-house model. The available evidence is the company's own description and its Kickstarter listing.

Source: The Verge

Regional and early signals

08/10

Anonymous model Space Bunny tops OpenRouter's daily call chart

An unclaimed anonymous model reached number one on both OpenRouter's and OpenCode's daily call charts within days, and 量子位 ran a hands-on coding test.

量子位 pulled the API from OpenRouter, wired it into DeepSeek Harness and ran four randomly picked cases. Asked to build an interactive cube-style Temple of Heaven in Three.js, the model finished within one to two hours and added details nobody requested: a bottom control bar with night, dawn and noon presets plus rain and smoke toggles, a day-night cycle of about 3.5 minutes, and Chinese double-hour labels running from 子时 to 亥时. A second task, an Apple-style alarm clock, gained a per-month ring toggle and a 7:23 wake time. In a third task, the first version of an app window could not be dragged and had no close or minimize buttons.

Limitations: the model is still anonymous, so parameters, pricing and ownership wait on a vendor claiming it, and both charts rank call volume rather than capability. The test is one person's uncontrolled run, and the model itself said synthetic mouse events are partly dropped in that environment, so it could not assert pointer-synchronous dragging in a headless setup. Chinese-language source only (量子位); no English-language report covers this.

Source: 量子位 (Chinese-language source) · @Fei2411 on X

09/10

FermiQLLM 1.0 adds quantum methods to a Qwen base

Tsinghua-linked startup 费米宇宙 says it has built a full-pipeline quantum-enhanced large model on top of an open Qwen base.

  • Company: 费米宇宙, focused on Q4AI (Quantum for AI); the project started in March 2026 with a 4B pilot model and the company was formally founded in May
  • Funding: an RMB 100 million seed round at a post-money valuation of about RMB 1 billion, with the money still being transferred
  • Approach: tensor networks, quantum simulated annealing and gauge degrees of freedom are embedded across data representation, architecture, training, reinforcement and evaluation; no fault-tolerant quantum computer is needed, and training and deployment run on existing GPUs
  • Internal numbers: against a conventional model of comparable size, inference performance up more than 15% and continual-RL training cost down more than 25%; on MATH-500, GPQA-Diamond and BBH, overall gains of 10%–20% versus the same-size open base

Limitations: every figure comes from company-internal testing reported exclusively by 量子位, with no public weights, technical report or third-party reproduction, and neither the parameter count nor the base version is disclosed. Chinese-language source only (量子位).

Source: 量子位 (Chinese-language source)

10/10

Doubao phone agent users get kicked out of Honor of Kings

On the nubia NaviX Ultra with Doubao's phone assistant installed, logging into Honor of Kings triggers a device-environment warning and a forced sign-out, and the two sides have not reconciled their accounts.

Per IT之家 on September 27th, ByteDance's Doubao phone assistant posted a statement in its official community on September 25th: from the evening of September 24th it received multiple user reports that on this handset, logging into Honor of Kings or entering matchmaking produced a device-environment-abnormal prompt and a forced sign-out. The statement says the assistant performed no operation on Tencent's game system and that the AI made no illegal clicks and ran no cheat or simulation behaviour; the team was still in contact with Tencent and had received no clear reply. It advised users not to retry logins on that handset repeatedly and to appeal banned accounts through official game support. The nubia NaviX Ultra is billed as the second-generation Doubao phone and starts at RMB 5,999.

Limitations: the accounts do not line up — Doubao says it touched nothing in the game system, Tencent has issued no public response, and a person familiar with the matter says the security-risk policy exists to keep matches fair and was not adjusted to target anything. No figures on affected users or bans are public. Chinese-language source only (IT之家).

Source: IT之家 (Chinese-language source)

Generate one yourself with JIUFENG

Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.

Generate free