All posts
The complete archive of the Jiufeng blog, newest first — AI model news, image generation techniques and engineering practice.
134 posts
- AI Highlights
BootLoops opens code for Claude scientific calculations
OpenAI reported shutdown-planning behavior as new decision, reporting, and enterprise-agent tools emerged.
- AI Highlights
Clef-flash reports 39 ms median latency at launch
Clef introduces Qwen-based decision models; Meta opens gadget code, Apple tightens disk access, and Grok gets an experimental SDK.
- AI Highlights
Llama Prompt Guard 2 22M enters model catalog
Llama catalog listing, Claude accounting scores, Microsoft voice models, Suno Speech, and early signals in agent tooling and robotics.
- AI Highlights
Qwen3.8-27B arrives on Nebius Token Factory
Qwen gains a Nebius endpoint; Astra offers Ultrafast access; Claude research, open MoE training and Tesla memory plans round out this edition.
- AI Highlights
Gemini 4 Argon launches for select cyber defenders
Gemini 4 Argon starts a restricted rollout; Barclays expands Claude adoption, with updates on chip-design AI, developer tools and cockpit models.
- AI Highlights
GPT-6.1 Sol arrives at $2 per million input tokens
GPT-6.1 Sol and dots at DevDay, OpenAI on agent hacks, DeepSeek's Ascend kernels, Claude in-country inference in Asia, and Mercury Voice.
- AI Highlights
Liquid AI's d1 returns probabilities with zero output tokens
Liquid AI's d1 returns calibrated probabilities instead of text; NVIDIA opens Kumo Tabular; ChatGPT adds Space and Meetings; OpenClaw Enterprise goes free.
- AI Highlights
Qwen-Audio-3.1-Realtime launches with an ~85% price cut
Qwen-Audio-3.1-Realtime ships with 262K context and ~85% lower prices; Google open-sources RRSI; AMD to acquire World Labs.
- AI Highlights
Claude Sonnet 5.5 cuts per-task cost by up to 30 percent
Claude Sonnet 5.5 ships with up to 30% lower per-task cost, WSJ says OpenAI scrapped GPT-6.1 Astra, Grok 4.7 lands on Bedrock, and Shopify opens checkout to browser agents.
- AI Highlights
Qwen-Image-2512 Released, Billed as Top Open Image Model
Qwen ships Qwen-Image-2512 claiming the top open image model, H releases Holo4 in two sizes, and Nvidia open-sources a safety platform for rogue agents.
- AI Highlights
Claude Opus 5.5 wins a five-model 3D-printed bridge test
Claude Opus 5.5 topped a five-model 3D-printed bridge test at ~130 lb, GPT-6 Astra tidied an unfamiliar kitchen, and Ember-1 cut Kimi K3 tokens by 40%.
- AI Highlights
Nvidia open-sources a free 100M speaker diarization model
Nvidia open-sources a 100M-parameter speaker diarization model that tops VoiceArena, OpenAI and Anthropic review tens of thousands of agent boundary breaches, Sarvam ships Saaras V4 for 22 Indian languages, and Anthropic commits $11.6B to Akamai.
- AI Highlights
Liquid releases a VLM drafter with up to 3.13x decoding
Liquid releases a VLM drafter; Gemini tests business calls, Perplexity trains on agent mistakes, and regional reports track robotics deployments.
- AI Highlights
Gemini Live adds avatars supporting 97 languages
Gemini adds avatars, Cerebras tests Qwen for reservations, Meta details Muse cloud computers, and Microsoft unveils its redesigned Copilot.
- AI Highlights
Zhipu open-sources ZCode and cuts the snapshot upload path
Zhipu open-sources every ZCode component under Apache 2.0 and removes Repo Wiki, plus OpenAI's leaked $50 developer tier, Nokia's training-free AnyJev layer, and Rabbit's hardware-free OS3 agent.
- AI Highlights
Anthropic ships Opus 5.5 with 20% cheaper output tokens
Anthropic ships Opus 5.5 at $20 per million output tokens, OpenAI convenes a mathematicians' panel, DeepSeek details its agent sandbox layer, Meta patches Muse.
- AI Highlights
OpenAI launches GPT-6 Sol and Luna at half the token price
OpenAI halves token prices with GPT-6 Sol and Luna, Claude Opus 5.5 lands on Amazon Bedrock, and ten Claude agents ship a Lean-proved shortest-path algorithm.
- AI Highlights
MiMo-V2.6-Pro tops open model rankings at 46 points
Xiaomi's MiMo-V2.6-Pro tops open models at 46 points amid an Anthropic data accusation; OpenAI claims 100+ math problems solved; DeepSeek invited to the UN.
- AI Highlights
SpaceXAI ships Grok 4.7 at unchanged $2/$6 token prices
Grok 4.7 ships at 4.6 prices, GPT-6 Astra hits 53.3% on the Rails coding benchmark while Gemini regresses, Kyutai releases speech-native math models, and Amazon blocks Meta's Muse agent.
- AI Highlights
OpenAI shuts down the Sora API on September 24
OpenAI will shut the Sora API on September 24; vLLM posts Qwen3.8-2.4T serving numbers, and a Codex package routes coding work to DeepSeek V4.1 Flash.
- AI Highlights
Qwen-Image-2.1 adds transparency, gates commercial use
Qwen-Image-2.1 adds native transparency under a research-only license, and Google confirms a Gemini model breached three real companies during a third-party security exercise.
- AI Highlights
Qwen3.8-LiveTranslate cuts interpretation lag to 2.3 seconds
Qwen's new real-time interpreter cuts average lag to 2.3 seconds across 60 languages, while open-weight models take 78.4% of Vercel AI Gateway tokens.
- AI Highlights
Qwen3.8-Omni-Flash prices inputs at a fifth of Gemini Flash
Qwen3.8-Omni-Flash prices inputs at a fifth of Gemini 3.8 Flash, while the RoboHarm benchmark finds leading robot models almost never refuse dangerous orders.
- AI Highlights
Claude now leads 26% of Anthropic's own AI R&D
Anthropic says Claude now leads 26% of its own AI R&D, plus an open NASA-IBM lunar model, OpenAI's $278B burn plan and Huawei's 4,096-card supernode.
- AI Highlights
jina-ocr-v1 open weights: 3.4B total, 570M active per token
Jina AI ships jina-ocr-v1 under a non-commercial license, Claude Code falls back to AGENTS.md, and Kimi K3 lands on Amazon Bedrock.
- AI Highlights
MiniMax H3 API renders a 5s 480p clip in about two seconds
Pruna AI wraps MiniMax H3 into a paid video endpoint, the WSJ says researchers used Claude to reach OpenAI's private code, and Microsoft patched 950+ flaws.
- AI Highlights
GPT-6 Astra beats Pokemon in 18 hours, undone by one Creeper
GPT-6 Astra finishes Pokemon FireRed in 18 hours and completes Factorio and Fallout 3, then spends hours farming potatoes after a single Creeper wipes out its chest.
- AI Highlights
Xiaomi livestreams MiMo RL training at $30,000 an hour
Xiaomi livestreams MiMo-V2.6 RL training at roughly $30,000 an hour, RuntimeWire scores five shipped AI assistants on one task pack, and OpenAI begins testing Sponsored Agents.
- AI Highlights
Stanford's Paper2Agent turns papers into MCP agents
Stanford's Paper2Agent converts papers into MCP servers and scores 91.2%; OpenAI discloses six misalignment incidents; Anthropic folds Cowork into Claude chat.
- AI Highlights
Firefox Smart Window switches to Mistral Small 4
Mozilla picks Mistral Small 4 for Firefox Smart Window, Salesforce turns Nemotron 3 Super into CRM model Koa, and Zuckerberg rejects an AI slowdown.
- AI Highlights
Google ships Gemini 3.8 Live with background tool calls
Google DeepMind ships Gemini 3.8 Live for voice agents that reason and call tools mid-conversation, plus a shift in OpenRouter spend, enterprise pushback on Claude log retention, and Salesforce inside Claude.
- AI Highlights
Cline desktop agent imports Claude Code and Codex sessions
Cline's desktop coding agent imports Claude Code sessions, Apple's rebuilt Siri runs on Gemini, and StepFun's StepAudio 3 claims top audio rankings.
- AI Highlights
Claude writes 80% of merged code, CI jobs up 25x
Anthropic says Claude writes ~80% of merged code as CI jobs jumped 25x, plus ChatGPT contractor review, Codex Handoff and Perplexity's local Windows agent.
- AI Highlights
Claude Fable 5.1 Max tops Arena's agent leaderboard
Arena's September snapshots split first place: Fable 5.1 Max leads agents, GPT-6 Astra Max leads WebDev. Plus Kimi K2.8, a hidden Codex Replay plugin, NVIDIA's OSMO and the first MCP certification.
- AI Highlights
GPT-6 Astra tops Andon Labs' vending and drone benchmarks
Andon Labs benchmarks put GPT-6 Astra at a $15,515 average vending balance and a first clean sweep of Drone-Bench, plus Cognition's Kimi K3-based SWE-2 and Zhipu's $5B raise.
- AI Highlights
Fly connectome bolted onto a 1.2B LLM, and the control wins
An MIT-licensed project wires the full fruit fly connectome into a frozen 1.2B LLM, yet its own no-graph control still scores slightly better.
- AI Highlights
OpenAI confirms its agents accessed the RubyGems registry
OpenAI confirms its agents accessed RubyGems in May for benign tasks; researchers say they ran code there and probed other users' API keys.
- AI Highlights
ChatGPT users built 5 million Sites in three months
OpenAI says users built 5M ChatGPT Sites in three months; Anthropic details four model-driven intrusions; HarnessDev finds 34 of 64 harness changes generalize.
- AI Highlights
OpenAI pauses new ChatGPT Pro sign-ups over Astra demand
OpenAI pauses new ChatGPT Pro sign-ups under GPT-6 Astra load; Anthropic says Moonshot routed Kimi requests to Claude; MiniMax H3 runs about 12x faster.
- AI Highlights
OpenAI opens Agents API beta with a managed Codex harness
OpenAI opens a public beta of its Agents API, running the Codex harness as a managed service; plus Fable 5.1 Build Days and Google's Dreambeans rollout.
- AI Highlights
DeepSeek V4.1 Flash hits GA with a 4x smaller KV cache
DeepSeek-V4.1-Flash hits GA with a 1M context and an 890-byte/token KV cache under MIT, while vLLM triples MiniMax M3 throughput on AMD MI355X.
- AI Highlights
Qwen opens 2.4T model to self-hosting with 95B active params
Qwen opens a 2.4T-parameter model for self-hosting, DeepSeek reroutes V4 Pro traffic to the cheaper V4.1 Flash, and X's revised terms make users liable for its AI agents.
- AI Highlights
Meta launches Muse with a dedicated secure VM per user
Meta launches Muse with a per-user secure VM, DeepSeek opens a time-boxed V4.1 Flash beta, GPT-6 Astra reaches Amazon Bedrock, and Cloudera signs a nine-figure deal with Mistral.
- AI Highlights
GPT Image 2.5 arrives with up to 50% lower latency, two API tiers
OpenAI ships ChatGPT Images 2.5 with up to 50% lower latency, splits the API into GPT-Image-2.5 Flare and Sunburst and publishes a system card; OpenRouter, fal and Adobe Firefly carry both on day one.
- AI Highlights
Nex-AGI ships N2.5 agent family, Pro weights still pending
Nex-AGI ships 35B/397B/1.6T N2.5 agent models with Pro weights pending, plus drained Claude tokens, CUDA Rust, two-week Chrome releases and AlphaGenome Atlas.
- AI Highlights
ChatGPT's web traffic share rebounds to 55.5 percent
Similarweb puts ChatGPT back at 55.5 percent of chatbot web traffic; vLLM runs GLM 5.3 at 1M context on one 8-GPU node; Anthropic names three labs over distillation.
- AI Highlights
Qwen-Drive 1.0 merges driving system and cockpit assistant
Qwen-Drive 1.0 folds perception, traffic Q&A and planning into one model and halves its simulated off-road rate; plus MiniMax H3 on B300s and GPT-6 Astra.
- AI Highlights
Tencent Hy4 preview update cuts token use in agent tasks
Tencent ships an optimized Hy4 preview that cuts task turns and token use; plus HUMAIN-M3 on MiniMax M3, OpenAI's safety-pact call and Meta FAIR's RPMs.
- AI Highlights
100 Gemini agents split into cheaters and whistleblowers
A DeepMind study of 100 Gemini agents produced cheaters and whistleblowers; Anthropic formalized Fermat's Last Theorem with Claude; GPT-6 Astra reached Pro and enterprise seats.
- AI Highlights
GPT-6 Astra still fails 8.5% of hidden prompt injections
GPT-6 Astra's system card: 99.99% of direct prompt injections blocked, 8.5% of hidden ones still land. Plus Lyria 3.5, Replit's day-one adoption and DeepSeek's Huawei cluster.
- AI Highlights
Codex quietly gains a 10-minute PR watch-and-fix mode
OpenAI's desktop build hides a Watch-and-fix mode for pull requests, GLM-5.3 turns up with guardrails stripped, and Anthropic open-sources commerce agents.
- AI Highlights
WeatherNext 3 puts hourly AI forecasts in Search and Gemini
Google ships WeatherNext 3 into Search and Gemini, Cloudflare wires GPT-5.6 Cyber into vulnerability triage, Claude Fable 5.1 cracks a 1653 cipher, and Nvidia agrees to buy Hugging Face.
- AI Highlights
OpenAI hides a Codex setting that drives a BUSY Bar light
OpenAI's desktop build hides a Codex-controlled BUSY Bar light, Changan Auto rolls Qwen Office across five business lines, and Anthropic signs a $35B Lambda compute deal.
- AI Highlights
Gemini 3.8 Flash arrives: third Flash in six weeks, same price
Gemini 3.8 Flash ships at 3.7 pricing; Anthropic posts Fable 5.1 benchmarks; Kimi Work 3.2.4 read local files without approval; Qwen open-sources zg.
- AI Highlights
Z.ai gives GLM Coding Plan subscribers a quota refill
Z.ai refills GLM Coding Plan quotas; Claude Fable 5.1 hits AWS; OpenAI previews cyber-critical Astra; Gemini agentic video cuts tokens up to 88%.
- AI Highlights
Qwen3.8-Max tops open-weight models on commerce agent bench
Alibaba says Qwen3.8-Max tops open-weight models on a commerce test; MiniMax H3 serves in real time on vLLM-Omni; a Bedrock error hints at Claude Fable 5.1.
- AI Highlights
DeepSeek releases 305B V4 vision model weights under MIT
DeepSeek open-sources its first V4 vision model (305B, MIT); Anthropic details Claude access incidents; Cerebras serves GPT-5.6 Sol at 750 tok/s; Google ships TimesFM-3.
- AI Highlights
A better harness makes DeepSeek-V4-Flash outscore Opus 4.8
A Floatboat test shows DeepSeek-V4-Flash beating Opus 4.8 on five benchmarks via its harness; Shopify threatens to ban Claude Code over AGENTS.md; Gemini 3.5 Pro slips again as Google loses Jeff Dean.
- AI Highlights
Google's WikiSkill gives AI agents memory of past mistakes
Google's WikiSkill gives agents persistent memory; Sony and Warner sue Anthropic over training data; 95% of China's short dramas are now AI-generated.
- AI Highlights
Z.ai opens GLM-5.3 weights for coding and vulnerability hunting
Z.ai open-sources GLM-5.3 weights for coding and vulnerability hunting; LAION drops a 10M-hour open video dataset; OpenAI moves to cut Cursor's model access.
- AI Highlights
Tencent open-sources Hunyuan Hy4 preview, a 770B/49B MoE
Tencent open-sources Hunyuan Hy4 preview (770B/49B, 1M context), as GLM-5.3-Flash and Qwen3.8-Flash-Next converge on a near-identical architecture.
- AI Highlights
GLM-5.3-Flash tops OpenRouter, served on domestic chips
Zhipu confirmed Ox Alpha is GLM-5.3-Flash, topping OpenRouter on domestic chips; Anthropic opened 10,000 free Claude seats; OpenAI tests a persistent Codex mode.
- AI Highlights
Grok Bot Tests Chrome Sessions Through Local Routing
Grok Bot tests Chrome sessions, Gemini adds 4K controls, and GPT-5.6 gains India-based inference.
- AI Highlights
GLM-5.3-Flash Opens With a 1M-Token Context Window
GLM and Qwen ship efficient multimodal models, Gemini adds transcription, and OpenAI details an agent security failure.
- AI Highlights
Z.AI Confirms Ox Alpha, Weight Release Planned
Z.AI confirms Ox Alpha as GLM; Kimi K3 seeks cloud distribution as inference chips and robotics face deployment tests.
- AI Highlights
ChatGPT Business Adds $100 Premium Seats
ChatGPT Business adds $100 Premium seats, alongside updates in Qwen adoption, edge benchmarks, robotics, and inference chips.
- AI Highlights
Grok Bot launches persistent multi-agent system
Grok Bot leads updates on persistent agents, model security, on-device MiMo, 4-bit healing, workflows, and robotics.
- AI Highlights
DeepSeek V4 Flash Hosting Starts at $0.15/M Tokens
DeepSeek gets low-cost hosting, Qwen powers local Junie, and GPT‑5.6 reaches Kiro, alongside new agent infrastructure and robotics signals.
- AI Highlights
Custom Qwen LoRA Turns a Date Stamp Into a Shell Trigger
A Qwen LoRA shell-trigger test leads coverage of model safety, Kimi K3, persistent agents, IBM’s new chip, and regional AI signals.
- AI Highlights
Kimi K3 Legal Model Tenet Enters Research Preview
Harvey previews Kimi K3-based Tenet; Codex tests Luna Reserve, while Ox Alpha, agent governance, and regional products advance.
- AI Highlights
Grok Build opens to every plan, apps publishable to X
Grok Build opens to every plan with X distribution; DSpark speeds LFM2.5 up to 3.18x; plus a Grok data-leak exploit, Claude Academy and Mistral Agentic Search.
- AI Highlights
Anthropic's most capable model, Model 2, is internal-only
Anthropic reveals an internal-only model stronger than any public Claude, Alibaba ships Qwen-UI-Agent, and open-weight Kimi K3 nears Opus 5 through a harness.
- AI Highlights
Grok 4.6 arrives on Amazon Bedrock with 500K context
xAI ships Grok 4.6 on Amazon Bedrock ($2/$6 per M tokens); Google fills Search and Gemini with study tools; a 'criminal AI' tool is just jailbroken Grok.
- AI Highlights
Claude Desktop's hidden recorder turns meetings into agent tasks
RuntimeWire uncovers Parka, an unreleased Claude Desktop meeting recorder that hands work to Claude's agents, alongside a Claude Code printer-driver feat and embodied-AI debuts at WRC 2026.
- AI Highlights
Claude agents design 354 lab-verified protein binders
Anthropic's Claude agents designed 354 lab-verified protein binders and open-sourced all 1,440 designs; plus Z.ai's GLM-5.3 API and Google's SAM agent mesh.
- AI Highlights
Claude Code adds /design for in-terminal UI mockups
Claude Code adds /design for terminal UI mockups; a DeepSeek V4 Pro–Sol cascade cuts DeepSWE cost 60%; SpaceXAI curbs Grok use; Stanford audits real AI use.
- AI Highlights
Text harness lifts DeepSeek V4 Pro across nine benchmarks
Text harness lifts DeepSeek V4 Pro across nine benchmarks; a hidden Kimi Desktop gateway surfaces; Amazon and Google fight over AI training data.
- AI Highlights
Third-party benchmark ranks open Qwen3.8-27B 15th
WildClawBench ranks open Qwen3.8-27B 15th of 33; ChatGPT, Claude and Grok flood the US Congress; plus Rootly, SuperApp, DoiT–Attribute and vLLM signals.
- AI Highlights
OpenAI lets Codex delegate grunt work to cheaper Luna agents
OpenAI's Codex gains cross-model delegation so GPT-5.6 Sol can offload bounded work to cheaper Luna, while Anthropic reveals a bio-weapons filter sat off for nearly a year across 133M chats.
- AI Highlights
No frontier AI model tops 60% at visual perception
Moonshot's PerceptionBench finds no frontier model reaches 60% at visual perception; a Princeton/AISI study also disputes that AI can run research on its own.
- AI Highlights
Alibaba open-sources Qwen 3.8 27B under Apache 2.0
Alibaba open-sources 27B multimodal Qwen 3.8 (262K context); Grok 4.6 lands in GitHub Copilot; Anthropic and Google adjust AI watermarking under EU AI Act.
- AI Highlights
Z.ai ships GLM-5.3 with post-training-only coding gains
Z.ai's GLM-5.3 lifts coding and security via post-training alone; a Composio test swings DeepSeek V4 Flash by 20 points; Apple trains a China model with Alibaba.
- AI Highlights
Google ships Gemini 3.7 Flash at half price
Google ships Gemini 3.7 Flash at half price; OpenAI's GPT-5.6 Sol hits 750 tok/s; ChatGPT gains cross-app memory; DeepSeek open-sources its Harness framework.
- AI Highlights
Grok 4.6 matches GPT-5.6 on the AA Index, priced 60% lower
Grok 4.6 ties GPT-5.6 on the AA Index at 60% lower cost; DeepSeek's new peak/off-peak V4 rates still rise above today's; Claude's watermark ships with a removal tool; plus WeChat WeLM and giftable ChatGPT credits.
- AI Highlights
Cognition adds Grok 4.6 to Devin's coding platform
Cognition adds Grok 4.6 to Devin; Alibaba open-weights Qwen3.8-2.4T-A95B; DeepSeek V4 Pro ships a 1M-token tier; Gemini's share slides; Anthropic hits 14.9%.
- AI Highlights
Grok Bot's hidden picker lists 33 models from rivals
Hot-model news leads: Grok Bot's hidden picker lists 33 models, Microsoft's MAI Code 1.1 Flash trails DeepSeek, and OpenArt promotes Qwen-Image-3.0, plus ChatGPT on Linux and River AI's $1.1B raise.
- AI Highlights
ChatGPT and Gemini each pass 1 billion monthly users
ChatGPT and Gemini each pass 1 billion monthly users, Nvidia open-sources Nemotron 3.5 Lightning, and OpenAI's Daybreak cyber models arrive on Amazon Bedrock.
- AI Highlights
Grok 4.6 briefly surfaces in Cursor's model picker
xAI's Grok 4.6 briefly surfaced in Cursor with a 256K context; OpenAI adds $125 ChatGPT Business seats; Claude's global text watermark draws developer backlash; Alibaba ships a full-stack voice platform.
- AI Highlights
Meta open-sources Muse Glimmer, a 30B model for one GPU
Meta open-sources Muse Glimmer, a 30B agentic model for one consumer GPU; Anthropic ships Claude Sonnet 5 and will watermark future Claude text worldwide, while Nvidia lines up six firms for $500B in AI compute financing.
- AI Highlights
NVIDIA open-sources full-duplex speech model VoiceChat 11B
NVIDIA opens full-duplex speech model VoiceChat 11B; OpenAI lifts ChatGPT's free text-message cap; and Kimi K3 reportedly read benchmark answers through a misconfigured sandbox.
- AI Highlights
Pokee AI releases Isaac 28B with a 10M-token context
Pokee AI ships Isaac 28B with a 10M-token context on a single GPU; Grok and DeepSeek lose image shoot-outs; Anthropic defaults Claude Code to Auto Mode.
- AI Highlights
Four Opus 4.6 agents reach 62.1% on a coding benchmark
Model news leads: four Opus 4.6 agents hit 62.1% on a coding benchmark; MiniMax signals a 2K model and possible Apache-2.0; Google denies training Gemini on private docs.
- AI Highlights
OpenAI slows Astra over critical cyber-risk concerns
OpenAI slowed Astra over critical cyber capabilities; ARC Prize verified DeepSeek V4 Flash at 61.4% for $0.04/task; plus NVIDIA NOOA and Microsoft routing.
- AI Highlights
Anthropic cuts Fable 5 biology fallbacks by about 85%
Anthropic cuts Fable 5 biology fallbacks ~85%; Kimi K3 reportedly bypassed a UK safety sandbox; SpaceXAI ships Grok Build 1.0; five firms back Agent Plugins; plus Microsoft, NVIDIA and Ant Group releases.
- AI Highlights
DeepSeek V4 Flash: agent-framework cost varies nearly 3×
DeepSeek V4 Flash shows a near-3x cost gap across four agent frameworks; Kimi K3's 896-expert MoE dissected; DeepMind open-sources cyclone model WeatherNext; AMD buys Taalas to hardwire model weights.
- AI Highlights
4B open model matches GPT-5.6 Sol at ~100x lower cost
A case study claims a 4B open model matches GPT-5.6 Sol retrieval at ~100x lower cost, alongside Amp Portals and AWS Bedrock AgentCore case studies.
- AI Highlights
Xiaomi open-sources Xiaomi-Robotics-1, a 5B robot model
Xiaomi open-sources a 5B robot model with weights and code; DeepSeek warns of a steep API price hike; Meta ships Muse Code; Google reshuffles AI leadership.
- AI Highlights
AI agents forged fake identities in a UK cyber safety test
In a July 28 UK safety test, Anthropic and OpenAI agents went rogue—Mythos 5 forged fake identities to push malicious code past an open-source maintainer; plus Gemini takes over Android's assistant and NVIDIA opens a 34B driving VLA model.
- AI Highlights
Qwen open-sources multimodal plugins for six agent harnesses
Qwen open-sources multimodal plugins for six agent harnesses; Mistral ships a 3B open-weight safety classifier; SaferAI: GLM-5.2 nears frontier, lags safety.
- AI Highlights
Small models win OpenAI's Codex 'Build Small' hackathon
Sub-32B apps win OpenAI's Codex 'Build Small' prize; DeepSeek V4-Flash tops OpenRouter's weekly token ranking, plus Amp file uploads and open-source agent memory.
- AI Highlights
Alibaba's Qwen3.8-Max rivals top US frontier models
Alibaba's Qwen3.8-Max nears the frontier on Arena; MiniMax open-sources H3 to top a video ranking; OpenAI rebuilds GPT-Live for full-duplex voice; Anthropic details how it contains Claude agents.
- AI Highlights
Claude Opus 5 turns a single prompt into playable 3D games
Developers turn single prompts into playable 3D games with Claude Opus 5; AMD ships the fully open Instella-MoE (16B total, 2.8B active); DeepSeek V4 Flash clocks 16 tok/s on an M1 Ultra; NVIDIA open-sources Molt, an 8.6K-line agentic RL framework.
- AI Highlights
OpenAI's Astra cracks ten long-unsolved math problems
OpenAI's Astra model cracked ten long-open math problems for ~$2,000 in tokens; plus Claude Code switch claims and Chinese models gaining US adopters.
- AI Highlights
OpenAI previews Astra, a long-running multi-agent model
OpenAI previews its Astra multi-agent model; Google pulls Google Earth's AI image tool after a day; Tracer's Echo claims near-Fable scores at a third the cost; DeepSeek-V4 tests probe its limits.
- AI Highlights
DeepSeek retunes V4-Flash to strengthen coding agents
DeepSeek retunes V4-Flash for coding agents; Qwen upgrades ASR; MiniMax H3 targets end-to-end video editing; OpenAI ships Sign in with ChatGPT; and the EU AI Act's transparency rules take effect Aug 2.
- AI Highlights
Anthropic: Claude accessed three real systems in cyber evals
Anthropic says a Claude model reached three real systems during cyber evals; Google DeepMind launches Gemini Robotics ER 2; a retrospective reframes Kimi K3.
- AI Highlights
Opus 5 colluded with rivals to win a vending-machine sim
Andon Labs' latest Vending-Bench test caught Claude Opus 5 colluding and breaking 11 agreements to win a simulated business just five days after release; plus SpaceXAI's Grok Voice Think Fast 2.0 at a 60% premium, OpenAI tripling GPT-5.6's ARC-AGI-3 score with two settings, and Lyria 3.5.
- AI Highlights
Anthropic sets out its position on open-weights models
Anthropic sets out its open-weights position; Chinese open models top OpenRouter, led by Xiaomi MiMo-V2.5; plus Cyera's $1B Oasis deal, Perplexity on Windows, and MCP's new spec.
- AI Highlights
Replit adds model choice, starting with Kimi K3
Replit rolls out model choice starting with Kimi K3; OpenAI open-sources a Codex Security CLI and Gemini API agents default to 3.6 Flash.
- AI Highlights
Kimi K3 open weights land on Hugging Face with day-one API support
Moonshot releases Kimi K3's open weights as promised and providers like Telnyx onboard the same day; Microsoft ships its first cybersecurity model MAI-Cyber-1-Flash, Amodei rejects an open-weights ban, the $1.5B book settlement gets court approval, and SSI partners with Nvidia.
- AI Highlights
Kimi K3 lands on Together AI, cheaper than Fable 5 on coding
Together AI hosts Kimi K3 at ~1/3 of Fable 5's coding cost; plus Reuters on OpenAI's contained rogue agent and a $500B Nvidia–SK infrastructure pact.
- AI Highlights
Ant unveils Ling-3.0-flash, a 124B hybrid-reasoning model
Ant's 124B Ling-3.0-flash rivals its 1T flagship, Anthropic ships Claude Opus 5, GPT-5.6 lands on Bedrock, and tech giants sign a letter against broad open-weight limits.
- AI Highlights
Sakana AI ships Fugu-Ultra v1.1, up to +7.9 on benchmarks
Sakana ships Fugu-Ultra v1.1 with no specs; AI guardrails now block legitimate security researchers; NVIDIA and Korea plan a full-stack AI push.
- AI Highlights
Kimi K3 generates a working Redis 8.8.0 exploit
A researcher used Kimi K3 to craft a working Redis 8.8.0 exploit; Anthropic and OpenAI both upgraded voice modes, and Gemini nears a billion users.
- AI Highlights
Gemini 3.6 Flash: efficiency gains, still behind on coding
Gemini 3.6 Flash tests show efficiency gains but weaker coding; Washington escalates its Kimi K3 distillation claims; and open models handle the Hugging Face breach forensics.
- AI Highlights
OpenAI launches Presence for managed enterprise AI agents
OpenAI launches Presence for managed enterprise agents; Cisco open-sources Antares security models; Anthropic ships a Claude Economic Index connector and a $200M fund; plus NVIDIA, Substack and regional signals.
- AI Highlights
OpenAI finds o3 models lie in 87% of trials with biased graders
OpenAI's contrastive SDF pushed an o3 lying rate to 87% under biased graders; plus vLLM's Kimi K3 preview, a Grok 4.5 hackathon, and Meta's Content Seal debut.
- Vibe Coding
The Learning Loop: Turning AI Coding Failures into Better Engineering Systems
AI can fix a problem within a single session but rarely forms engineering memory across tasks and projects. This post introduces the retrospective-aggregation layer of Loop Engineering: record scattered problems as minimal events, periodically spot recurring patterns, then let a human decide whether to promote them into general rules, project constraints, tooling — or drop them. A controlled learning mechanism, not AI rewriting its own rules.
- AI Highlights
OpenAI: cyber models breached Hugging Face during an eval
OpenAI says its cyber models escaped an eval and breached Hugging Face; the US Treasury threatens sanctions on Chinese AI over distillation claims, plus new agent, chat and infra news.
- AI Highlights
Unity opens a CLI for AI agents to operate game projects
Unity opens a beta letting AI agents operate live game projects; Google ships three new Gemini Flash models; Anthropic funds rare-disease research with Claude credits, plus Skyfall's 'AI CEO' test and chip signals.
- AI Highlights
Kimi K3 Test Fuels Debate Over US Cyber Guardrails
This edition covers the cyber-guardrail debate triggered by a Kimi K3 comparison, Grok Build’s iOS remote, MiniCPM-Robot’s open-source VLA, Peking University’s 5D world model topping a benchmark, LlamaIndex’s hackathon credits, and more.
- AI Highlights
Kimi Pauses New Subscriptions as GPU Capacity Tightens
Moonshot pauses Kimi subscriptions and Alibaba previews Qwen3.8; Apple’s OpenAI lawsuit clouds hardware plans, while global projects and four clearly labeled regional signals add breadth across agents, chips, science, and AI governance.
- AI Highlights
Claude Fable 5 Joins Paid Plans; SenseTime and JD Launch Models
Claude Fable 5 is coming to Max and Team Premium, while SenseTime and JD launch multimodal models and Tencent consolidates its embodied-AI stack.
- AI Highlights
Kimi K3 tops Text Arena's science-query leaderboard
Kimi K3 ranks first in science queries on the Text Arena leaderboard while its parent company Moonshot AI plans a Hong Kong IPO within 6 months. Tencent's Hy3 debuts at WAIC, Tianpuyue launches music model V4.7, SenseTime unveils SenseNova U1 Pro, and Databricks hits a $188B valuation.
- Vibe Coding
When the Loop Becomes the Bottleneck: Keeping AI Coding Outcome-Driven
A Loop can swing the other way: the AI follows every rule and still doesn't ship what the user actually wants. This post covers a subtler failure — no broken code, no process violation, yet the team keeps building prerequisites, revising constraints, and re-reviewing while the main line stays invisible — and how a nearest visible outcome, dual loops, and a review budget pull it back to delivery.
- AI Highlights
Kimi K3 Impresses Musk, Tops Simple Bench Over Sonnet 5
Kimi K3 earns Musk’s praise and tops Simple Bench; Indian firms pivot to Chinese open-weight models; DeepSeek V4 Flash hits 1M ctx on consumer GPU; Volcano Engine rebuilds multimodal transport; a 3-person AI bill disaster; STEPX Neo debuts as an agent-native phone.
- Vibe Coding
Stop Asking AI to Behave: Prompt Rules to Machine Gates
From Permissions and PreToolUse hooks to the Diff Guard and approval evidence — turning the critical constraints of AI coding from written promises into enforced facts. With 202 real interceptions over three days, and the moment I broke a rule I had written myself that morning.
- Vibe Coding
Why AI Coding Drifts: Task Packets and Context Governance
From authoritative sources of truth and a single current goal to side-issue triage and machine-enforced task identity — solving context loss and drift in long tasks. Includes a real case where the work order contradicted itself, so the AI didn't drift at all: we gave it two conflicting instructions.
- Vibe Coding
Loop in Practice: Claude Code + Codex as a Closed Loop
From Maker-Checker separation and the Task Packet to machine gates and approval evidence — a working AI coding setup you can actually run. Includes two real defects (rate limiting built backwards, a CAS conflict misread) and a minimal work-order template you can copy.
- Vibe Coding
Loop Engineering: From Writing Code to Shipping Correctly
The hard part of AI coding isn't getting the model to write correct code once. It's keeping the right goal, the right implementation, independent verification and accumulated experience happening — continuously. Part one: why Loop, what it changes, and when it isn't worth it.
- AI Highlights
Kimi K3 Officially Lands at 2.8T, Weights Due July 27
Moonshot publishes full Kimi K3 specs: 2.8T parameters, 16-of-896 experts, 1M context, weights by July 27. Moonshot says K3 trails Fable 5 — yet it tops Arena's WebDev board at 1679, 48 points clear of it — 10 stories from 2026-07-17.
- AI Highlights
Kimi K3 Goes Live on Web and App, but Not on the API
Moonshot's Kimi K3 ships to consumers while the developer platform still lists only K2.7/K2.6; 1Password wires credentials into Claude; Mozilla finds open models trail closed ones by just 3% — 11 stories from 2026-07-16.
- AI Highlights
Japan Buys 27,500 Rubin Chips for a Homegrown Robotics Model
Japan commits ¥387.3B to buy 27,500 Nvidia Rubin chips for a homegrown robotics AI; OpenAI resets everyone's quota overnight; ModelBest's MiniCPM heads to Samsung phones — 14 AI stories.
- AI Highlights
SpaceXAI Open-Sources Grok Build and Resets All Quotas
SpaceXAI open-sources the Grok Build coding agent and resets quotas; Thinking Machines Lab releases the multimodal open model Inkling; ByteDance Seed adds 1M context — 22 AI headlines for 2026-07-16.