ReviewRadar

GLOBAL AI REVIEW RADAR

2026.09.20 · Sunday

Issue 040

Key Updates 3 | Leaderboard Flash 6 | Highlights 4 | Tomorrow's Watch 4

KEY UPDATES

Update 1Research institutions have established "AI causing incidents in the physical world" as an independent subject

Research institutions have established "AI causing incidents in the physical world" as an independent subject. Micro1 released a report on September 18 stating that once models start driving browsers, calling APIs, and operating devices, the cost of errors shifts from "misreading a sentence" to "doing the wrong thing." Refunding incorrectly, placing wrong orders, or sending incorrect device commands all count. We agree with the direction, but need to think twice about the source; the report was published by a company selling AI engineer agents themselves. Who holds the ruler for the risk list matters more than how scary the title sounds. The action item for regular users is very specific: for any AI that can spend money, modify things, or toggle devices, keep permissions as minimal as possible, and retain manual confirmation for critical operations.

What it means for you|Keep permissions minimal and retain manual confirmation for critical operations when using AI that can spend money or modify devices.

Update 2Someone has written "Managing AI Employees" into three open-source spreadsheets

Someone has written "Managing AI Employees" into three open-source spreadsheets. Three governance templates (MIT license, free) just appeared on GitHub: one decision table clarifying "what AI decides on its own vs. what must stop and ask humans," a complete example, plus a Microsoft Agent 365 implementation template. It addresses the current real dilemma: AI is already on the job, but permission lists are still stuck in product managers' heads. A reminder: single-person repos, no adversarial testing—the tables might look smooth, but you need to see if they actually block models that truly overstep boundaries. Regular users can also ask their tools these three questions: What can it do for me right now? Which actions require my prior approval? Can I revoke access with one click?

What it means for you|Ask your tools what they can do now, which actions require prior approval, and if you can revoke access with one click.

Update 3For the first time, AI coding bills allow checking for "wasted retries." A new tool called AgentMeasure on Show HN directly reads local Codex/Claude Code logs, listing which steps kept stumbling in place, then converts this into a settlement based on subscription fees—data never leaves your machine

For the first time, AI coding bills allow checking for "wasted retries." A new tool called AgentMeasure on Show HN directly reads local Codex/Claude Code logs, listing which steps kept stumbling in place, then converts this into a settlement based on subscription fees—data never leaves your machine. Users of monthly-subscription coding agents can finally use the ledger to verify if it actually got the work done. It's a prototype born just one day ago; run it on your own project for a week first, don't rush to use it as reimbursement proof.

What it means for you|Users of monthly-subscription coding agents can finally use the ledger to verify if it actually got the work done.

LEADERBOARD FLASH

●  General Intelligence Leaderboard: Anthropic's Claude Fable 5.1 and OpenAI's GPT-6 Astra tie for the lead at 53 points. Same score, different prices: GPT-6 does the same work for less money.

●  Text Blind Test Leaderboard: Anthropic occupies seven of the top ten spots. Ranks 1 through 4 differ by only 6 points, essentially a tie; Rank 4 Meta has only 3,227 battles, not even a fraction of the leader's count.

●  Coding Blind Test Leaderboard: Top spot GPT-6 Astra scores 1800 but has the thinnest sample size on the board (2,281 battles, ±16 point fluctuation); Chinese contenders hold three or more seats, with Alibaba's Qwen3.8 occupying three chairs (Ranks 4, 6, 9), and Kimi K3 at Rank 5 clinging close behind.

●  Agent Practicality Leaderboard: Tests actual work. Kimi K3 has 108,000 matches behind it, the thickest sample on the board; Tool hallucination rate (the proportion of fabricating interfaces before calling them) is generally 0.37 for the top ten, but Claude Opus 4.8 achieves a unique 0.12, though its overall rank is dragged down to 6th by its "task completion rate."

●  Speed Leaderboard Changes Hands: Celeris-1 runs at approx. 1,567 words per second, more than three times faster than second place; the top open-weight model remains Zhipu GLM-5.3 (45 points).

●  Overall Summary: Intelligence, speed, cost-saving, open-source, and practical work—today, five different models each take first place in one category. The question "Which is strongest?" is outdated; "Which fits your specific task best?" yields the answer.

HIGHLIGHTS IN ONE SENTENCE

● 

  • The top three in Visual Blind Test differ by only 9 points; image recognition and table reading are effectively tied among leaders, so pick the cheaper one.

● 

  • Meta Muse Spark rolled out three versions across multiple leaderboards simultaneously, betting heavily, but sample sizes are generally thin.

● 

  • The most important number in leaderboard news isn't the rank, but the battle count in parentheses.

● 

  • A Spanish-speaking vendor launched a self-promotion page for "AI Boundary Protection," with no customer cases or verifiable data; treat it as marketing material.

TOMORROW'S WATCH LIST

●  Arena snapshot stopped on Sept 19 (page claims earlier updates); watch for official refreshes, specifically if the Coding Leaderboard champion's match count exceeds 5,000.

●  The growth rate of Meta Muse Spark series; ranks aren't stable until they accumulate enough samples (tens of thousands).

●  Whether Micro1's physical risk report gets third-party re-testing or named rebuttals; until it withstands peer scrutiny, it's just a proposal.

●  Full leaderboard tables and expert commentary are in the daily issue page. Feel free to chat about how much wasted money you've spent letting AI do your work.

📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.

ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)