ReviewRadar

GLOBAL AI REVIEW RADAR

2026.09.11 · Friday

Issue 031

Key Updates 1 | Leaderboard Flash 3 | Highlights 11 | Tomorrow's Watch 4

KEY UPDATES

Update 1**1

1. Who watches the watchers when AI goes wrong? A new "examiner" test specifically checks how many bad acts supervisory programs miss. Many companies have installed "internal monitors" on their AI—a piece of code that watches internal activities and raises an alarm if something looks off. ObserverBench, launched on September 10, doesn't do anything else but test these supervisors. It feeds a batch of cases where AI actually crossed the line into various monitoring methods to see who can correctly rank the severity of harm. The focus is on what real misdeeds were missed, not how many alarms were raised. A supervisor that misses big trouble but loves raising small alarms is more dangerous than one with low sensitivity. Currently, it's a login-free online demo that anyone can try. Security Expert Lao Zhou says | Last issue I said "the security gate itself needs to pass inspection first," and now someone has finally issued exam papers for the inspectors. Scoring based on missed harmful judgments forces monitoring methods to admit what they failed to catch, which is more honest than just looking at false positive rates. But the exam venue is newly built, and no one has vetted the depth of the question bank yet. I'm keeping tabs but not jumping in. Editor Xiao He says | At first, I didn't get the concept of internal monitors, but then it clicked: it's like installing cameras at home—you're afraid it won't record when something actually happens. The demo page requires no registration; I'll find some time to click through it for everyone and report back. 2. Legacy systems running for decades moved to new code by AI in two weeks; engineers shared every step of the crash-and-burn process. Many companies maintain a batch of antique code, such as old business systems written in Oracle Forms 6. Fewer people know how to modify them, but they can't be thrown away either. On September 10, Vaadin engineers wrote an experiment log where they had AI migrate an entire application to Java. What industry norms say takes 6 to 8 months was done in two weeks. The article isn't bragging; much of it records where the AI got stuck, how many rounds it took to get it right, and which parts still needed human fixes. Programming Expert Lao Xu says | Few dare to write about failure rounds in such detail; this honesty is worth more than the "two weeks" number. Let me throw some cold water: that's his personal speed, not the team average. There's still a whole suite of testing, regression, and gray-scale deployment between "it runs" and "we dare to launch." If you want to copy this trick, pick a small module you wouldn't mind losing and roll it out for a week first. Editor Xiao He says | What surprised me most is that he really dared to write out the AI's crash process. Our old office system quote from outsourcing was half a year; I forwarded this article to our tech lead. Repeating the old advice: before letting AI perform major surgery, back up first, and let it list changes for humans to review. 3. New Siri finally remembers who you are; Wired reviews new features after the September 10 event. After Apple's autumn event, Wired published a review of new features. The biggest change is that the new Siri begins to understand personal context. If you say "send that photo to Mom," it needs to know which photo you meant and which contact is "Mom." It also supports chaining multiple commands, and you can type to it in situations where speaking isn't convenient. This addresses the old problem of voice assistants failing to keep up with conversation flow. Product Expert Azhe says | In Issue #040, I reviewed this beta version: entry point does not equal usage. The only truly critical item in this feature list is "remembering you." Personal context is one of the few trump cards voice assistants can't easily copy because the iPhone is in your pocket. As usual, don't trust demo presentations; wait until the features actually land on devices, give it a week of real feedback, and then draw conclusions. Editor Xiao He says | Last time I said my inability to use it wasn't my fault; this time I'll stick to the old method and force myself to use it continuously for a week before reviewing. My only concern is that it will record my family and habits in its little notebook. Where is this notebook stored? Can it be used for training? I'll wait for the privacy policy to clarify before turning it on—don't rush to hit "Agree."

LEADERBOARD FLASH

●  Intelligence Index (Artificial Analysis, scraped Sept 11). Claude Fable 5.1 (max/xhigh: Intelligence Index (Artificial Analysis, scraped Sept 11). Claude Fable 5.1 (max/xhigh) and GPT-6 Astra (max/xhigh) tie at 53 points. With equal scores, look at the wallet: GPT-6 costs $2.31–$3.26 per task, while Claude costs $5.98–$7.63. The top three open-source models are all domestic: GLM-5.3 (45), Kimi K3 (44), and GLM-5.3-Flash (42).

●  Human Blind Test (Arena, data as of 2026-09-10, snapshot not updated). Anthropic holds: Human Blind Test (Arena, data as of 2026-09-10, snapshot not updated). Anthropic holds four of the top five spots in the text leaderboard, with Claude Fable 5 leading at 1507 points. A reminder: 3rd place Fable 5.1 (max) has only 2906 battles, making the sample size over twenty times thinner than the second place, so rankings may fluctuate. In the agent leaderboard, Fable 5.1 (Max) is first, winning on getting work done and self-repairing after command errors; GPT-6 Astra (Max) is second, but ironically has higher post-win reputation scores.

●  Speed & Cost (AA Today's List). Celeris-1 tops the speed chart at 1382 words per secon: Speed & Cost (AA Today's List). Celeris-1 tops the speed chart at 1382 words per second, more than 4x faster than the smartest model, Fable 5.1. However, fastest and smartest are not the same set of models; assign tasks accordingly. The cheapest single task drops to the $0.01 tier, sufficient for high-volume, simple scenarios, but don't skimp on this cost for difficult problems.

HIGHLIGHTS IN ONE SENTENCE

●  . Vendors stopped announcing this upfront; instead, switches are on by default upon registration, and sneaky terms changes are now tracked. Check this leaderboard before handing over data to AI.

●  . PromptSign applies software signing concepts to CLAUDE.md / AGENTS.md to prevent prompt injection. The direction is right; waiting for real interception cases.

●  . Doesn't care about nominal data center size, but rather how much paid-for compute you actually get. Next, watch to see if anyone dares to use it to publicly score cloud providers.

●  . Privacy-first approach gets bonus points; location accuracy is mediocre. If you're considering installing it, use this as a baseline draft and compare with established brands.

●  A developer experienced this firsthand: facts he remembered were confidently denied repeatedly by the AI until he checked logs to confirm the AI was lying. One crude solution: archive important conversations.

●  . First let AI generate an outline and start building; check docs only when stuck. AI provides the map; docs act as the referee.

●  . Are the questions accurate? Does grading distinguish real skill? No public data yet; waiting for candidate testimonials.

●  . Zero game dev experience, web racing game DRIVE NEXT—the barrier for "one person + AI" has dropped to your feet.

●  . c2anime's The Last Light claims agents completed everything from script to final cut. Choosing 34 seconds is smart; short duration is currently the capability boundary.

●  . Coletivo allows multiple agents to assign and hand off tasks via MCP. Agent collaboration is the next major testing ground; there aren't even decent benchmarks yet.

●  Anthropic dominates the top of Arena; AA leaders are tied but GPT-6 is half the price; top three open-source models are domestic; the 200-tool training data scoring site hits HN front page; new benchmark for testing AI monitors launches.

TOMORROW'S WATCH LIST

● 

  • Will Arena update the Sept 11 snapshot? Will newcomers with thin battle counts drop in ranking?

● 

  • Third-party call tests for DeepSeek V4.1 Flash after launching on the National Supercomputing Internet. Only bills from real work count.

● 

  • First wave of real word-of-mouth after the new Siri public beta rolls out.

●  Full version (with all source links) available at daily.physixfrontier.com/review/

📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.

ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)