ReviewRadar

GLOBAL AI REVIEW RADAR

2026.09.24 · Thursday

Issue 044

Key Updates 0 | Leaderboard Flash 4 | Highlights 10 | Tomorrow's Watch 3

LEADERBOARD FLASH

● 

  • Third-party intelligence index (Artificial Analysis, real-time scrape on Sept 24): the top three are all different tiers of Claude Opus 5.5, at 58 / 56 / 54 points; fourth and fifth are Claude Fable 5.1 and GPT-6 Astra, both at 53. The strongest open-source one is Xiaomi MiMo-V2.6-Pro, at 46.

● 

  • Coding blind test leaderboard (Arena, data as of Sept 11): top is gpt-6-astra-max at 1800 points, but it only played 2,281 matches with a 16-point swing; third place has over twelve thousand matches with only a 7-point swing.

● 

  • Getting-things-done leaderboard (Arena, data as of Sept 15): top two are Claude Fable 5.1 (Max) and GPT 6 Astra (Max); it compares whether it gets the job done, confirms it's done, and can rescue itself after a mistake.

●  All three Arena leaderboards haven't updated in over ten days, and none of the new models launched this week made it on. Don't use the rankings as a basis for model selection yet — look at the match counts in parentheses first.

HIGHLIGHTS IN ONE SENTENCE

●  OpenAI put out a fixed test paper for AI's mental-companion skills (MentalHealthBench)

●  It tests whether AI answers reliably on topics like mental health: does it go along with dangerous thoughts, does it tell people to see a doctor when it should. Old Zhou's take: mental health topics are the easiest place for things to go wrong — AI has no license, yet its tone is as confident as someone who does; the first version is usually the same shop writing the questions and grading them, so wait until the questions can be downloaded and others can rerun it before treating it as a basis. Xiao He adds: that "don't worry" line — take it as a reminder, not a conclusion.

●  Nvidia released a real-time voice model that can tell "who's speaking" (Nemotron 3)

●  In multi-person conversations, it figures out who said which sentence while listening, not after the recording is done. The most common mistake in meeting transcription is attributing Zhang San's words to Li Si. A-Zhe reminds: in voice comparisons, the dividing line stopped being "can you hear it clearly" a long time ago — it's now "can it guess who should be speaking" accurately; official demos don't count, first try it on your own company's noisy meeting recordings.

●  Someone kept records for five months and listed 48 ways an AI assistant can go wrong

●  An author had an AI assistant run their shop for them, and over five months logged every mistake. The worst one: the assistant said at 2 a.m. that tickets couldn't be sold anymore — and it had stayed silent about this for weeks. Old Xu says: mistakes aren't the scary part — the scary part is it keeps going without a peep; the most practical thing is to use this list to fix your own acceptance checks.

● 

  • Issue a verifiable credential for every step AI takes, $0.10 per verification (rubric-attest): the credential can prove that step wasn't altered, but not that the step was done right.

● 

  • Open-source robot test-paper site VSArena: give it a 128×128 camera frame plus one sentence, and where the blocks are depends entirely on eyesight.

● 

  • Have AI write Connect Four on the spot; in the second round of the coding contest it handed in 38 minutes late: the problem and the finished product are both public, more trustworthy than a vendor demo reel, but it's only one round.

●  The rest of the picks and the full commentary on the three highlights are on the review issue page. The preview will update along with tonight's official release.

TOMORROW'S WATCH LIST

● 

  1. After Arena's next refresh, will the newly launched Claude Opus 5.5, GPT-6 Sol, and Luna reshuffle the rankings.

● 

  1. For Nvidia's robot simulation tutorial, wait until someone follows it and produces a time-cost comparison before deciding whether it's worth messing with your GPU and drivers.

● 

  1. DeepSeek's new training method currently only has one report; wait until the official writeup says how much it saves and where it's better before treating it as a basis.

📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.

ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)