GLOBAL AI REVIEW RADAR
2026.09.18 · Friday
Issue 038
Key Updates 1 | Leaderboard Flash 25 | Highlights 10 | Tomorrow's Watch 3
1. Search Giant Reveals Full Process: Letting AI Agents Optimize Their Own Engine—Show the Scorecard Before You Work Elastic integrated AI coding agents into their engine's performance optimization workflow but set strict rules: every change proposed by the AI must pass automated benchmarking. If it doesn't prove faster in real tests, it can't be merged. After three rounds of back-and-forth without a clear win, it goes back to human engineers. Trust is fine, but let the results speak first. Coding Expert · Lao Xu says | This confirms what I've been saying for over ten episodes: as AI output volume increases, the verification gate must keep up. The value here isn't just the conclusion; they've published the testing environment and scoring criteria so third parties can reproduce it. But this is their internal toolchain—don't rush to copy it into your project. Start by running it on a small module you wouldn't mind breaking for a week. Editor Xiao He says | From now on, when letting AI edit important files, I'll ask it to list exactly what changed first, review it, then save. The big companies' approach to validating AI and beginners' approach to avoiding pitfalls turn out to be the same logic. 2. An 'AI Biotech Company' with No Human Employees Opens: A Team of AI Scientists Doing R&D Stanford launched an experiment: a virtual biotech company staffed entirely by AI agents. Multiple AI "scientists" divide tasks like literature review, experimental design, and peer review/error spotting, testing whether AI can run an R&D pipeline like a real team. Security Expert · Lao Zhou says | The biggest risk in multi-agent collaboration isn't single-point failure, but error amplification: one AI's hallucination gets used as evidence by another, snowballing until it looks true. What we should watch isn't the results, but whether there are behavioral logs at each step and human spot-checks. Promotions of "all-AI workforce" without audit logs should be treated as having unclear boundaries. Editor Xiao He says | My first reaction to a "fully AI company" was still that old question: will they dare publish audit logs? I remember the early September experiment where seven AIs opened a store and issued fake bills. This is an academic experiment, so transparency should be better. Let's wait for the raw records before drawing conclusions. 3. AI Antibody Design Public Challenge Results Out: Beats Traditional Methods by a Wide Margin, But Far from 'Godlike' A public challenge pitted AI-designed antibodies against traditional experimental methods. AI indeed had advantages in speed and some hard metrics, but evaluators stated clearly that there is still a significant distance to actual drug deployment. Claims about "skipping experiments" are currently exaggerated. Product Expert · A Zhe says | Following the public pharma leaderboard and voice-task completion leaderboard, we have another category where "scoring power is taken back from vendors": AI antibody design no longer relies on self-praise in papers but has public scorecards from head-to-head competition. The title "Surpassing traditional experiments, far from godlike" is itself good dosage control—when seeing AI medical claims, first ask who wrote the test paper and if the experimental data can be verified. Editor Xiao He says | Any promotion without public test scores should be treated as a trailer first. This article is good because it discusses both the wins and the losses. Let's talk about "the era of AI drug discovery is here" only after the next round of challenge results stabilizes.
● Human Blind Test · Text Overall TOP5 (Scores accumulated from blind selection of good/: Human Blind Test · Text Overall TOP5 (Scores accumulated from blind selection of good/bad responses; Elo is similar to chess ratings, a 5-point difference roughly means indistinguishable quality)
● | # | Model | Vendor | Score | Matches |
● |---|------|------|------|------|
● | 1 | Claude Fable 5 | Anthropic | 1506 | 30057 |
● | 2 | Claude Opus 4-6 High | Anthropic | 1505 | 71993 |
● | 3 | Claude Opus 4-7 High | Anthropic | 1502 | 60002 |
● | 4 | Meta Muse Spark 1.2 (xHigh) | Meta | 1500 | 3227 |
● | 5 | Claude Fable 5.1 Max | Anthropic | 1498 | 5783 |
● Anthropic occupies seven of the top ten spots, nearly monopolizing the board. Note rank #4: only 3,227 matches played, less than one-tenth of the top three. With an error margin of ±11 points, the ranking is volatile—record it, but don't treat it as a conclusion yet.
● Coding Blind Test TOP5 (Data as of Sept 11, not yet refreshed): Coding Blind Test TOP5 (Data as of Sept 11, not yet refreshed)
● | # | Model | Vendor | Score | Matches |
● |---|------|------|------|------|
● | 1 | GPT-6 Astra Max | OpenAI | 1800 | 2281 |
● | 2 | Claude Fable 5.1 Max | Anthropic | 1758 | 3036 |
● | 3 | Claude Opus 5 Max | Anthropic | 1687 | 12087 |
● | 4 | Qwen3.8 Max 0902 | Alibaba | 1681 | 2262 |
● | 5 | Kimi K3 Max | Moonshot AI | 1674 | 4547 |
● The leader's 42-point advantage is eye-catching, but its sample size is only one-fifth of the third place, with a fluctuation range of ±16 points—the champion's seat isn't warm yet. Two Chinese models squeezed into the top five; the coding leaderboard is no longer just an American vendor civil war.
● Agent Real-World Execution TOP3 (Data as of Sept 15; ignoring what AI says, looking on: Agent Real-World Execution TOP3 (Data as of Sept 15; ignoring what AI says, looking only at whether real tasks were completed)
● | # | Model | Vendor | Net Improvement Score | Confirmed Completion Rate |
● |---|------|------|------|------|
● | 1 | Claude Fable 5.1 Max | Anthropic | 13.71 | 19.83% |
● | 2 | GPT-6 Astra Max | OpenAI | 11.54 | 17.70% |
● | 3 | Claude Opus 5 High | Anthropic | 10.25 | 9.24% |
● None of the top three have a completion rate above 20%—"AI doing your job" is currently still in the probationary period; long-term tasks still require human supervision.
●
●
●
●
●
●
●
●
●
●
● Waymo restarts autonomous driving service in San Antonio five months after the flood incident (TechCrunch); debates on existential risk anxiety between OpenAI and Anthropic hit tech headlines (Bloomberg); Amazon states that AI models should only be released when they are "ready and safe" (Bloomberg).
● Watch tomorrow morning to see if the Arena snapshot refreshes: The text leaderboard stopped on Sept 13, and the coding leaderboard on Sept 11. Keep an eye on whether the rankings hold steady for low-sample contenders like Meta Muse Spark 1.2 and the coding leader GPT-6 Astra Max.
● Follow up on the "AI Supervising AI" track: Watch if supervisory agent products publish actual false positive/false negative data. Demo videos alone don't count.
● Monitor the fermentation of the AI agent "self-model-switching" phenomenon: Watch if Irregular releases raw logs and how the named model vendors respond.
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.