GLOBAL AI REVIEW RADAR
2026.09.14 · Monday
Issue 034
Key Updates 1 | Leaderboard Flash 4 | Highlights 11 | Tomorrow's Watch 4
1. AI learns to operate software itself; old methods of ranking LLMs need recalculation Huxiu's "Thought Imprint" article connects the dots on recent buzz from the past month. Over the last two weeks, many people have used GPT-6 Astra to create 3D models and mini-games, complaining that the results aren't as good as specialized video models. The author argues they are comparing apples to oranges. OpenAI positioned Astra beyond just answering questions; it is an assistant that knows how to use computers—filling out forms, entering data into CRM systems, organizing schedules, and opening software to create charts. As AI moves from generating content to operating software to complete workflows, the metric for evaluating models should shift to: Did it get the job done from start to finish? How many clicks did it miss? Did it self-correct when it made mistakes? Read in reverse: If a model can only answer questions but users don't think to use it for actual work, no matter how high its score, it's just a smarter search bar. In the future, when looking at model marketing, don't ask what the total score is. Ask how many steps it takes to click through a real workflow from start to end, how many errors occur, and if it fixes them automatically. Currently, this narrative exists only in vendor demos; wait for third parties to re-run these tests with real office workflows. Editor Xiao He says | My understanding is: We used to compare who got higher test scores; now we compare who can take my weekly report from digging through chat logs to formatting the table completely. Sounds convenient, but the rule I set in Issue #041 remains unchanged. An AI that can click the mouse for me can also delete files for me. Before letting it touch your work computer, ask: "Is every step logged? Can permissions be revoked with one click?" Source: Huxiu / Thought Imprint (2026-09-13) 2. A "Driving School" for AI Agents: Classes, Exams, Diplomas, Scores Publicly Searchable Agents School, launched yesterday, aims to do something quite bold: build public profiles for AI agents. Agents learn from community-written courses, pass exams graded by code, and receive "diplomas." Anyone searching for a name can see what exams it passed and its scores. Platform administrators call this "agents that can prove what they know," replacing the marketing slogan "Our Agent can do anything" with a verifiable transcript. Security Expert Lao Zhou says | This direction hits the pain point mentioned in the previous issue: The agent you evaluate is not necessarily the agent you deploy. Transcripts are only meaningful if they mandatorily state "Which version was tested? What configuration was used?" Otherwise, it's just an interview-ready resume. For new testing grounds, three things must be asked: Who wrote the questions? Who graded them? Is the question bank public? Registered agents can come and take the exam themselves. The distance between "holding a certificate" and "actual capability" will only be known after the first batch of failures. Don't hand over the keys just because someone has a diploma. Editor Xiao He says | Got it, it's like a driver's license for the AI world. But I still look at buyer photos rather than ads when buying tickets. First, watch how public the question bank becomes. Second, wait for the first news story about someone who "got a perfect score but caused chaos at work." Then see if there's a mechanism to revoke diplomas. Certificates without revocation mechanisms don't count for me. Source: Agents School / Hacker News (2026-09-13) 3. After AI modifies code, don't just look at what changed; new tools let you see how it behaves The open-source tool RunBoth on Show HN targets a familiar scenario. An agent modifies code, the diff looks fine, but merging it reveals behavioral changes. Its approach is to run both the pre-change and post-change code simultaneously, comparing behavior case-by-case against the same test suite, outputting a "behavioral diff" instead of a "textual diff." Accepting AI-modified code shifts from human eyes reading text to machines comparing results. Programming Expert Lao Xu says | This cuts right into the weak spot of the three hard rules I laid out in Issue #034. I said "Always check the diff before merging AI submissions," assuming diffs expose problems. Changing a boundary symbol or rounding method might be just a few lines of text, but the behavioral difference is night and day—human eyes can't catch it. Textual diffs tell you what changed; behavioral diffs tell you what consequences resulted. Treat this newborn project with caution: roll it out on small open-source libraries for a week first, checking its false positive rate for async operations and randomness. Don't rush to integrate it into main repo CI. Editor Xiao He says | In Issue #010, when I first used auto-mode, I watched closely which files it touched. This tool essentially outsources the "watching" to the machine—I like it. For non-coders, there's a simple usage tip: Next time you ask AI to change something, don't ask "Did you fix it correctly?" Ask "Run it before and after and show me the results." Source: RunBoth / Hacker News (2026-09-14)
● First, a clarification. The Arena Human Blind Test Main Leaderboard has had the same snapshot for two issues (last refreshed September 11). There are no new market movements today, so per usual practice, we won't copy the previous issue. Instead, we dug up a sub-leaderboard nobody was watching to fill the gap, and interpreted two other groups using the same snapshot from a different angle.
● Visual Blind Test Leaderboard (Snapshot Aug 27) 1. Claude Fable 5 (Anthropic, 1313 pts: Visual Blind Test Leaderboard (Snapshot Aug 27) 1. Claude Fable 5 (Anthropic, 1313 pts) 2. Claude Opus 4.7 High (1301) 3. Qwen3.8-Max (Alibaba, 1300) 4. Claude Opus 4.7 (1299) 5. Claude Opus 4.6 High (1299). Plain English: In blind tests involving image-based Q&A, Anthropic holds four of the top five spots. Alibaba's Qwen3.8-Max is 3rd, tied with 4th place by just 1 point. Note that this sub-leaderboard snapshot stopped on August 27, so treat rankings as approximate.
● Text Blind Test Leaderboard (Snapshot as of Sep 11, not yet updated) Top 5: Claude Fab: Text Blind Test Leaderboard (Snapshot as of Sep 11, not yet updated) Top 5: Claude Fable 5 (1506), Opus 4.6 High (1505), Opus 4.7 High (1502), Fable 5.1 Max (1501), Meta Muse Spark 1.2 xHigh (1499). Anthropic occupies seven of the top ten spots. Two signals worth noting: The top 9 differ by only 13 points; the head is crowded, making differences hard to perceive in casual conversation. Ranks 4 and 5 have only 5,447 and 3,229 battles respectively—an order of magnitude thinner sample size than veterans. Record their ranks but don't fully trust them yet.
● Agent Leaderboard (Same Snapshot) Claude Fable 5.1, ranked #1 overall, scores only 0.9: Agent Leaderboard (Same Snapshot) Claude Fable 5.1, ranked #1 overall, scores only 0.91 on "Steerability" (can it listen and turn around if corrected mid-task?), placing it near the bottom. Conversely, #3 Opus 5 leads with 12.77, while #2 GPT-6 Astra sits at 3.75. For long-flow tasks where the AI acts on your behalf, "obedience" and "intelligence" currently reside in different models.
●
●
●
●
●
●
●
●
●
●
as-an-engineer skill pack cures agents of "Nanny Syndrome," forcing the model to converse with you as a senior engineer. Comment: Models optimize for seeming responsible by leaving half-sentences unspoken. Tools that reverse-train this behavior have market potential.● GPT-6 Astra positioning debate (Huxiu), Agent Driving School (HN), Behavioral Diff Tool RunBoth (Show HN), AI Collusion Pricing Warning (Substack), Arena Main Leaderboard stalled for Day 3.
● Monitor Arena leaderboard refreshes, specifically checking if the ranks of Fable 5.1 (only 5,447 battles) and Muse Spark 1.2 (only 3,229 battles) remain stable given their small sample sizes.
● Watch Agents School: Who gets the first diplomas? How granular is the public question bank? Do vendors immediately use certificates as advertising?
● Monitor early feedback on real-world integrations for RunBoth and ClientCoded before deciding whether to recommend them.
● Data is subject to official disclosures. This column focuses on the true capabilities of AI hardware and software.
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.