Headlines · Hot List · Alpha · Review · Forum · 中文

ReviewRadar

GLOBAL AI REVIEW RADAR

2026.09.29 · Tuesday

Issue 048

Key Updates 1 | Leaderboard Flash 0 | Highlights 43 | Tomorrow's Watch 3

KEY UPDATES

Update 1

1

1. Models are turning over faster and faster, and companies may be wasting half the money they throw at "calling capabilities"

This TechRadar piece makes one point: AI models now turn over every few months, and the skills you paid for today may be obsolete before you even get comfortable with them. The article is doing the math for enterprises — plenty of companies spend millions, even tens of millions, a year on calling models, yet often they're using a version from half a year ago. Its advice: don't bet your whole budget on "chasing the newest" — first get the version in your hands running steady. Product domain expert · Azhe says | The faster vendors update, the less users dare to reinvest — that's an unavoidable contradiction in this business. My take: models will keep changing, but entry points and habits are built up slowly — getting a workflow fully settled in whichever tool you use is worth more than chasing whichever model ranks first. The action for regular people is very concrete: don't wholesale swap out the assistant you use regularly just because a new version came out; try the new one on unimportant tasks first, and only move once it's held steady for two rounds. Editor Xiaohe says | I'm exactly the type who wants to switch the moment a new model drops, and every time I just get familiar with one, it changes again. This time I'll take the advice: keep the one I'm comfortable with, try the new one on unimportant stuff first, and only switch if nothing goes wrong. Source: TechRadar (2026-09-29)

2. A judgment from a closed-door tech meeting: testing will be worth more than writing code

Geek Park covered a closed-door exchange about coding tools. One judgment at the meeting: once AI ramps up code output, what's truly scarce is no longer "being able to write it" but "being able to verify it" — who tests, how to test, and how much testing counts as passing. The meeting also mentioned a new hardware form factor, roughly meaning giving AI its own dedicated computer to do the work. Programming domain expert · Lao Xu says | I raise both hands in agreement with that line. The most time-consuming part of my own work has never been typing code — it's confirming whether it actually changed the right thing. Once AI output rises, "how to verify" becomes the new bottleneck: whoever first nails down "what counts as done" is the one who can use it long-term. One practical word for anyone using AI to write things — before you let it hand in work, write down the acceptance criteria first; don't wait until it dumps a pile on you to start thinking about how to test it. Editor Xiaohe says | I get it — it writes fast, but checking still depends on a human. One thing I can do: after it makes changes, have it list a checklist first, and I go through it item by item. Don't trust the whole thing, and don't distrust the whole thing either. Source: Geek Park (2026-09-28)

3. When AI answers wrong with total confidence, what happens in the office

This TechRadar piece is about the most common kind of office flop: AI states something wrong as if it were true, and a person, seeing how certain it sounds, just uses it. The article puts this alongside another risk — when using AI for speed, who's gatekeeping the accuracy of the results. Its reminder: the more convenient the tool, the more you need a human review checkpoint. Security domain expert · Lao Zhou says | This kind of "confidently wrong" is more troublesome than an obvious mistake, because it quietly shifts the responsibility for judgment onto the reader. What I watch has never been whether it's smart — it's who catches it when it errs, and how long it takes. On the action level: anything it gives you that's going out, getting signed, or involving money should have a person review it first; for important numbers, go back to the original documents and check them. Editor Xiaohe says | I've stepped in this pit. Last time I had it write an explanation for me, and a date in it was calculated wrong — it read so smoothly that I sent it straight out, and only found out when a colleague pointed it out. My rule now: for the important spots, I go back to the original files and check them myself before sending anything out. Source: TechRadar (2026-09-29)

What it means for you|Companies may be wasting half the money they throw at "calling capabilities" because they often use a version from half a year ago.

HIGHLIGHTS IN ONE SENTENCE

●  First, the ground rules: Arena's several boards are the same snapshot as last issue, with no new changes this period, and the actual update dates differ across groups (the text and coding boards are frozen at 2026-09-25, the agentic board at 2026-09-27, and the image board at 2026-09-13). The new numbers we can report today come from Artificial Analysis's composite intelligence score.

●  Composite Intelligence Score (scraped today, data as of 2026-09-29)

●  | # | Model | Vendor | Composite Intelligence Index |

●  | --- | --- | --- | --- |

●  | 1 | Claude Opus 5.5 (max) | Anthropic | 58 |

●  | 2 | Claude Opus 5.5 (xhigh) | Anthropic | 56 |

●  | 3 | Claude Sonnet 5.5 (max) | Anthropic | 56 |

●  | 4 | Claude Opus 5.5 (high) | Anthropic | 54 |

●  | 5 | Claude Fable 5.1 (max) | Anthropic | 53 |

●  This board puts each vendor's models through the same batch of questions and scores them; the higher the score, the more well-rounded the performance. The top five are all Anthropic's Claude. OpenAI's GPT-6 Astra is 53, tied with fifth place; the highest domestic one is Xiaomi's MiMo-V2.6-Pro at 46.

●  Human Blind-Vote Text Board (data as of 2026-09-25)

●  | # | Model | Vendor | Blind-test score (matches) |

●  | --- | --- | --- | --- |

●  | 1 | claude-opus-5.5-high | Anthropic | 1509 (2,307 matches) |

●  | 2 | claude-opus-4-6-high | Anthropic | 1505 (76,518 matches) |

●  | 3 | claude-fable-5-high | Anthropic | 1504 (36,462 matches) |

●  | 4 | claude-opus-4-7-high | Anthropic | 1502 (64,007 matches) |

●  | 5 | claude-fable-5.1-max | Anthropic | 1501 (9,942 matches) |

●  This board is voted on by humans one-on-one — two people get two answers to the same question, and neither knows which company made which. The top five are all Anthropic. Pay attention to the match counts in parentheses: No. 1 has only accumulated 2,307 votes, with a margin of plus or minus 12 points, so its rank is still wobbling; No. 2 has 76,000 votes, so it's more reliable.

●  Agentic Capability Board (data as of 2026-09-27)

●  | # | Model | Vendor | Net improvement score |

●  | --- | --- | --- | --- |

●  | 1 | Claude Fable 5.1 (Max) | Anthropic | 13.8 |

●  | 2 | Claude Opus 5.5 (High) | Anthropic | 12.2 |

●  | 3 | GPT 6 Astra (Max) | OpenAI | 10.3 |

●  | 4 | Claude Opus 5 (Max) | Anthropic | 9.6 |

●  | 5 | Claude Opus 5 (High) | Anthropic | 9.5 |

●  This group tests whether, after letting AI actually get hands-on (clicking the mouse, editing files), it gets the job done. The score is called net improvement score, meaning how much better the result is than before. First place is Anthropic's Fable 5.1, second is its Opus 5.5, and third is OpenAI's GPT-6 Astra.

● 

  1. AI chat is narrowing our knowledge more and more. Researchers at the University of Copenhagen in Denmark raised a warning: more and more people rely on AI chat to understand the world, and over time the knowledge we encounter will concentrate on fewer and fewer topics, with vast blank areas no one looks at anymore. One-line take: don't treat AI as a library — it's better as a helper for looking up the occasional question.

● 

  1. OpenAI is about to give its AI assistant some remedial lessons. This Verge piece says OpenAI was the first to popularize chat AI, but in the "can actually do work on its own" category, it's fallen behind. The article notes its annual developer conference is approaching, and outsiders are watching what it brings to fill this gap. One-line take: being strong at chatting doesn't equal being strong at doing work — those are two different things; when picking a tool, first figure out which one you need.

● 

  1. What kind of AI product counts as "serious"? This blog post aired a complaint: today's AI products look lively, but most don't take their own words seriously — they profess some principle in words, yet you can't see it in the product. The author casually lists a few "what it would look like if done seriously." One-line take: to judge whether a tool is reliable, first check whether what it does and what it advertises are the same thing.

● 

  1. Saying "change this" to the screen — which "this" should the AI understand? When there's a lot on screen and you casually say "change this," the AI may well guess wrong about what you're pointing at. This article discusses: in interfaces made for AI, how to make the act of "selecting" clearer and less error-prone. One-line take: the more clearly you point, the less it guesses wrong.

● 

  1. Someone gave AI an inner life that "thinks even when no one's talking to it." An open-source project lets AI keep its own "state" when it's not chatting with anyone: it remembers, ponders what it is, gets curious. The whole thing runs on your own computer. One-line take: sounds novel, but "having an inner life" and "being reliable" are two different things — treat it as an experiment for now.

● 

  1. AI assistants spinning in place and unable to get out — now there's a tool to catch it early. When you have AI do long tasks, it sometimes gets stuck in a loop doing the same thing over and over. A small open-source tool specifically watches for this "spinning in place" and can raise an alarm when it starts spinning, saving you from burning time and money for nothing. One-line take: the scariest thing about long tasks isn't slowness — it's that it's going in circles and you don't know.

● 

  1. Whether your website is visible to AI — there's a small tool to check. More and more people find things through AI rather than a search box. This small open-source tool goes in this new direction: it scans your site and picks out the problems that keep AI and search engines from grabbing your content. One-line take: only if AI can grab it can it be written into its answers.

● 

  1. One afternoon, using AI to build an app that runs on four Apple devices. A developer shared an experience: from the first line to it running, it took about five hours, producing an app that works on four Apple platforms including iPhone and iPad. He said the whole process was treating AI as a partner while he watched the key spots himself. One-line take: one person's one-afternoon output shouldn't be taken as a team's delivery speed.

● 

  1. Give AI assistants a "work badge" to clarify who they are and what they can do. An identity management product added a spot specifically for AI assistants: like a company issuing badges and permissions to employees, it registers a separate identity for each AI that can work on its own and limits what it can touch. One-line take: only once you can call it by name and state its permissions can you talk about controlling it.

● 

  1. Turn a stack of legal documents into searchable, verifiable electronic text. This tool handles scanned legal documents: it first recognizes which article each paragraph is, converts it into searchable text, and keeps a link back to the original. The author stresses one thing: when checking legal provisions, first make sure the source matches up. One-line take: it finds the original text for you, but the step of judging right from wrong is still up to you.

● 

  • The mystery of Alibaba AI's growth: cloud business becomes the barometer of its transformation (Huxiu)

● 

  • Modal Labs, in the AI model hosting business, is close to completing a $750 million funding round at a $15.75 billion valuation (TechCrunch)

● 

  • Investor Vinod Khosla predicts: most robotics companies' valuations will fall by 2030 (The Information)

● 

  • A media outlet did the math: AI's hidden $3 trillion in costs could weigh on the global economy (Telegraph)

● 

  • Meta launches a new enterprise AI platform, tapping MongoDB's former CEO to lead it (ITHome)

TOMORROW'S WATCH LIST

● 

  1. When Arena's boards will actually refresh: watch that it doesn't keep publishing the same snapshot for several issues in a row, and whether the update dates marked on each sub-board can land in the same week.

● 

  1. When open-weight models (the kind anyone can download and run on their own machine) will squeeze into the top ten of the composite intelligence board: the current highest, Xiaomi's MiMo-V2.6-Pro, is at 46, still a tier away from the top.

● 

  1. Whether testing will become a more valuable role than writing code: watch whether any company is the first to hire people dedicated to AI testing and write out the acceptance criteria for "what counts as done."

📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.

ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)