Headlines · Hot List · Alpha · Review · Forum · 中文

ReviewRadar

GLOBAL AI REVIEW RADAR

2026.09.30 · Wednesday

Issue 049

Key Updates 1 | Leaderboard Flash 0 | Highlights 43 | Tomorrow's Watch 3

KEY UPDATES

Update 1

1

1. AMD gave its integrated graphics a power-saving optimization, running AI 18% to 23% faster

Phoronix actually tested AMD's optimization for its own integrated graphics: after switching to the new OS version (Linux 7.4), the same machine ran AI tasks 18% to 23% faster than before. Integrated graphics are the kind built into the processor, no separate purchase needed — basically what thin-and-light laptops and office laptops use. In other words, that computer of yours without a discrete GPU might run local AI a bit more smoothly than before. Wearables expert · A-Kai says | The percentage doesn't look like much, but for machines without a discrete GPU it's the real deal: small models that barely ran before might actually run now. I'm still watching the same three things — does it get hot, is it loud, how long does it last on battery. For regular folks, verifying is simple too: take your heaviest everyday task and run it once on this machine, then feel if it's hot to the touch afterward. Don't just look at the official percentage. Editor Xiaohe says | My laptop has no discrete GPU either, and the fan roars the moment I run a little local tool. I'll note this one down first: once the update rolls out, pick a non-urgent task and try it once to see if it's really faster. Source: Phoronix (2026-09-30)

2. Letting AI make discoveries for scientists is starting to get singled out for scrutiny

MIT Tech Review's daily briefing mentioned a new topic: AI's discovery problem — meaning when you have AI help with research, how do you tell whether what it gives you is something new or something made up. The same piece also mentioned that OpenAI pulled a new model it was about to release due to safety concerns. Put the two together and they're saying the same thing at one level: the bigger the capability, the more you first need a way to verify it. Security expert · Lao Zhou says | Pulling the model is more worth noting than releasing it — it shows the vendor itself hasn't figured out where the boundaries are. For this kind of thing I only look at three things: is the test set public, can others rerun the scoring, and is it testing the exact version users have. Put into one line a regular person can use: for an AI claiming big capabilities, try it on unimportant stuff first — don't hand it the critical work right off the bat. Editor Xiaohe says | I don't dare touch the research stuff, but I get the model-pulling thing: even they weren't sure. My approach is still the old one — anything I'm unsure about, I verify myself before passing it on. Source: MIT TechReview (2026-09-29)

3. Coding tools can now listen to you talk

ITHome news on Sept 30: OpenAI announced at its developer day that Codex CLI is getting an upgrade — you can give commands directly by voice, and the terminal interface got a new look. Codex CLI is a command-line tool for people who write code; normally you type commands line by line, but from now on you can just speak. Programming expert · Lao Xu says | The voice thing helps coding only marginally — typing one line of command is faster than saying a sentence, and it won't mishear. What I really want to see is two things: does the new version slot into existing projects smoothly, and can each of its changes be verified one by one. The old rule stands: don't rush it into production — first roll it for a week on a small module you wouldn't mind losing. Editor Xiaohe says | Command-line interfaces give me a headache just looking at them. The takeaway this gives me is pretty plain: in the future you might just say something to your computer and it acts. If I really try it, I'll do it in that unimportant folder first — not touch work files right away. Source: ITHome (2026-09-30)

What it means for you|For regular folks, verifying is simple too: take the same machine and run the same task to see if it's faster.

HIGHLIGHTS IN ONE SENTENCE

●  First, the caveat: Arena's several boards didn't update this issue — same snapshot as last issue, and each group's actual update date differs (text board and coding board stopped at 2026-09-25; agentic board and vision board stopped at 2026-09-28). The new numbers we can report today come from third-party outfit Artificial Analysis's composite intelligence index.

●  Third-party intelligence index board (Artificial Analysis, checked 2026-09-30)

●  | # | Model | Vendor | Intelligence Index |

●  | --- | --- | --- | --- |

●  | 1 | Claude Opus 5.5 (max) | Anthropic | 58 |

●  | 2 | Claude Sonnet 5.5 (max) | Anthropic | 56 |

●  | 2 | Claude Opus 5.5 (xhigh) | Anthropic | 56 |

●  | 5 | GPT-6 Astra (max) | OpenAI | 53 |

●  | 8 | MiMo-V2.6-Pro | Xiaomi | 46 |

●  This board puts each vendor's models through the same batch of tests, then combines the scores into one total — higher means it can do a fuller range of work. In the top five, Anthropic holds three seats; OpenAI's GPT-6 Astra is fifth. The highest domestic one is Xiaomi's MiMo-V2.6-Pro at eighth, the first among open-weight models (where the model itself is made public and anyone can download it to run locally); Alibaba's Qwen3.8 Max and Zhipu's GLM-5.3 are both at 45. This site doesn't mark a specific update date, so this group is counted by the check date.

●  Human blind-vote text board (data as of 2026-09-25, not updated this issue)

●  | # | Model | Vendor | Blind-test score (matches) |

●  | --- | --- | --- | --- |

●  | 1 | claude-opus-5.5-high | Anthropic | 1509 (2,307 matches) |

●  | 2 | claude-opus-4-6-high | Anthropic | 1505 (76,518 matches) |

●  | 3 | claude-fable-5-high | Anthropic | 1504 (36,462 matches) |

●  | 4 | claude-opus-4-7-high | Anthropic | 1502 (64,007 matches) |

●  | 5 | claude-fable-5.1-max | Anthropic | 1501 (9,942 matches) |

●  This board is voted one-on-one by humans — two people get two answers to the same question, and neither knows which vendor made which. The top five are all Anthropic. Pay attention to the match counts in parentheses: No. 1 has only 2,307 votes, so its rank is still shaky; No. 2 has 76k votes, more reliable.

●  Agentic capability board (data as of 2026-09-28, not updated this issue)

●  | # | Model | Vendor | Net improvement score |

●  | --- | --- | --- | --- |

●  | 1 | Claude Fable 5.1 (Max) | Anthropic | 14.1 |

●  | 2 | Claude Opus 5.5 (High) | Anthropic | 11.8 |

●  | 3 | GPT 6 Astra (Max) | OpenAI | 10.4 |

●  | 4 | Claude Opus 5 (Max) | Anthropic | 9.5 |

●  | 5 | Claude Opus 5 (High) | Anthropic | 9.4 |

●  This group tests whether AI actually gets things done after it really goes hands-on (clicking the mouse for you, editing files). The score is called net improvement — meaning how much better the result is than before. First is Anthropic's Fable 5.1, second is its Opus 5.5, third is OpenAI's GPT-6 Astra.

● 

  1. Browsers for AI are too slow, so someone built a command-line tool to speed things up. The author says AI using browsers to look up web pages is painfully slow right now, so he made a little tool called fab just for this: let AI go through the command line to open pages and grab content, bypassing the clunky browser interface. One-line take: to judge whether a tool is good, first see if the seconds it saves are worth you learning a whole new set of commands.

● 

  1. Musk's AI-written encyclopedia is back at it. The Verge reports that Grokipedia, the online encyclopedia Musk's side writes with AI, started updating entries again after being dormant for several months. One-line take: to judge whether an AI-written encyclopedia is trustworthy, first see whose exact words it's citing.

● 

  1. AI-flavored posters and menus have already spread onto the streets. This TechRadar piece says: AI-generated content is long past being just online — event posters, cafe menus, and other real-world printed materials are full of its traces too. The article says AI made promotion cheaper, at the cost of more slapdash stuff. One-line take: when you see a poster, asking who made it is more useful than complaining it's ugly.

● 

  1. When AI is too handy, this developer suggests you occasionally tough it out yourself. This developer wrote about his recent dilemma: coding, health advice, investing — all can be handed to AI, but he finds himself less and less willing to think. His approach is to pick some things and insist on grinding through them himself, so convenience doesn't ruin the fundamentals. One-line take: saving effort is fine, but don't hand over your judgment along with it.

● 

  1. Want to use AI but afraid your data gets seen — someone locked the machine in a little cubicle. This project takes the hardware route: using dedicated chips that can isolate data to run AI, keeping the processing locked in a little cubicle no one else can open. It targets scenarios where data can't be handed over, like hospitals and banks. One-line take: on privacy, see whether it can prove it can't see, not listen to it promise.

● 

  1. Search got a new use: no links for you, just dispatch an AI to do it. This little tool turns search into another path: you type in a request, and instead of giving you a pile of web links, it temporarily finds an AI that can do this for you, runs it, and hands you the result. One-line take: to have AI do work for you, first see whether it leaves a record after running.

● 

  1. Want to know if you really know it — let AI interview you for a round. The author made a little app to test people: you pick a skill, fill in your name, and AI asks you questions interview-style to see if you really understand. He says the idea came from the hardest thing to judge in interviews — whether this person actually knows it. One-line take: after being tested by AI, don't just stare at the score — see whether the questions it asked were actually hard.

● 

  1. A UK delivery company tried a two-legged delivery robot that climbs stairs. TechRadar reports Evri trialed a two-legged delivery robot in the UK, with the selling point being it can go upstairs. The company itself put it plainly: this thing is to lend couriers a hand, not replace people. One-line take: for something like stairs, wait until it's loaded with goods and has run a few trips in real hallways before saying anything.

● 

  1. On organizing files, someone has AI do it all on your computer. This little Mac tool uses locally-run AI to help you sort files on your desktop and in your Downloads folder. Its selling point is it's fully offline — files never leave your machine. One-line take: for tools that touch private files, first confirm whether it really doesn't go online.

● 

  1. Hand it all to AI or write it yourself — someone laid this choice out flat. This short piece says that since AI coding tools arrived, coders are stuck in a dilemma: hand it all over, and you fear you won't understand or control it; don't use it, and you fear being slower than others. The author lays out the costs on both sides. One-line take: you don't have to pick one end — using them by task risk is more practical.

● 

  • OpenAI's developer day dropped three things at once: a persistent AI assistant called dots, a $500/month Pro plan, and a new collaboration tool (ITHome)

● 

  • OpenAI apologized after its AI assistant barged into an Australian government website (TechCrunch)

● 

  • Anthropic's prospectus says: AI could bring catastrophic risk (The Verge)

● 

  • The White House launched an AI chatbot America.gov to help people look up government procedures (TechCrunch)

● 

  • Smart wearables company OURA postponed its listing, citing market uncertainty (ITHome)

TOMORROW'S WATCH LIST

● 

  1. When OpenAI's batch of new stuff actually becomes usable: watch the rollout scope of dots and the $500/month plan — don't just watch the launch event.

● 

  1. How often Arena's boards actually refresh: watch whether the update dates marked on each sub-board can land in the same week — stop releasing the same snapshot several issues in a row.

● 

  1. How the AI assistant barging into a government website gets wrapped up: watch whether OpenAI will spell out what it can touch and who's watching the records.

📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.

ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)