ReviewRadar

GLOBAL AI REVIEW RADAR

2026.09.21 · Monday

Issue 041

Key Updates 1 | Leaderboard Flash 0 | Highlights 39 | Tomorrow's Watch 4

KEY UPDATES

Update 1

I

I. Two AIs Poison Each Other, Humans Judge: Prompt Wars Opens an Arena for "AI Scamming AI"

Adapted from the "Nuclear War" game played by programmers last century. You write a prompt to hand over to your AI warrior. It doesn't answer questions; it focuses on one thing: trying to brainwash the opposing AI's instructions, making the opponent forget its task and work for you instead. Every round of offense and defense is archived on the leaderboard. Which model is most easily fooled and which is most resistant to scams becomes clear at a glance. The reason to watch this is practical. As permissions granted to AI grow larger—reading emails, modifying files—the attack vector where "someone inserts a piece of text to manipulate it" is the most deadly. This public battle arena is essentially a public health check for the security baselines of various models. Security Expert Old Zhou says | The old problem of "inability to distinguish data from instructions" is demonstrated here in the most straightforward way, moving injection attacks from papers to leaderboards. But the usual three questions remain: Are the test cases public? Can the scoring be reproduced? Is it testing the version users actually have? Until the author clarifies whether the models in the arena are the versions you are currently using or if they have been secretly hardened, the rankings should only be taken as reference. Editor Xiao He says | I understand it as the AI world's prank contest, messing with each other to see who falls for it first. It looks fun, but when I think about how I actually delegate tasks to AI in daily life, I can't laugh anymore. The old rule still applies: Don't click or feed unknown files and links to humans or AI alike.

II. How Many Can Pass App Store Review with One-Sentence App Generation? StoreReady Puts AI App Builders on the Same Stage to Submit Answers

If you want to use AI to make an app and list it on the store, look at this case file first. It asks every AI app generation tool the same five questions: Can it pass Apple Store review? Where does it get stuck? How many times was it rejected? How long to fix it? Did it finally launch? Each entry must include evidence sources, presented as a flip-through comparison table, with a handwritten human comment closing out the last row. Everyone brags about AI helping you make apps, but Apple's review process is the real hurdle. Many people get stuck at "easy to generate, hard to list." Laying out the results of the same gatekeeper side-by-side is a rare horizontal real-world test for these tools. Programming Expert Old Xu says | The question setter is them, the answers come from public records of the tested tools, and each item has a verifiable source. This is better than launch event demos; clear documentation earns bonus points. However, the criteria for the five questions were set by the site owner. Passing review doesn't equal usability. Crash rates after listing, or whether AI can handle subsequent revisions, aren't tested in this table. If you want to switch tools, try a small project you wouldn't mind losing for a week first. Don't rush to hand over your main product. Editor Xiao He says | I checked those five rows for everyone; the gaps are significant. Some got rejected in the first round. The commoner's rustic method remains the same: Don't trust "one-click listing," look at others' actual rejection case files first. If you try it yourself, remember that developer accounts cost over 600 yuan/year, so factor that into your costs.

III. Text-to-Speech Tool Speechify Reviewed Again: Expensive, But It Truly Changed How We Read

TechRadar journalists conducted long-term real-world testing. Documents, images, web pages, and e-books can all be read aloud to you, with speed adjustments and voice changes. The conclusion is blunt: The subscription price isn't cheap, but it did change how the author digests information—listening to reports during commutes, listening to long articles while lying down. Looking at it today, there's something to say. The audio quality of AI reading crossed the threshold of "not sounding like a robot" just in the past two years. These products are moving from "barely listenable" to "changing habits." Whether the price can hold up is being repriced by the market. Product Expert Azhe says | The rule of this industry has always been my mantra: Entry point equals model. Previously, these listening/reading products won by voice tone. Now, large models have raised the baseline for reading quality overall. The moat has shifted from "sounding good" to "being embedded in your habits." For scenarios like commuting, housework, and eye protection, whoever welds the usage inertia in place survives. The journalist's phrase "expensive but changed how I digest information" is the signal that habits have been welded in. Editor Xiao He says | I've stepped into pitfalls with voice tools before. I mentioned earlier that my meeting notes were ruined by nearby chitchat. Tools that only listen without speaking are actually safer; they don't spy on you, they just read to you. But don't buy an annual card yet. Take a long report and trial it for a week. Only pay if you can really absorb it.

What it means for you|As permissions granted to AI grow larger, this public battle arena is essentially a public health check for the security baselines of various models.

HIGHLIGHTS IN ONE SENTENCE

●  3 key updates (with dual-expert commentary), 3 groups of leaderboard news flashes, 10 curated section highlights, plus "What Everyone is Watching" and "Tomorrow's Focus."

●  Let me tell you straight. Upon checking today, the latest snapshots for the four Arena blind test leaderboards are still from 2026-09-20, the same version used in the previous review. So the leaderboards haven't been updated; data is as of 2026-09-20. The Intelligence Index column was freshly scraped from Artificial Analysis today. Compare the two together.

● 

Text Blind Test Leaderboard: Anthropic Takes Seven of the Top Ten Spots

●  | # | Model | Vendor | Elo (Battle Count) |

●  |---|------|------|----------------|

●  | 1 | Claude Fable 5 (High) | Anthropic | 1506 (30,057 battles) |

●  | 2 | Claude Opus 4-6 (High) | Anthropic | 1505 (71,993 battles) |

●  | 3 | Claude Opus 4-7 (High) | Anthropic | 1502 (60,002 battles) |

●  | 4 | Muse Spark 1.2 (xHigh) | Meta | 1500 (Only 3,227 battles) |

●  | 5 | Claude Fable 5.1 (Max) | Anthropic | 1498 (5,783 battles) |

●  Plain language interpretation. In blind tests judged by humans, Claude's parent company nearly monopolizes the field, taking seven of the top ten spots. Meta's new model reaching #4 looks impressive, but it has only fought over three thousand battles. Its score could fluctuate by 11 points, so don't take the ranking seriously yet. The newly scraped Intelligence Index aligns with this conclusion: Claude Fable 5.1 and GPT-6 Astra tie at 53 points, occupying the head of the pack. The domestic tier formed by Alibaba's Qwen3.8 Max, Moonshot AI's Kimi K3, and Zhipu's GLM-5.3 is biting closely at 44 to 45 points; the gap is already very small.

● 

Code Blind Test Leaderboard: The Top Spot Has the Thinnest Sample Size

●  | # | Model | Vendor | Elo (Battle Count) |

●  |---|------|------|----------------|

●  | 1 | GPT-6 Astra (Max) | OpenAI | 1800 (Only 2,281 battles, ±16 points) |

●  | 2 | Claude Fable 5.1 (Max) | Anthropic | 1758 (3,036 battles) |

●  | 3 | Claude Opus 5 (Max) | Anthropic | 1687 (12,087 battles) |

●  | 4 | Qwen3.8 Max (0902) | Alibaba | 1681 (2,262 battles) |

●  | 5 | Kimi K3 (Max) | Moonshot AI | 1674 (4,547 battles) |

●  Plain language interpretation. The top score is scary, but the sample size is less than one-fifth of the third place. Check if new rankings wobble; wait for stability over one or two rounds before drawing conclusions. What's solid is the third place, backed by twelve thousand battles.

● 

Agent Execution Leaderboard: Getting Things Done is the Score

●  | # | Model | Vendor | Task Completion Rate | Self-Correction on Error | Tool Hallucination Rate |

●  |---|------|------|---------|-----------|-----------|

●  | 1 | Claude Fable 5.1 (Max) | Anthropic | 19.83 | 12.62 | 0.37 |

●  | 2 | GPT-6 Astra (Max) | OpenAI | 17.70 | 7.28 | 0.37 |

●  | 3 | Claude Opus 5 (High) | Anthropic | 9.24 | 11.98 | 0.33 |

●  | 4 | Claude Opus 5 (Max) | Anthropic | 12.41 | 13.08 | 0.35 |

●  | 5 | Claude Fable 5 (High) | Anthropic | 5.97 | 9.06 | 0.37 |

●  Plain language interpretation. This leaderboard has no total score, only six report cards. Ranked 6th, Claude Opus 4.8 has the lowest tool hallucination rate in the field at 0.12, meaning it makes up fake tools the least, but it was dragged out of the top five by its task completion rate. When choosing an AI to do work for you, look at dimensions first, then rankings. The 8th ranked Kimi K3 is the only domestic model in the top ten, with the thickest sample size exceeding 100,000 battles. No new snapshot for the image leaderboard this period, so it's not listed.

● 

  1. Apache releases Casbin Gateway, installing a security gate for AI coding assistants on computers. All coding AIs share one security gateway, uniformly managing which directories they can touch and which commands require human approval. Hanging under the Apache name means the open-source world has created a standard component for "don't hand the whole keychain to AI." Old Zhou reminds us: The direction hits the nail on the head; environment control is more effective than prompts, but a newly installed security gate needs its own inspection. Wait for real deployment cases.

● 

  1. AgentTrace installs a dashcam for AI agents. An open-source observation and self-healing engine. Every step an agent takes is archived; errors can be replayed and located, and automatic fixes/reruns can be triggered based on policies. Editor Xiao He's old saying rings true again: Being able to replay work is like having video footage—it makes it usable—but don't skip the frames you need to monitor yourself.

● 

  1. nitpicker, an open-source AI code reviewer where you keep the keys. A self-hosted PR review tool claiming code never leaves your machine. The slogan directly states "Keys in your hands." As AI output volume rises, the tool war for acceptance begins. Old Xu's stance: Verify AI code once before trusting it. The fact that the review tool itself is open-source and self-auditable counts as a bonus.

● 

  1. TechRadar real-world test: Letting AI guess what wine I'm drinking yields mixed results. Journalists fed taste descriptions to AI to guess the wine. Sometimes right, sometimes wrong. Flavor words match the framework, but details start getting confidently fabricated. A daily-life version of "verify AI before trusting": The more confident it sounds, the more you should ask "on what basis?"

● 

  1. Axel, an AI that gossips about you with others, shocks the privacy circle. The newly launched chat AI features gossiping about users, turning things you say to it into talking points brought into other conversations. Developers call it a social experiment. Old Zhou reminds us: Everything said to AI may default to being heard by a third party. Don't feed sensitive info. This experiment deserves a slap of sobriety.

● 

  1. Delightful Cells hands batch spreadsheet tasks to AI, automatically rerunning failed rows. Upload a spreadsheet, describe what each row should do in one sentence, and it runs them in batches. Failures auto-retry; completed rows aren't billed twice. Suitable for testing with unimportant lists first. Keep a copy of the original sheet before batch operations so you can roll back if things break.

● 

  1. Hideo Kojima responds to Sony withdrawal controversy: PHYSINT is still in development, no falling out. In August this year, Sony suddenly withdrew from the long-time collaborative new project PHYSINT, causing ugly rumors. Kojima recently responded directly: Amicable separation, project handed over to continue development. Big company withdrawal doesn't equal a death sentence for the project. Authors and investors part ways, games get made anyway—this will happen more often.

● 

  1. OpenCloak replaces personal info in prompts with codes before sending to AI. An open-source browser extension that automatically replaces names, phone numbers, and addresses with codes before sending, restoring them upon return. The idea is correct: Rather than demanding deletion afterwards, better not send it at all. Note that restoration must happen locally; avoid extensions that require returning to the source online.

● 

  1. Base Browser, a hardcore Firefox fork free of AI flavor. A new project on Codeberg that directly deletes all AI features from the browser for pure browsing. Going against the grain while AI floods settings pages, it hit the front page of trending lists within a day of launch. A control experiment: When everyone adds AI to browsers, the absence of AI itself becomes a selling point.

● 

  1. Engadget teaches you how to make WeChat and iMessage stickers with ChatGPT; simpler than expected. Describe the pattern, generate, cut into sticker packs. The main hurdle is the export size step. Lightly practical. Making one for elderly family members is more intimate than saying ten nice things. After generating, confirm platform sticker specs first; uneven white borders are the most frustrating part.

EVERYONE IS WATCHING

●  Newly scraped Intelligence Index today: Claude Fable 5.1 and GPT-6 Astra tie at 53 points, occupying the head. The domestic tier formed by Qwen3.8 Max, GLM-5.3, and Kimi K3 bites closely at 44 to 45 points. Faraday Future released nine robots at once, the most expensive exceeding 920,000 yuan. Gurman says Apple may release a smart home screen device as early as next month.

TOMORROW'S WATCH LIST

●  The four Arena leaderboards are still on the 09-20 version today. Watch if the 09-21 snapshot refreshes, and see if the battle count for Meta's new models exceeds 5,000.

●  Zhipu's MaaS data non-retention mechanism was just announced for recent launch. Wait for official access rules and real-world feedback.

●  Alibaba's open-sourced Qwen-Image-2.1 image model is hot today. Wait for third parties to run it against the same topics as the popular image leaderboard from the last issue. Don't just look at official demos.

●  The above is the preview for Issue #054. The formal version updates at 17:00. The full illustrated version is on the review publication page.

📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.

ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)