ReviewRadar

GLOBAL AI REVIEW RADAR

2026.09.22 · Tuesday

Issue 042

Key Updates 1 | Leaderboard Flash 0 | Highlights 39 | Tomorrow's Watch 4

KEY UPDATES

Update 1

1

1. The claim that "80% of AI projects fail" traces back to just one footnote

This conclusion, often heard in meetings, has been treated as industry common knowledge for years. Independent consultant Jamie Watters traced it back and found that consulting firms, vendor presentations, and industry media all cite it, but the earliest source is just a single footnote with no sample size and no definition of which projects were counted. In other words, this "80%" feels more like an impression repeated through hearsay than a publicly verifiable statistic. The relevance to ordinary people is real. It's the kind of number that decides whether your proposal lives or dies. If a boss says "80% will fail," your budget might vanish. Next time you hear it, ask who calculated it and what projects they included—this works better than arguing. Programming Expert Lao Xu: I've seen numbers like this too many times. After circulating for a few years, nobody remembers the source. Unlike benchmark scores, you can't rerun it on the spot; once someone cites it, you can't even find an entry point to refute it. If you really want to know if AI projects work, count how many your own team finished last quarter and put the success rate on the table. Editor Xiao He: My pitfall was this. Hearing "80% fail," I agreed to a small-step pilot plan, only to find the bottleneck wasn't that AI didn't work, but that no one defined who would accept the deliverables. Now I only look at two things: who accepts it, and whose head rolls if it fails.

2. Adding an "ingredient label" to AI-written reports to catch mismatches with raw data instantly

A new open-source free tool called factlabel acts like the ingredient list on food packaging. When AI writes a conclusion, it cross-checks every sentence against the cited tables and figures, flagging and blocking anything fabricated or mismatched. The code is on GitHub; deploy it yourself to run it without paying fees. This addresses a very daily annoyance. The most frustrating thing about AI getting numbers wrong isn't the error itself, but how smoothly it writes them as if they're true, leading lazy humans to forward them blindly. With this check, you at least know which sentences require verifying the original source first. Security Expert Lao Zhou: This direction hits the nail on the head. My old rule remains: if AI reports a number, where is the raw table? If you can't find it, don't use it yet. Now there's a tool that automatically checks this, saving half the manual verification effort. But it only verifies consistency with cited data; if the data itself is wrong, it can't stop that—you still need humans. Editor Xiao He: Last time I had AI organize expense reports, it fabricated amounts and dates. I spent hours checking against original receipts. With this tool, I'd at least know which lines to check first instead of rereading the whole thing.

3. Don't wait for cloud prices to drop; someone open-sourced the full workflow for running frontier models on your own PC

Use your own GPU-equipped computer as a server, run models locally, keep data private, and avoid monthly subscriptions. Tutorials for this have always been scattered. Tim Dettmers, a widely recognized clear explainer in this area, released the entire environment, dependencies, and pitfalls from his course, including how to downgrade performance when GPUs are insufficient and which steps cause bottlenecks. Whether it's worth doing depends on your machine. This path saves subscription fees and privacy concerns but costs tinkering time. Wearables Expert A Kai: I care about noise, heat, and electricity bills. High benchmark scores mean nothing if the machine buzzes in the living room and gathers dust. Whether the tutorial specifies "which card runs at what level and stays stable after an hour" matters much more than peak speed. Editor Xiao He: My old rule stands: bookmark these tutorials first, don't act yet. Once my old PC successfully runs a small task, then we'll talk about adding RAM.

What it means for you|If a boss says "80% will fail," your budget might vanish; next time you hear it, ask who calculated it and what projects they included.

HIGHLIGHTS IN ONE SENTENCE

●  3 Key Updates (with dual-expert commentary), 3 Leaderboard Snippets, 10 Curated Picks, plus "What Everyone Is Watching" and "Tomorrow's Focus."

●  Let me be honest. Today's check confirms the latest snapshots for the four Arena blind test leaderboards are still from 2026-09-21, the same version used in the previous review. So, leaderboards are not updated; data is as of 2026-09-21. No new leaderboards published within 72 hours were found to replace them. We prefer to state facts rather than pass off old data as new.

● 

Text Blind Test Leaderboard: Anthropic takes seven of the top ten spots

●  | # | Model | Vendor | Score (calculated by human votes; parentheses show match count) |

●  |---|------|------|----------------|

●  | 1 | Claude Fable 5 (High) | Anthropic | 1506 (30,057 matches) |

●  | 2 | Claude Opus 4.6 (High) | Anthropic | 1505 (71,993 matches) |

●  | 3 | Claude Opus 4.7 (High) | Anthropic | 1502 (60,002 matches) |

●  | 4 | Muse Spark 1.2 (xHigh) | Meta | 1500 (only 3,227 matches) |

●  | 5 | Claude Fable 5.1 (Max) | Anthropic | 1498 (5,783 matches) |

●  Plain language interpretation. In blind tests judged by humans, Anthropic dominates, taking seven of the top ten spots, with the top three all theirs. Meta's Muse reaching #4 looks impressive, but it has only played over 3,000 matches, meaning its score could fluctuate by 11 points. Don't take the ranking seriously yet.

● 

Coding Leaderboard: The top spot has the thinnest sample size

●  | # | Model | Vendor | Score (parentheses show match count) |

●  |---|------|------|----------------|

●  | 1 | GPT-6 Astra (Max) | OpenAI | 1800 (only 2,281 matches, ±16 points) |

●  | 2 | Claude Fable 5.1 (Max) | Anthropic | 1758 (3,036 matches) |

●  | 3 | Claude Opus 5 (Max) | Anthropic | 1687 (12,087 matches) |

●  | 4 | Qwen3.8 Max (0902) | Alibaba | 1681 (2,262 matches) |

●  | 5 | Kimi K3 (Max) | Moonshot AI | 1674 (4,547 matches) |

●  Plain language interpretation. The top score is scary, but the sample size is less than one-fifth of #3. Watch if the new ranking wobbles before drawing conclusions. The solid third place is backed by 12,000 matches. Domestic models Qwen and Kimi both made the top five.

● 

AI Assistant Practicality Leaderboard: Did it get the job done?

●  | # | Model | Vendor | Sample (match count) |

●  |---|------|------|----------------|

●  | 1 | Claude Fable 5.1 (Max) | Anthropic | 13,320 |

●  | 2 | GPT 6 Astra (Max) | OpenAI | 10,372 |

●  | 3 | Claude Opus 5 (High) | Anthropic | 24,794 |

●  | 4 | Claude Opus 5 (Max) | Anthropic | 19,934 |

●  | 5 | Claude Fable 5 (High) | Anthropic | 38,293 |

●  Plain language interpretation. This leaderboard has no total score, only six individual report cards. Ranked #6, Claude Opus 4.8 has a tool hallucination rate of only 0.12, the lowest overall, meaning it fabricates tools the least, but it fell out of the top five due to its lower "actually completed tasks" ratio. When choosing an AI to do work for you, look at dimensions before rankings. #8 Kimi K3 is the only domestic model in the top ten, with over 100,000 matches—the largest sample.

● 

  1. PokerTools Arena, a public testing ground where AI plays Texas Hold'em together. Several models are placed on the same No-Limit Texas Hold'em table, with every step—dealing, calling, folding—recorded on the webpage for replay to see where it made mistakes. Open source and free. Poker is typical of "making continuous decisions with incomplete information," closer to real-world work than multiple-choice questions; but it only tests poker, so don't treat it as a comprehensive capability ranking.

● 

  1. Agent Harness Replay, records and replays the entire process of AI writing code. It captures a real AI programming session, laying out every request received by the model and every tool call. Previously you could only see the final code; now you can inspect every intermediate step. If something goes wrong, you at least know which screen to look at.

● 

  1. Callwitness, records exactly what AI retrieved from tools. Which tool was called, what content was returned—all receipts are kept. This record helps troubleshoot cases where "it clearly called the API, but the answer was fabricated." The downside is it only records; it doesn't judge correctness.

● 

  1. Real-world log of using AI for bug bounties. A security researcher publicly shared their complete process of using TypeSafe AI to find vulnerabilities, detailing where it helped, where human judgment was needed, and even unsuccessful attempts. Publishing failure logs is more valuable than only sharing successes, but this is a single-person sample; don't generalize it.

● 

  1. FearGate, a website dedicated to tracking "unfulfilled promises" in the AI circle. Grand AI claims made by vendors and media are listed chronologically; expired unfulfilled ones are marked overdue. The homepage currently lists several expired claims. Checking past fulfillment before believing AI hype is useful for ordinary people; but inclusion and judgment are decided by one site admin, so treat it as a lead only.

● 

  1. New algorithm for "slimming down" large models, solved as a physics problem. Formulating which layers to keep and which to remove as an optimization problem in physics allows the model to shrink, enabling larger versions to run on the same hardware. The direction is interesting, but this is a research blog, not a public exam. How much capability is lost after slimming down lacks third-party data.

● 

  1. Jev's web playground, an AI that only asks for "Yes / No" judgments. Input a question, and it gives only two buttons, e.g., "Is this request urgent?" No beating around the bush. Giving conclusions without chitchat suits processes requiring binary decisions; but you still need to verify if the answer is correct.

● 

  1. First domestic open-source model sold on AWS. Moonshot AI's Kimi K3 is listed on Amazon Web Services, allowing overseas developers to call it directly or download the model for self-deployment. Whether open-source models can sustain themselves financially affects your future access to them more than their benchmark scores.

● 

  1. deslop-skills, a skill pack to "de-grease" AI outputs. A set of mini-skills for AI coding assistants targeting two common issues: uniform-looking interfaces and long, empty text. Installed, it makes AI self-check before delivering work. It treats common ailments, and the idea is copyable; but it provides a checklist—whether the model listens is another matter.

● 

  1. Practical compliance guide for AI assistants. This long article breaks down EU AI Act provisions step-by-step, explaining what materials and records an AI assistant needs for launch, including when human oversight is mandatory. Aimed at companies, not individuals, but readers with overseas business can learn what documents are required.

EVERYONE IS WATCHING

●  AMD's market cap hit $1 trillion for the first time, driven by AI chip demand. Amazon blocked Meta's Muse shopping assistant from its platform after negotiations failed. California's governor signed seven bills specifically regulating water and electricity usage for AI data centers. Jensen Huang reiterated that predictions of AI exterminating humanity within ten years are completely wrong.

TOMORROW'S WATCH LIST

●  The Yunqi Conference opens today; Alibaba's new Qwen lead Liu Dayiheng presents his first report.

●  EU data center energy disclosure rules enter the legislative process; costs for hosting overseas need recalculation.

●  Samsung forms a robotics team for chip factories; what is the first task?

●  Above is the preview for Issue 055. The official version updates at 17:00. Full illustrated version available on the review journal page.

📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.

ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)