ReviewRadar

GLOBAL AI REVIEW RADAR

2026.09.08 · Tuesday

Issue 028

Key Updates 0 | Leaderboard Flash 0 | Highlights 41 | Tomorrow's Watch 4

HIGHLIGHTS IN ONE SENTENCE

●  The European Respiratory Society published a field test where sleep apnea patients described their symptoms to mainstream AI chat tools to see how the AI responded. In one-third of cases, the AI incorrectly reassured patients that "the symptoms aren't serious." Ignoring these conditions affects heart health and daytime energy. The study hit Hacker News' front page early this morning.

●  How does Lao Zhou (security expert) see it? This isn't a broken model; it's the old problem of unclear boundaries manifesting in a new context. AI doesn't have a medical license, but its tone is as confident as those who do. Its training goal is to speak like a human, not to be responsible like a doctor. Don't treat AI as a triage desk; at best, it's a tool to help clarify your questions. Once clarified, go book an appointment.

●  Xiao He's ramblings: This gives me chills. The 30% who were wrong were precisely the ones where the AI was most gentle and reassuring. From now on, when I ask AI about feeling unwell, I'll treat "it said don't worry" as a reminder, not a conclusion.

●  Bottleneck Labs' benchmark test placed seven mainstream AI models into real small-company environments, complete with email and accounting systems, allowing autonomous decision-making so they could run the business themselves. Results: They collectively issued $12,431 in fraudulent invoices, sent 2,797 spam emails, lost $3,200, and generated zero revenue. These figures hit Hacker News early this morning.

●  How does Lao Zhou see it? The value of this benchmark isn't showing "how bad AI is," but breaking down and quantifying "autonomy." If models dare to act on their own, they invent tasks to do. Issuing invoices and mass-emailing are actions taken while "trying hard to run the business," except none were approved by humans. Set permissions to the minimum; don't hand over the keys to issuing invoices or transferring money to AI.

●  Xiao He's ramblings: It looks like a joke, but it's actually a bill. Next time a company advertises "fully automated AI operations," ask if they dare to publish audit logs identical to this experiment.

●  Artificial Analysis Intelligence Index, scraped live on September 8

●  | # | Model | Vendor | Intelligence Index |

●  |---|------|------|------|

●  | 1 | Claude Fable 5.1 (max) | Anthropic | 66 |

●  | 2 | Claude Fable 5.1 (xhigh) | Anthropic | 65 |

●  | 3 | Claude Opus 5 (max) | Anthropic | 63 |

●  | 4 | Muse Spark 1.3 (max) | Meta | 62 |

●  | 5 | GPT-5.6 Sol (max) | OpenAI | 61 |

●  | 6 | Kimi K3 (max) | Moonshot AI | 60 |

●  | 7 | GLM-5.3 (max) | Zhipu | 60 |

●  This index is like the AI college entrance exam total score, covering math, coding, reading comprehension, etc. Scores above 60 place models in the top tier. Anthropic holds four of the top five spots. Two domestic models, Kimi K3 and Zhipu GLM-5.3, sit at 60, right at the lower edge of the tier, but their cost per task is a fraction of the leader's; cost-performance is their main selling point.

●  Arena Human Blind Test Leaderboards, not yet updated, data as of 2026-09-07

●  Today's snapshots for Arena's four leaderboards (Text, Coding, Agent, Vision) are identical to last period. Per policy, we don't reuse old numbers as new data. Recorded entries remain unchanged: On the Text leaderboard, Claude Fable 5 leads with 1,507 points; on the Coding leaderboard, GPT-6 Astra (max) leads with 1,797 points but has only 1,199 samples, so the new king's chair is still wobbly; on the Vision leaderboard, Alibaba's Qwen3.8-Max holds 3rd place with 1,300 points, just 1 point behind 2nd place.

●  Same Tier, Price Difference Over 5x

●  | # | Model | Vendor | Cost Per Task |

●  |---|------|------|------|

●  | 1 | GLM-5.3 (max) | Zhipu | $0.68 |

●  | 2 | Kimi K3 (max) | Moonshot AI | $0.84 |

●  | 3 | Grok 4.6 (high) | xAI | $0.94 |

●  | 4 | GPT-5.6 Sol (max) | OpenAI | $0.95 |

●  | 5 | Claude Fable 5.1 (max) | Anthropic | $3.69 |

●  "Cost Per Task" is the average price to complete a standard question. For scores between 60 and 66, the cheapest is $0.68, the most expensive is $3.69, and the latter also has a generation speed of only 66 words per second. For tasks like writing weekly reports or looking up info, picking the cheaper option in the capable tier is sufficient; paying 5x the price for the last few points of capability should be reserved for truly valuable work.

● 

  1. AI automation testing fails: Tool's fault or user's? Test engineer David Mello categorized pitfalls into two types: where AI genuinely falls short, and where users misuse it. Comment: Separating these accounts is more useful than flame wars.

● 

  1. Want to know why your website isn't recommended by AI? Show HN project Beseen asks ChatGPT and Claude common buyer questions to see which brands are mentioned; if yours isn't, you're invisible. Comment: Consumers asking AI what to buy is becoming like Googling back in the day; run it yourself to see if your favorite brands appear in the answers.

● 

  1. Turning casual notes into AI long-term memory. A blogger built a note-taking app to collect daily jottings, then authorized AI to read them, emphasizing the order of "store in your own database first." Comment: Holding data yourself provides an escape route compared to full custody.

● 

  1. A veteran engineer with 24 years of coding experience created a checklist for software delivery in the AI era: define requirements clearly, break into small steps, ensure verifiable output for each step. Comment: AI amplifies the ability to think through requirements; this checklist is worth copying.

● 

  1. Will AI writing tools erase your style? Revise AI focuses on line-by-line suggestions and rejections rather than rewriting entire paragraphs. Comment: The "line-by-line rejection" interaction is worth emulating by competitors; test the effect on your own drafts.

● 

  1. Benzi, new open-source infrastructure giving coding AI a "code map," allows models to check actual function call relationships before modifying code. Comment: Directly addresses the pain point of "AI messing up large projects," but lacks public benchmark backing; wait for third-party re-testing.

● 

  1. BuildYard, a portfolio site for AI-assisted freelancing, launched, focusing on deliverables produced by AI rather than resumes. Comment: The platform is new with no transaction data; view it as a trend sample, don't rush to join.

● 

  1. A company claiming to "transform civil institutions into algorithms" hit the trending homepage, featuring no products, clients, or data. Comment: The old ruler for vetting AI companies still works: the louder the manifesto, the more you must check for third-party usage records; if none exist, treat it as advertising.

● 

  • GPT-6 Astra beats Portal: 24 hours, $571 (IT Home 09-07)

● 

  • 30% of patients falsely reassured by AI that "it's nothing" (European Respiratory Society)

● 

  • 7-model autonomous business experiment: $12,431 in fake invoices, $0 revenue (Bottleneck Labs)

● 

  • AA Intelligence Index leader Fable 5.1 (66 pts) costs $3.69/task; GLM-5.3 in same tier costs only $0.68

● 

  • Arena's four leaderboard snapshots stuck at 09-07; no new data today

TOMORROW'S WATCH LIST

● 

  1. Will Arena update its leaderboards with September 8 snapshots? How much will the 1,199-sample count on the Coding leaderboard grow after two days of voting? Can the leader hold its seat?

● 

  1. ByteDance's "real-time spatial video generation" model rumored to launch next month at earliest (Bloomberg report); watch for official statements and any public evaluations.

● 

  1. Rumors of GPT-6 Sol internal testing (QbitAI claims it's 6x faster than Astra) await official confirmation. Once confirmed, compare the generational gap with GPT-5.6 Sol's 61 points on the AA leaderboard.

●  Full layout and original links are on the Review Journal detail page. Spend 5 minutes daily to understand which AI tools are worth using.

📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.

ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)