ReviewRadar

GLOBAL AI REVIEW RADAR

2026.08.30 · Sunday

Issue 020

Key Updates 10 | Leaderboard Flash 0 | Highlights 0 | Tomorrow's Watch 0

KEY UPDATES

Update 1Google Gemini 3.7 Flash Climbs Three Leaderboards in Two Days: Chat #9, Coding #8, Task Execution #20

What it can help you do: This is Google's volume-oriented AI released on August 13, focusing on cheap and fast. New data published by Arena's official account on August 28: Chat blind test score 1490, ranked #9 (leaderboard snapshot as of 08-26, 5,718 battle samples); WebDev coding leaderboard jumped from #19 to #8; Agent leaderboard (testing if AI can run tasks and get things done) entered at #20, with "net improvement" +3.4 and "task completion rate" +10%, ranking #5 overall. Compared to predecessor Gemini 3.6 Flash: Agent #35, Coding #21—a leap with generational change. Pricing is aggressive too: Reading 1 million words costs ~$0.75, writing ~$3.75 (AI billed per token, one token ≈ 0.75 Chinese characters), limited-time price. The August LLM weekly report also noted it outputs 343.7 tokens per second, fastest overall, ranking #10 in comprehensive scores (compiled by third party, cited from Zhihu). Expert Commentary · Lao Xu (Programming): Last issue I warned: wait for rankings to stabilize for a round or two before judging. Today's report still has thin samples—Chat leaderboard has 5,718 battles, while veterans often have tens of thousands; Coding leaderboard has only 2,552 for it. The direction is truly strong: predecessor Gemini 3.6 Flash ranked #21 in coding, yet one generation lifted it into the top ten, showing Google got the "fast and cheap" line right. But old rule: treat vendor launch stats and newly entered ranks as unverified until tested. Run real tasks for a week before integrating into projects to see where it breaks. Editor Xiao He: Price is indeed attractive, but I continue the approach from Issue #031: watch for volatility in newly ranked models. Also, "speed" is most intuitive for ordinary users like me—no waiting for spinning wheels when replying. I'll try writing weekly reports with it first, keeping important tasks for the provider I've used for six months.

Update 2Someone Created a "Criminal Record" Database for AI Hallucinations: Latest Entry from Aug 27 US Court Records

What it can help you do: A website called the AI Hallucination Case Library collects real records of AI "confidently making things up": fabricated legal precedents, non-existent citations, invented facts, recording parties, times, and sources item by item. On August 29, it surfaced on Hacker News; the latest entry is from an August 27 lawsuit in Florida, USA, where someone used AI-fabricated material as evidence. The value of such libraries lies in aggregating scattered failure cases into a searchable ledger—next time you see AI confidently listing "sources," check against this ledger first. Expert Commentary · A Zhe (Product): I position these as "buyer review archives." Model quality is largely defined by vendors; but cases exposed in court or debunked by journalists cannot be PR'd away. Reclaiming scoring power from vendors relies not on another benchmark list, but on negative lists with case numbers. It doesn't tell you which model is strongest; it tells you all models may bury landmines in your most serious report. Editor Xiao He: In Issue #028 I said "check buyer reviews before buying"; now even failure showcases are archived. When using AI to write, verify every material, link, and number provided before forwarding—my old habit: verify before trusting AI.

Update 3Simurg: Free Search Tool for AI – Prefers Interruption Over Hallucination. Newly launched open-source tool providing free web search for AI agents, selling point: "abort hallucination." If no source supports the answer, it stops rather than continuing to fabricate

Simurg: Free Search Tool for AI – Prefers Interruption Over Hallucination. Newly launched open-source tool providing free web search for AI agents, selling point: "abort hallucination." If no source supports the answer, it stops rather than continuing to fabricate. One-line comment: "Stop rather than lie" is a targeted approach; tool is new, interception effectiveness awaits third-party testing.

Update 4Documentation.ai.md: Proposes Standard for "Docs Written for AI." Open proposal: Software docs shouldn't just be for humans, but for users' AI assistants to operate directly—software now has two readers: those deciding usage, and agents doing the work

Documentation.ai.md: Proposes Standard for "Docs Written for AI." Open proposal: Software docs shouldn't just be for humans, but for users' AI assistants to operate directly—software now has two readers: those deciding usage, and agents doing the work. One-line comment: Docs are evolving a new species "read by machines"; whoever standardizes first owns the entry point.

Update 5Programmer's Long Post: Beyond Coding, First Real Time-Saving AI Use is Grocery Shopping. A developer reviewed that outside coding, the only stable time-saving AI use is planning weekly grocery lists—considering budget, dietary restrictions, and expiring fridge items

Programmer's Long Post: Beyond Coding, First Real Time-Saving AI Use is Grocery Shopping. A developer reviewed that outside coding, the only stable time-saving AI use is planning weekly grocery lists—considering budget, dietary restrictions, and expiring fridge items. One-line comment: AI's practical value often hides in unglamorous chores; personal logs are more credible than press releases.

Update 6BoqCalc: AI Prices 500-Line Engineering Quotes, Focuses on "No Fabricated Unit Prices." Construction industry AI pipeline: upload bill of quantities, it calculates costs, flags hidden risks, locking quote unit prices in verifiable data, preventing model improvisation

BoqCalc: AI Prices 500-Line Engineering Quotes, Focuses on "No Fabricated Unit Prices." Construction industry AI pipeline: upload bill of quantities, it calculates costs, flags hidden risks, locking quote unit prices in verifiable data, preventing model improvisation. One-line comment: Another design sample preferring "no calculation over wrong calculation"; engineering errors are costly, wait for real project cases before use.

Update 7AI Begins Touching Cryptography Foundations. Security podcast invited cryptographer Chris Peikert to discuss recent progress: AI beginning to automate proofs on mathematical problems like "lattices"—lattice hardness assumptions underpin post-quantum cryptography

AI Begins Touching Cryptography Foundations. Security podcast invited cryptographer Chris Peikert to discuss recent progress: AI beginning to automate proofs on mathematical problems like "lattices"—lattice hardness assumptions underpin post-quantum cryptography. AI can help verify, theoretically also find flaws. One-line comment: AI moving from "solving problems" to "proving theorems" shocks then delights crypto circles; "AI breaking encryption" remains a hypothetical drill, not reality.

Update 8Security Researcher Public Call: Send Me Your AI Agent's Weirdest Logs. Researcher soliciting "weirdest agent logs" from users—records of AI going off-track or acting autonomously while working, to study real behavioral boundaries

Security Researcher Public Call: Send Me Your AI Agent's Weirdest Logs. Researcher soliciting "weirdest agent logs" from users—records of AI going off-track or acting autonomously while working, to study real behavioral boundaries. One-line comment: Agents' most honest capability material isn't demos, but failure process logs; folk samples are scarcer than papers.

Update 9Veteran E-book Manager Calibre Adds AI-Generated Covers. Select a book, let AI generate cover art based on title/content without design skills

Veteran E-book Manager Calibre Adds AI-Generated Covers. Select a book, let AI generate cover art based on title/content without design skills. One-line comment: AI image gen is becoming default feature in daily software; check copyright terms for public use.

Update 10"Ask AI" Button: Websites Starting to Embed AI Entries Like Social Shares. No-registration widget: Site owners generate "Ask AI What This Site Says" buttons; visitors click to chat with their own AI assistant

"Ask AI" Button: Websites Starting to Embed AI Entries Like Social Shares. No-registration widget: Site owners generate "Ask AI What This Site Says" buttons; visitors click to chat with their own AI assistant. One-line comment: Users already habitually ask AI before clicking websites; sites are embedding AI...

📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.

ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)