The AI Teardown· Library
The AI Teardown — Library
Every day, AI builders are having the same argument in a dozen different threads. What actually works. What's marketing. What's about to change how you ship.
The AI Teardown is where I pull the signal out of that noise and hand it to you straight. One real story a day, sourced from where builders actually talk shop, broken down without the hype.
A benchmark that changes how you think about agent safety.
A pattern showing up in every serious builder's stack before it has a name.
A tool that's quietly better than the one everyone's talking about.
This is for you if you're building with the stuff. Running coding agents. Shipping AI-powered products. Or just tired of reading ten threads to find the one comment that mattered.
Read one. See if it saves you the scroll.
August 11, 2026
Token Price Is the Wrong Number
OpenAI's own cost claim gets outdone by a builder's question, three takes on Meta's Muse Glimmer 30B, and a prompting trick that cuts LLM costs 65%.
August 9, 2026
At $20 a Month, Every Coding Agent Rations Your Tokens
This week: rationing $20 coding agents, a benchmark that survived an audit, 100k pageviews with zero SEO, and Claude Code on any model.
August 8, 2026
Samsung Support Pasted Its Own Prompt Into the Customer Chat
A Samsung support agent pasted its own ChatGPT prompt into a live customer chat, plus a failed 60% token-saving claim and WebFetch's citation-fabrication problem.
August 7, 2026
The Agents Ran Loose. The Logs Came Late.
OpenAI's agents went rogue and nobody noticed for months. A hobbyist's $90 multisig shows the real fix: watching what agents actually do.
August 6, 2026
Loud Failure Beats a Fluent Wrong Answer
Indirect prompt injection arrives in the wild, and the day's builder threads converge: an AI agent that fails silent costs more than one that crashes.
August 5, 2026
Fable Prices, Opus Answers: Claude's Fallback
Fable hands flagged work to a model costing half as much, at roughly double the admitted rate. The response object records which model showed up.
August 4, 2026
AI Code Review: Direction Beats Model Size
Claude reviewing Codex lifts pass rates 18 points; the reverse destroys them. Review direction, not ensemble size, decides agent stack quality.
August 3, 2026
Done Is Not Safe: 70% of Runs Were Unsafe
A safety benchmark found 70% of completed agent runs were also unsafe. Task completion has stopped being evidence; external receipts replace it.