Writing
Notes on verifying what AI systems do.
Essays and open data from the lab: where AI systems behave, where they fail, and how to prove the difference. Every piece links to something you can check or run yourself.
The AI everyone talks about is not the one developers run
The conversation is owned by a few frontier labs and their five-dollar models. The board developers actually route traffic on has DeepSeek at number one for nine cents, and the celebrated flagships mid-table at fifty times the price. Every number on it is checkable.
Model collapse is a provenance problem, and provenance is checkable
Over half the open web is now AI-generated, and the next training run will scrape it. The fix is not bigger models. It is data provenance, recorded at the source and checkable later. Why the scaling answer fails, and what a verifiable data discipline looks like.
The rating gap in the top AI assistant apps
Every leading AI assistant shows a near-perfect lifetime App Store rating. Among recent reviewers, several sit more than a star lower. Here is the cited population data, and what the gap can and cannot tell you.
The Friction Matrix: how top Entertainment apps handle their backlash
Every top streaming app keeps a near-perfect lifetime rating while recent reviewers revolt. Plotting that backlash against how often developers reply sorts them into Firefighters and Ghost Ships.
The Friction Matrix: the Productivity chart's silent giants
AI assistants now dominate the App Store Productivity chart, with near-perfect lifetime ratings, unhappy recent reviewers, and almost no developer replies. Here is who responds when users turn, and who stays silent.
The Friction Matrix: the Productivity chart's 30-second take
Population truth versus recent mood for the most-rated US Productivity apps, in brief.
Google externalised the cost of renaming Gmail
Google shipped Gmail address renaming and never shipped a webhook. 124 open-source projects still key OAuth identity on email. What breaks, and who pays.
124 repositories. Four ecosystems. One broken assumption.
The reproducible data behind the essay: 2M+ repositories scanned, severity tiers, ecosystem breakdown, and the full methodology. A complete audit.