Essays and articles on AI reliability. Select any card to read.
An agent can hand you a decision that is well-formed, confident, and wrong, and the check between it and the money is the last thing that can help. Asking the agent how sure it is does not help. Re-deriving the answer from a record it does not own does.
A classifier that is well calibrated on its eval can invert under shift, where the wrong answers come back more confident than the right ones. A single accuracy number never shows it.
Loop engineering moved the work from prompting the agent to designing the system that prompts it. It also moved the risk. When the loop runs while you sleep, “done” becomes the most expensive word in the system.
An accuracy score is an average, and the average is where the failure hides. One public autopsy, every number rerunnable.
Why every claim we make ships with code you can run.
The silent cost of AI-assisted engineering.
Why the most dangerous failure mode in machine learning is also the most dangerous failure mode in thinking.