Done Means Verified
An agent is not finished because it acted. It is finished when it has evidence that the user’s objective was achieved.
Read the essayEssays on reliability, human agency, evaluation, and the systems that turn capable models into dependable software.
Long-form arguments, grounded in the product and its code.
An agent is not finished because it acted. It is finished when it has evidence that the user’s objective was achieved.
Read the essayA public place to try an agent must be safe enough to control—and ordinary enough that the agent cannot see the controls.
Read the essayFrom a failed run to a better question.
Sixteen turns. A premature finish. An honest retry. Why I want an internal benchmark—even when I’m also building the agent.
Read the essayNot model capability in isolation, but the system around it.
Verification, read-backs, traces, and the difference between an action taken and an outcome achieved.
Mandates, reversibility, approval boundaries, and preserving the consequential decisions for people.
Harnesses, tools, recovery, evaluation, and the architecture required for dependable work on the web.