◇ EXPERIMENT
AI AGENT · COST & EVALUATION
RepoCoach
Two paper estimates, both wrong — only measurement got it right. A vertical slice of a source-reading agent whose real output is data on the counter-intuitive parts of agent engineering: cost grows super-linearly with call count, 28% fewer tool calls raised tokens by 17%, and a 65.6% cache-hit rate inflated the budget metric roughly threefold. It also pinned down how safety gates have to be checked in both directions, how state survives across processes, and why a constraint written only into a prompt still needs a gate at the exit. 638 tests, 22 merged PRs — planned by Claude, implemented by DeepSeek, reviewed by GPT. Development has since been wound down, with the measured parts written up as reusable findings.








