The way you audition an agent is almost certainly broken in the same way your hiring process is broken. If you're running vibe checks on outputs, you're not evaluating — you're confabulating.
The Farnam Street essay on job interviews lands one clean punch: unstructured conversations feel predictive and aren't. We walk out of an interview convinced we've read a person, when mostly we've read our own preferences back at ourselves. The research it draws on is consistent — structured questions, work samples, and base-rate thinking outperform the confident narrative we construct in real time. Hiring conviction built on impression is noise that happens to wear a blazer.
Swap "candidate" for "agent pipeline" and the essay reads like a warning you should have gotten six months ago. Founders eval their agents the way bad hiring managers eval candidates — a few impressive outputs, a smooth interaction, and suddenly there's trust. What the essay pushes toward instead is structured, repeated, task-specific testing against known ground truth, plus honest accounting for base rates: how often does any agent get this class of task right? Charisma in an agent is fluency. Fluency is not accuracy. Build the rubric before you run the demo.
- Define eval criteria before you see any output, not after
- use work samples on tasks the agent will actually face in production, not showcases it was tuned to perform
- track base rates across runs so a single impressive result doesn't become unearned institutional trust.
