If you're making decisions about who — or what — to trust with consequential work, the research on interviews is a cold bucket of water: confidence and competence barely correlate. That applies the moment your agent roster starts expanding.
The Farnam Street essay on why job interviews fail makes a bracing case that the format most companies treat as a cornerstone is closer to a coin flip dressed in business casual. Unstructured conversation lets interviewers anchor on likeability, narrative smoothness, and physical presence — none of which predict job performance with any reliability. What does work: structured questions asked consistently, work sample tests, and taking base rates seriously before a single candidate walks in the door. The lesson is that conviction formed from a single vivid conversation is almost always noise.
Replace "candidate" with "agent" and the essay reads like a design checklist. Founders routinely evaluate agent behavior the way bad interviewers evaluate humans — they run one impressive demo, feel the pull of fluency, and ship. The actual discipline is structured evaluation: consistent test sets, real task samples from your domain, and a sober prior about failure rates before you see any output. An agent that converses smoothly is not an agent that performs reliably. Build the rubric first, then run the conversation, and weight the rubric every single time.
- gut confidence after one demo is as unreliable for agents as for people
- work-sample evals on your actual tasks beat general benchmarks
- base rates about failure modes deserve more weight than any single impressive run.
