The Agent Builds What You Specify
An eval harness tells you whether an agent did the job correctly. It cannot tell you whether the job was worth doing. Fitzpatrick's core argument is that founders systematically collect worthless signal — compliments, hypotheticals, politeness — and mistake it for validation. The same failure mode now runs at machine speed. If the task definition fed to your agent rests on an untested assumption about customer behavior, every successful eval confirms a mistake. Garbage in, verified garbage out.
Ask About Their Life, Not Your Idea
Fitzpatrick's method is behavioral: ask what people actually do, how often, what they tried, what it cost them. This is the exact input that should shape your task design and tool surface before you build anything. A customer who says they struggle with a workflow three times a week and have already paid for a workaround is telling you the shape of a real job. That specificity is what makes a well-scoped tool — one with a clear name, tight permissions, structured output — possible in the first place.
Commitment Signals Where Judgment Lives
Fitzpatrick draws a hard line between conversation and commitment: time, reputation, and money are the only signals that matter. For an AI-native founder, commitment also reveals where human judgment must stay in the loop. A customer who will not hand over a real decision — even a small one — to an automated workflow is telling you something about trust and accountability that no eval can surface. That hesitation is design input. It tells you where the human stays in the seat, and building against it rather than with it is how autonomous products fail.
