Mastodon Skip to content
Artificial Intelligence

The Verification Asymmetry

I can audit a model's output on leadership and must take its tax advice on faith. That asymmetry is the whole problem, and it is not psychological. The third essay in the Prior Obligations series argues that AI does not deskill the competent—it unmasks the fake.

Six men in ragged clothing move in a diagonal line across a village landscape, each holding a staff and touching the man ahead; the leading figure has fallen backwards into a ditch.
Error travels down the line unchecked. Each man holds the shoulder of the one ahead with complete confidence, and not one of them is equipped to notice where the file is heading—which is the condition of any organisation where output passes from hand to hand and meets nobody able to verify it.
Published:
audio-thumbnail
The Verification Asymmetry
0:00
/494.616

I left a promise unpaid at the end of the first essay in this series: the combined harvester never claimed to think, and the AI model does, which is the one genuinely novel element in the present anxiety. This essay pays this debt, and the answer is less comfortable than the 'psychological debt' framework it displaces.

Taking myself as a worked example of this. If I put a language model to work on a question of delegated authority, or the conditions under which a team will tell its leader an unwelcome thing, I can audit the output claim by claim. I know where the literature is so thin as to be positively translucent, and where the confident summary generated has quietly elided a contested finding. I can tell when the machine has produced something that reads beautifully and means nothing, because I have spent years mastering the subject matter. Put the same model to work on the tax advice for a discretionary trust distribution and I am a different operator entirely. The prose has the same poise. The structure is as orderly. But, I have no way of knowing whether it is right, and every incentive—because of the fluency, the confidence, the absence of hedging—to assume it is guidance which won't have me fall foul of the ATO or my bank manager. Put another way, my competence to verify does not travel with me. It is domain-specific, non-transferable, and cannot be delegated to the AI model.

This is the whole problem, and it is not a psychological condition. It is an epistemic one, and it has a shape: the machine's output is least auditable at the point where the user's experience runs out, which is precisely where the user most needs it validating. Poor performers are systematically the worst judges of their own performance because the competence required to produce good work is largely the same competence required to recognise it. Add a tool that manufactures the surface features of expertise on demand—the orderly structure, the confident register, the plausible citation—and the person who cannot tell good work from bad is handed something that looks exactly like the good work.

What the Automation Literature Already Knew

None of this required the arrival of the language model to become visible. Lisanne Bainbridge set it down more than forty years ago in a paper on the ironies of automation, and the irony was precisely this: automating a process raises the skill demanded of the human overseeing it, while removing the daily practice by which that skill was acquired and maintained. The operator is asked to supervise a system they can no longer perform, and to intervene in exactly the abnormal conditions the automation could not handle—the hardest cases, met with the most atrophied judgement. Later work on human–automation interaction added the mechanism: whether people use, over-rely on, or refuse a system tracks their trust in it far more closely than its actual reliability, and calibrating that trust requires knowledge of the system's failure modes. Trust in an aviation autopilot can be calibrated because its failure modes are documented, drilled, and survivable in a simulator. Trust in a model that fabricates a case citation in flawless legalese cannot be calibrated by anyone who does not already know the case law.

So the disanalogy I flagged in Part I holds, but it does not vindicate the framework. What Champniss labels competency debt—the sense that using the tool makes one less capable—is not manufactured by the technology at all. The competent do not experience it in this form; they experience the tool as an instrument they can drive, argue with, and correct, precisely because they retain the capacity to catch it out. What the tool does is reveal, with unusual clarity and at speed, who could have caught it out and who could not.

The Certificate and the Competence

Which brings us to the more uncomfortable half of the problem. If verification is domain-specific, then an organisation's capacity to use these tools safely is simply the sum of the actual competence sitting in its roles. Not the competence its org chart asserts. The competence that is actually there.