Worth Confirming

Most AI deployments this year have settled on the same safeguard. A person sits in the loop and confirms before the system acts. The policy took about a year to become standard and it is a good one.

What the person is looking at when they confirm has had less attention. In most of the systems I have seen, that part was never designed at all.

I wrote a few weeks ago that removing the person was never the goal, that the goal is giving the person something worth confirming. It was one sentence and it deserved more than that, because the gap between a person in the loop and a person who is right to confirm is where most of these systems quietly fail.

The Size of the Click

The first mistake is what the button asks.

Most of the time it asks the person to agree that the system got the answer right. Here is the cause. Here is the fix. Click to confirm. The person looking at the screen usually did not run the analysis and often could not. They are being asked to sign off on something they have no way to check.

People do one of two things with that. Some click yes, because the system sounded sure and arguing with it would take an hour. Others go do the whole investigation themselves, because their name is on it now and they are not signing something they do not understand.

Both defeat the purpose. The first one means nobody actually checked. The second one means the system saved no work at all.

The better question is smaller. Not “is this the answer” but “is this real, and should someone look at it.” That is a question a person can answer in a minute. They do not need to know the cause. They need enough to say “yes, this matters, send it on” or “no, this is noise,” and to pass along what they saw. That is the decision people in charge are good at. It is also the moment where the expensive work should get handed to the right person instead of landing on whoever happened to be looking.

What the Person Is Owed

The second mistake is what gets put in front of them.

I did not see this one until someone said it to me directly. We were looking at the same screen, one I thought was clear, and she said, “You’re an engineer, I’m not, and I don’t know what action to take on that page.” I had been reading the evidence and assembling the story without noticing I was doing it. She was being asked to confirm something the page had never actually said.

A proposal worth confirming contains three things. What the system thinks happened. What supports that read. What would change it. A claim, the evidence behind it, and the thing that would make it wrong.

Most systems show the middle one and skip the other two. They hand the person the evidence, the raw signals, the grouped events, the confidence scores, and leave the story to be assembled on the spot. The claim is implied. The falsifier is absent. The person is being asked to confirm a conclusion the system never stated, on the basis of material the system never explained.

The instinct is to solve this by showing less. That is wrong. The evidence has to stay, because nobody can judge whether something is real without seeing what happened and what got grouped together. The fix is not less information. It is the same information with the read attached. This is what I think. Here is why. Here is what would change my mind.

Humans do this for each other constantly. Someone walks a colleague through a mess and says “here’s what I think is going on” before they show a single log. Systems built to keep humans in the loop mostly stopped short of that sentence, and the sentence was the whole point.

The Rule Is Older Than the Technology

None of this is new to AI. It is the rule for any recommendation that lands on a decision-maker’s desk. The good ones arrive as a claim with its support and its risks. The bad ones arrive as a deck full of analysis and a slide that says “options for discussion.”

Leaders have been asking their teams to bring something worth confirming for as long as there have been teams. The machines arrived and got a pass on the same standard, because the output looked thorough and thorough felt like enough.

I keep coming back to the falsifier. The claim and the evidence are hard enough to get a system to state. The third part, what would change the read, is the one I have not seen done well anywhere, including by people. It might be the part that makes confirmation honest. I am not sure yet how to build it.