My AI Agents Were Wrong a Third of the Time. Their Citations Looked Perfect.

What I learned when I asked a second set of agents to prove the first set wrong.

Every one of them showed its work. It just wasn't the work I'd asked for.

- Me, closing a ticket I shouldn't have

This summer I sat down with a client's backlog of a few hundred tickets. A lot of them were old. Some described bugs we had long since fixed. Some asked for features that had shipped under a different ticket. Closing the dead ones by hand would have taken days.

So I handed the triage to AI agents. Each agent took a batch of tickets, read the code, and returned a verdict: still open, or obsolete. Forty tickets came back marked obsolete, each with evidence attached.

Before closing them, I stopped and asked one question: how confident am I in these?

My answer was 85 to 90 percent. Four to six wrong, out of forty.

That sounded fine. I asked for one more pass anyway.

This story involves forty agents, but the lesson doesn't need forty. If you've ever asked an AI "is this done?", "is this still used?" or "is this safe to delete?" and acted on the answer, it applies to you.


The Pass That Changed the Answer

The second pass used fresh agents with a different job. They weren't asked to check the verdicts. They were asked to refute them. Each one got a single "obsolete" verdict and one instruction: prove this wrong. If you can't find conclusive proof it's right, mark it not confirmed.

Thirteen of the forty didn't survive.

That's 32 percent. Not four to six wrong, but thirteen. My estimate missed by a factor of two to three, and I'd have stated it just as calmly as the verdicts that were right.

The refute pass cost a fraction of the original triage. It kept about ten real, unfinished pieces of work from disappearing into a closed-ticket graveyard where nobody would look for them again.


The Evidence Was Good. That Was the Problem.

Here's the part that surprised me.

I expected the wrong verdicts to look sloppy: vague reasoning, missing references, guesses dressed up as findings. They didn't. Of the forty verdicts, 35 cited a specific file and line number. 31 cited the commit that supposedly resolved the ticket. And when I went and looked, the cited code almost always said exactly what the agent claimed it said.

The facts were right. The conclusions were wrong.

Every miss was a scope miss. The agent found real evidence that part of the ticket was done and concluded that the ticket was done:

  • A ticket asked to fix a broken form. The agent found the commit that removed the dropdown causing the error. True. The form still couldn't save.
  • A ticket asked for test coverage across a service. The agent found the new tests. True. They covered 4 of 15 methods.
  • A ticket asked for a feature with a specific column the domain expert had requested. The agent found the shipped feature. True. The column wasn't in it.

None of those is a hallucination. The agent didn't invent anything. It answered a smaller question than the one the ticket asked, and it backed that answer with perfect citations.

This is the AI failure I run into most, at every scale. It isn't making things up. It's not going deep enough: stopping at the first real piece of evidence instead of checking the whole question.

And here's what stung: my own spot-check made the same mistake. I clicked through a handful of citations, saw that the code matched, and felt good. I was checking the evidence. I wasn't re-reading the ticket's actual requirements.


The Most Dangerous Claim Is "It's Gone"

One kind of verdict failed more often than the rest: the absence claim. "Nothing calls this anymore." "This code path was removed." "Zero references."

An absence claim is only as good as the search behind it. In one case, the agent searched a directory and its subdirectory, found nothing, and concluded the feature was dead. The live caller was in a third repository it never looked at.

A search that finds nothing looks exactly like a search pointed at the wrong place. You can't tell them apart from the result alone.

A cousin of this showed up a few weeks earlier. An agent searched for a set of fields, found them defined in a data migration, and reported that a parent-child relationship was already wired up. The fields did exist. But they were only populated by one path, an interactive form. The bulk migration, which was the path that mattered, never set them. "The code exists" and "the code is used" are different claims, and a text search only proves the first.


What I Do Differently Now

None of this made me trust agents less for triage. They did most of the work, and most of it was right. What changed is where I put my attention.

Before any bulk action on agent output, a fresh agent tries to refute each verdict. Not review. Refute. The instruction matters. An agent asked "is this right?" looks for reasons to agree. An agent asked "prove this wrong" looks for the gap. And it defaults to not confirmed when the evidence is short of conclusive. The burden of proof sits on closing, not on keeping open.

Re-read the original requirement before looking at the evidence. Line by line. The citation tells you what was done. Only the ticket tells you what was asked. If you read the evidence first, it frames everything after it.

Treat "nothing found" as unproven until the search has found something. Before I believe an absence, I want to see the same search find the thing when it is present. If you can't show the probe works, you haven't shown the thing is missing.

Spread the risky verdicts across agents. If one agent gets all the tickets of a similar shape, its blind spot shows up everywhere in that batch, and it all looks consistent. Mixing them up means one agent's bias can't hide inside one pile.

Don't treat a confidence number as a measurement. My 85–90 percent wasn't careless. It was a feeling, stated as a number. The number that mattered came from the refute pass.


You Don't Need Forty Agents for This

Most developers aren't running a fleet of agents, and none of the habits above need one. A fresh chat session asked to prove the answer wrong does the same job as a refute agent. Re-reading what was actually asked takes a minute. Asking the AI where it looked, and making it find something that is there, takes two.


The Question That Did the Work

It's tempting to tell this story as "the refute pass saved us." It did the work. But the thing that made it happen was a single question, asked before an action that would have been hard to undo: how sure am I?

The answer was a confident 85–90 percent. The real figure was 68.

So here's where I've landed on when to trust AI and when to push back. Trust the citations, check the search, and push back on the conclusion. Push hardest right before an action you can't easily undo.

AI agents are extraordinary at producing evidence. They're much less reliable at deciding whether that evidence answers the question. That judgment is still the job. The good news is that you don't have to do it alone. You can hand it to another agent, as long as you tell that one to disagree.