Judging the Judge: How Do You Know Your AI Evaluation Actually Works?


Most teams building agentic AI systems call API, synthesise an answer and rely on another AI to judge whether the output is any good. It's the standard approach and it's rarely questioned. In a recent Passion Academy session, we looked at a paper examining exactly that assumption: if you're using an LLM to grade an LLM, how do you know your grader is any good?
An agentic system is one where a user query triggers a chain of decisions rather than a single response. Ask "what time is my meeting with Tom?" and the system has to pick a tool (a calendar lookup), fill in the right parameters, execute it and turn the result into an answer. That means every response depends on an orchestrator that selects tools, executes them and synthesises the output.
The paper the session was based on, by Gurung and colleagues, looked at two questions:
It also examined how far an error travels through the system once introduced and whether the system can recover. The session focused mainly on the first question.
The most common way to evaluate a system that produces free text is to use another LLM as the judge. You give it a reference answer and the system's actual output and ask whether they match closely enough. The judge returns a simple true or false.
The paper's central point is that this judge should itself be validated and the way to do that is to check it against humans. Have two or three people independently annotate the same outputs as correct or incorrect and measure how much they agree with each other using Cohen's kappa (higher is better). In the paper's setup, human-to-human agreement came out around 83. When LLM judges were asked to do the same task, they scored significantly lower.
This matters more than it might sound. If you're tuning a pipeline and your evaluation scores stay flat, or get worse, it's tempting to assume the pipeline isn't improving. But it might be the judge that's broken, not the system it's judging. You can never trust your results beyond the accuracy of the instrument measuring them.
The researchers deliberately broke one tool in the pipeline so it always returned a wrong answer, then traced what happened downstream. Two findings stood out:
That second point turns out to be practically useful well beyond this one experiment. Whether a tool was actually called is a cheap, general-purpose signal for catching problems in an agentic system. The researchers built a lightweight runtime interceptor that checked basic facts about each run, whether a tool fired, how the answer was actually produced and used that alone to meaningfully cut down on errors.
A few practical critiques came to light, drawn from applied experience building these systems.
The orchestrator was arguably too naive. It doesn't appear to have been instructed on how to handle tool failure, such as a broken database connection or a suspicious-looking output. In practice, this is a standard part of building such a system: the orchestrator should be able to tell the user honestly that it can't answer right now because a tool is broken.
Honest failure was scored as a wrong answer. This is arguably the sharpest critique. In the paper's rubric, if the system correctly says "I can't answer this because the tool is broken," that gets counted as incorrect. That seems like the wrong call. There's a case for splitting this into two separate metrics: how often the system gives a correct answer (which includes an honest "I don't know" alongside answers that match ground truth), and separately, how often it's able to give a substantive answer at all. The second metric becomes a measure of tool robustness in its own right, distinct from answer accuracy.
Sometimes you don't need a judge at all. During development, it's often possible to force the LLM to output structured fields rather than free text and compare those fields directly against ground truth. That sidesteps the whole judge-validation problem for the cases where it's feasible.
The judge itself deserves more investment. The paper didn't put much effort into improving the judge or aligning it with human judgment. This is an active area of work: building a judge that encodes the actual process and rubric a human annotator would use, rather than a generic true/false comparison.
A few things worth carrying into how you build and evaluate your own agentic systems:
The paper's real contribution isn't a specific accuracy number. It's the validity audit itself: a discipline of asking how you actually know your measurement of an agentic system is sound before you use it to make decisions. That applies whether the "tool" is a simple database call or another agent, or an entire fleet of agents. Tool-call tracking, in particular, stands out as one of the cheaper and more general diagnostics available for tracking down where errors originate in a system that's otherwise a black box.
https://arxiv.org/pdf/2604.16706