What is an LLM judge?
An LLM judge is a prompt that grades an agent’s output by reasoning over it against one or more stated intents, the things the agent was supposed to do. Instead of matching a pattern, it reads the output, works through whether each intent was met, and returns a verdict, which is why it fits properties that need judgment: whether an answer is grounded in what was retrieved, whether it follows a policy, whether its tone fits the situation.
Its verdicts inherit the properties of the model running it. Run the same judge twice on the same input and it can disagree with itself, and swapping the underlying model changes what its verdicts mean even with the prompt held fixed. It’s also the most expensive grader per verdict, since it pays for an inference call every time it runs, which is why most teams point it at a fixed eval set before a change ships and a sampled or flagged slice of production traffic, rather than every output.