What's the difference between an LLM judge and a reward model?
An LLM judge is a general model, prompted at grading time with the criteria you want checked; a reward model is a separate model trained specifically to output one preference score, with the criteria baked into its training data rather than a prompt.
Change what an LLM judge should check and you edit the prompt; change what a reward model should check and you retrain it on new preference data, which is slower and needs a labeled dataset to start. A reward model also returns a single number with no explanation, while an LLM judge can be asked for a verdict plus the reasoning behind it, so its calls can be read and audited one at a time rather than trusted as a black box. Reward models are the common choice inside training, like RLHF, where the same scoring policy has to run millions of times cheaply; an LLM judge is the more common choice for evaluation, where criteria change often. A judge’s per-event cost is why it usually runs on a sample rather than every trace, a tradeoff a reward model’s cheap, fixed inference doesn’t face the same way.