An LLM judge is a second model that scores an agent's output against criteria you write in plain language. This guide explains where a judge fits beside deterministic tests and human review, then shows how to add one to an Ori Eval test, compare candidate models, calibrate the threshold against human-labeled examples, and keep evaluation cost under control.
Brief · Source report