If two qualified humans disagree on half the cases, an LLM judge cannot rescue the rubric.
Begin with a small calibration set containing clear passes, clear failures, and borderline examples. Ask reviewers to label independently and explain decisions. Discuss disagreements, rewrite vague criteria, and add examples until the team shares an operational definition of quality.
Then test the LLM judge against resolved labels. Measure false passes and false failures by category, not only overall agreement. Repeat difficult cases to understand variance and test pairwise comparisons in both answer orders.
Version the rubric, examples, judge prompt, and model. Recalibrate after material changes or when production exposes a blind spot.
Human review is not perfect ground truth; it is a process for creating accountable judgment. A trustworthy automated evaluator is one whose agreement, biases, and permitted decisions have been measured against that process.