Start with observable criteria
Terms such as good, relevant, or accurate are too broad by themselves. A usable rubric defines what the reviewer can observe, provides contrasting examples, and makes the correct action clear when evidence is incomplete.
Calibrate before volume
Reviewers should score the same controlled set, compare disagreements, and revise unclear instructions before production. This separates a training issue from a rubric issue and prevents inconsistency from scaling.
Review the review
Quality assurance should sample work, inspect high-risk categories, and track corrections without treating every error as identical. The goal is to understand why a decision failed and what control should change.
Keep an exception log
New edge cases are valuable operating data. Logging them creates better examples, exposes recurring ambiguity, and gives the client a visible record of how the workflow is learning.


