Human feedback needs a disagreement policy.
The useful signal is often in the reasons people disagree. A review process needs a way to preserve and resolve those differences.
Human feedback is often described as if it were a single input: collect ratings, combine them and use the result. But a rating can encode a factual judgment, an interpretation of instructions or a personal preference. These are different kinds of evidence. Our view is that a team should decide how it will handle disagreement before it begins collecting feedback, because the resolution policy shapes what the resulting data means.
Define what the feedback is meant to represent
Ouyang and colleagues demonstrated a process using human demonstrations and ranked model outputs to improve instruction-following behavior in the systems they studied. That result helped establish the practical importance of human feedback. It does not mean that every collection of preferences represents all users, or that preference alone proves factual correctness. A feedback process needs an explicit account of the task and the judgments it is collecting. [1]
Consider a hypothetical task comparing two explanations of a software feature. One explanation is concise; the other includes more detail. A reviewer preparing expert documentation may prefer the first. A reviewer helping a first-time user may prefer the second. Both judgments can be reasonable if the intended audience was never specified. Combining the votes produces a result, but it does not repair the missing task definition.
Begin by stating whose experience matters for the task, what context reviewers should assume and which dimensions they should assess. A reviewer may be qualified to judge whether an explanation is understandable without being qualified to verify a specialized technical claim. Separating those responsibilities makes the feedback easier to interpret. It also gives reviewers a legitimate way to say that a judgment is outside the information or expertise available to them.
Distinguish three reasons for disagreement
We propose a simple working distinction: an evidence error, an instruction ambiguity or a legitimate preference difference. An evidence error occurs when a judgment conflicts with the source material the task requires. Instruction ambiguity occurs when reasonable reviewers apply different interpretations of the same rule. Preference difference occurs when the task permits several acceptable outcomes and people value them differently. A real case may involve more than one category.
Each category suggests a different response. Evidence errors call for checking the source and correcting the judgment. Ambiguous instructions call for revising the rule and identifying records that may have been affected. Legitimate preference differences call for preserving the variation or explicitly choosing an audience-specific policy. Treating every disagreement as reviewer error can erase useful information; treating every disagreement as equally valid can preserve mistakes.
In the software-explanation example, an incorrect feature description should be checked against the relevant specification. A dispute over whether setup steps belong in the answer may reveal an incomplete rubric. A preference for one clear, accurate explanation over another may reasonably remain unresolved. The review interface should make these outcomes available instead of forcing every situation into a winner and a loser.
Make the reasons reviewable
A useful feedback record can contain the judgment, the criterion, a short reason and the evidence needed to inspect it. Ask for enough explanation to distinguish a grounded judgment from a guess. This does not require a long essay for every routine decision. A reference to an incorrect sentence, an absent prerequisite or a confusing instruction can be more useful than a paragraph of general impressions.
Keep original reviews separate from the final adjudicated decision. An adjudicator may choose an outcome to make a dataset usable, but that decision should not overwrite the fact that independent reviewers differed. Record the reason for resolution and whether the instructions changed. If the same ambiguity appears repeatedly, the appropriate improvement may be in task design rather than in individual reviewer performance.
Calibration should follow the same principle. Reviewers can assess a shared set independently, compare their reasoning and practice applying revised rules. Keep calibration examples separate from the examples used to measure independent agreement. Otherwise, familiarity with discussed cases can make the process appear more consistent than it is on unfamiliar work. The objective is transferable understanding of the task.
Be specific about representation
Santurkar and colleagues compared opinions expressed by language models with survey responses from US demographic groups and found substantial differences in the systems they studied. Their work illustrates why broad statements about what humans prefer need a defined population and measurement method. It does not establish how every current model behaves, nor does it imply that any single demographic group has one uniform opinion. [2]
For a practical review program, document the qualifications relevant to the task and the limits of the reviewer sample. Language fluency, domain knowledge and familiarity with the intended context may all affect what a judgment can support. Avoid turning a convenient sample into a claim about all users. When a task is audience-specific, evaluate that audience deliberately and report what was actually represented.
Representation also depends on the situations presented for review. Even a varied reviewer group cannot comment on cases it never sees. In the explanation example, a collection limited to straightforward questions says little about responses to incomplete requests or unfamiliar terminology. Review the sampling plan alongside reviewer qualifications. Both determine the scope of the evidence, and both should be revisited when the product's intended use expands.
Turn recurring disagreement into a design decision
At the end of a feedback cycle, inspect repeated reasons for disagreement. Some may point to missing reference material. Others may reveal that two criteria conflict, such as a preference for brevity and a requirement to explain every step. A team then has a product decision to make: offer different response modes, prioritize one criterion for a defined audience or gather more evidence. Additional votes alone will not settle an undefined objective.
Make changes to the review policy explicit. If the target audience changes, a previous label may no longer express the desired behavior. If a source is corrected, identify the judgments that depended on it. If the rubric changes, preserve the earlier version so comparisons across batches remain interpretable. These records help downstream teams understand why a dataset changed and which conclusions still hold.
The value of human feedback lies partly in what people can notice and explain. A good process gives those observations a route into clearer instructions, better examples and more thoughtful product choices. Agreement is useful when people are applying a well-defined rule consistently. Disagreement is useful when it exposes a question the team has not yet answered. Designing for both produces feedback that is easier to trust, inspect and improve.
Sources & further reading
Analysis from ADCO AI Labs, based on the research cited below. Examples are illustrative.
- Training language models to follow instructions with human feedbackLong Ouyang et al. · 2022
- Whose Opinions Do Language Models Reflect?Shibani Santurkar et al. · 2023