ADCO AI
All insights

Evaluation starts with the decision.

A score becomes useful when it is tied to a specific workflow, a clear acceptance rule and the evidence behind each judgment.

Before choosing a benchmark, ask what decision the result needs to support. A team comparing two models needs different evidence from a team deciding whether an assistant can handle a new workflow. A team investigating a failure needs something different again. Our position is that evaluation should begin with this decision, then work backward to the scenarios, criteria and review process capable of informing it.

Name the decision before choosing the metric

Imagine a hypothetical support assistant that drafts responses for a person to review. The team is deciding whether to expand from straightforward product questions to more ambiguous requests. A single measure of response quality leaves too much unspecified. The reviewer also needs to know whether the draft uses the available evidence, identifies missing information and stays within the task. These criteria describe the proposed use; they should be agreed before the team examines results.

Write the decision in a form that can be revisited: the system version, the workflow, the proposed change and the evidence needed to support it. Specify which failures prevent expansion, which require a narrower scope and which are tolerable within the review process. There is no universal threshold that turns an assistant into a dependable system. Acceptance depends on the consequences of the errors and the controls surrounding the task.

Look across dimensions and situations

The HELM project introduced a broad evaluation framework spanning scenarios and multiple dimensions, including accuracy, calibration, robustness and efficiency. Its research makes tradeoffs visible instead of treating accuracy as the complete description of a model. The particular scenarios and models in that study are historical; the useful design principle is to state what an evaluation covers and where coverage is missing. [1]

For the support-assistant example, a team could report ordinary requests separately from ambiguous, incomplete and conflicting requests. It could also distinguish response correctness from the ability to recognize that the available evidence is insufficient. This prevents strong performance on a frequent, easy category from obscuring poor behavior on a smaller category that determines whether the proposed expansion makes sense.

Include examples where asking a question or escalating is the desired result. An assistant that always produces a polished answer may appear helpful while ignoring an important boundary. Conversely, a system that refuses almost everything can avoid many errors while failing to serve users. Evaluation should make that tradeoff inspectable. The desired balance needs to be expressed through cases and acceptance rules, then reviewed by the people responsible for the workflow.

Make the rubric observable

A criterion such as helpfulness is a starting point, but reviewers need to know what they can observe. In the hypothetical workflow, a rubric might ask whether the response answers the actual question, whether each material claim is supported by the supplied reference and whether a missing prerequisite is identified. Pair each criterion with examples that show a clear success, a clear failure and a genuinely ambiguous case.

Keep the reference material and the instruction version with the judgment. If two people disagree about factual support, they should be able to compare the same source passage. If the task permits several acceptable responses, the rubric should say so. A single reference answer should not quietly become the only permitted wording when the intended task allows equivalent answers.

Before reviewing a large batch, ask reviewers to assess a small shared set independently. Discuss the reasons behind disagreements before changing the scores. A consensus reached through discussion can help improve the rubric, but it should not be reported as independent agreement. Keeping the original judgments and the adjudicated result allows the team to distinguish reviewer consistency from the effect of the resolution process.

Evaluate the evaluator

Zheng and colleagues investigated language models used as judges and documented limitations including position, verbosity and self-enhancement biases. They also found strong agreement with human preferences in the settings they tested. Together, these findings support a conditional approach: an automated judge can be useful, but agreement demonstrated on one benchmark should not be assumed for every new task, rubric or model. [2]

Our recommendation is to compare an automated judge with independent human review on a relevant sample before relying on its scores. Inspect disagreements by criterion, not just in aggregate. For pairwise comparisons, swap the order of candidate answers and investigate unstable results. Where concision matters, check whether a longer answer receives credit simply for sounding more complete. These checks should be repeated when the judge, its instructions or the evaluated task changes.

Some judgments can be checked directly: a required field is missing, a reference does not exist or a calculation does not match an expected value. Other judgments need contextual interpretation. Use the available form of evidence for each criterion and record which evaluator produced the result. Combining deterministic checks, automated review and human judgment is most understandable when their responsibilities remain explicit.

Preserve the comparison, then keep learning

A comparison should make it possible to understand what changed between runs. Keep the evaluation set, model identifier, instructions, retrieval context and scoring procedure connected to the results. If the system changes several components together, avoid attributing the outcome to one component without additional evidence. Where outputs vary, report the repeated-run procedure and the amount of variation observed rather than selecting a convenient response.

After reviewing a failure, add its general pattern to a separate collection of candidate tests. Avoid treating a fix for a known example as proof that the broader class of problems has disappeared. Maintain a stable set for regression comparisons and refresh a separate set to examine new behavior. This gives the team both continuity and a way to notice gaps that the original evaluation never anticipated.

The final report should return to the decision that started the work. State what the evidence supports, what remains uncertain and which next action follows: proceed within the defined scope, revise a control, gather more examples or postpone expansion. An evaluation becomes operationally useful when someone can read it and understand why a decision was made. The score is one part of that explanation; the cases, criteria and unresolved failures complete it.

Sources & further reading

Analysis from ADCO AI Labs, based on the research cited below. Examples are illustrative.

  1. Holistic Evaluation of Language ModelsPercy Liang et al. · 2023
  2. Judging LLM-as-a-Judge with MT-Bench and Chatbot ArenaLianmin Zheng et al. · 2023