A useful dataset is a maintained system.
The value of a dataset depends on what a team can explain, reproduce and improve after the first delivery.
A dataset can arrive in a perfectly formatted file and still be difficult to use. The columns may be consistent while the examples miss important situations. The labels may be complete while the instructions behind them are unclear. For teams building AI, the useful question is what decisions the data can support, and what evidence makes that answer credible. Our view is that dataset quality belongs in the operating model from the beginning.
Start with the use the data must support
Consider a hypothetical assistant that helps employees find an internal procedure. A collection of clean, well-written questions about common procedures may make a useful starting point. But its usefulness changes if the assistant also has to handle old documents, ambiguous requests, missing permissions or questions with no documented answer. Those situations call for different evidence and different expected behavior. The intended use should therefore become a written sampling plan before collection begins.
That plan can describe the task, the population of inputs, the contexts that matter and the situations deliberately excluded. For the hypothetical assistant, a team might distinguish current from superseded documents, straightforward lookups from conflicting sources, and answerable from unanswerable requests. These are proposed design choices, not a universal taxonomy. Their purpose is to make coverage discussable. A larger collection is useful only if the additional examples address a meaningful need.
Keep the context attached to the records
In Datasheets for Datasets, Gebru and colleagues propose documenting why a dataset exists, what it contains, how it was collected and which uses are recommended. The underlying contribution is a structured way to communicate information that a data file alone cannot carry. A datasheet describes a resource; it does not independently certify that the resource is suitable for a particular deployment. [1]
Our practical extension is to keep that documentation close to the work. A reviewer should be able to find the relevant instruction version, source reference and interpretation rule without reconstructing a conversation. A downstream team should know whether a missing value means unavailable, not applicable or not yet reviewed. If a record changes, the reason should travel with it. This makes an apparently small correction understandable months later, when the original contributor may no longer be available.
The same principle applies to exclusions. A team that removes difficult examples can produce a cleaner file while narrowing the problem it actually covers. Keep an exclusion reason and inspect the pattern across removals. An excluded language, format or source type may reveal a deliberate boundary; it may also reveal a gap that requires another collection method. The point is to make the boundary visible before someone mistakes it for complete coverage.
Inspect repetition before counting scale
Lee and colleagues studied duplicated text in language-model training datasets and showed that deduplication could reduce memorized output and overlap between training and evaluation material in their experiments. Separately, Kandpal and colleagues found that duplication increased susceptibility to the privacy attacks they studied. These results support inspecting repetition. They do not establish a single deduplication rule or guarantee privacy for every dataset. [2] [3]
For a new collection, distinguish identical records from examples that express the same content in different forms. A copied paragraph, a translated version and a genuinely independent example may require different handling. Preserve enough information to understand what was grouped and why. If examples share a document, authoring template or interaction, consider that relationship when separating development and evaluation data. Otherwise, a test can become easier because it resembles material the team has already used.
Deduplication also involves a coverage decision. Repetition may be accidental, or it may represent a frequent real-world pattern. Removing every repeated form without inspecting its role can alter the distribution the collection is meant to represent. A useful review compares the collection before and after filtering, records which groups changed and checks whether rare but important cases disappeared. The objective is a defensible composition, with a reason for the examples retained.
Review the rules as well as the labels
When two reviewers disagree, changing one label may resolve the record but leave the underlying problem intact. In our hypothetical procedure-search task, one reviewer may judge an answer against the newest source while another accepts any approved-looking document. The disagreement points to an instruction gap. Correcting the record is necessary; clarifying how document versions should be handled prevents the same conflict from recurring elsewhere.
A practical quality report can separate source problems, instruction problems and execution errors. Each category suggests a different response: recollect material, revise a decision rule or improve reviewer preparation. Sample some accepted records as well as obvious exceptions, because a process that examines only rejected work cannot tell much about what it routinely approves. Where uncertainty remains, preserve that uncertainty instead of silently converting it into a confident label.
Give every release a future owner
A delivery should answer four practical questions: what changed, why it changed, what was checked and who handles the next issue. An immutable release identifier connects a model run or evaluation to the exact data used. A short change log explains additions, removals, corrected labels and instruction revisions. A known contact gives downstream users somewhere to report a record that no longer behaves as expected.
Maintenance does not require changing everything continuously. Establish a stable release for reproducible comparisons, then maintain a separate queue of proposed corrections and new cases. Decide when those changes justify another release and which checks must be repeated. If a source disappears or a procedure changes, the team can assess the affected records rather than reopening the entire collection without direction.
This is where a dataset becomes a useful organizational resource: a new person can understand its purpose, trace its decisions and contribute an improvement without guessing. The first delivery demonstrates that data exists. The maintenance process makes it possible to keep using that data with a clear understanding of its boundaries. For teams planning their next collection, that continuing usefulness is a stronger design target than a record count alone.
Sources & further reading
Analysis from ADCO AI Labs, based on the research cited below. Examples are illustrative.
- Datasheets for DatasetsTimnit Gebru et al. · 2021
- Deduplicating Training Data Makes Language Models BetterKatherine Lee et al. · 2022
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace and Colin Raffel · 2022