Acceptance Criteria for Training Datasets: Deduplication, Evaluation Independence, and Version Traceability

A dataset can load successfully and contain every required field yet still be unsuitable for training. Using a constructed support-ticket classification example, this article explains how to define acceptance boundaries and check exact and near duplicates, temporal and entity leakage, label quality, and version traceability. Acceptance should establish fitness for the agreed task, rather than impose one universal pass rate on every dataset.

A batch of support tickets has been cleaned of missing values and its fields standardized. Is it ready for the training team? If different replies within the same ticket enter the training and test sets, or the input includes fields that become available only after resolution, a model may score well in evaluation yet fail on new tickets. The central argument of this article is that dataset acceptance requires a chain of evidence for a specific task: the samples must be usable, the evaluation independent, and the version reproducible. Complete files are only the starting point.

1. Define the task contract: what the model can see, and when

The constructed example used throughout this article is an enterprise after-sales support-ticket classification task. At initial submission, a ticket is classified as "Delivery inquiries," "Equipment faults," or "Other" to help staff route it. The model receives only the title and description available at submission, while the final category confirmed by a human serves as the target label for supervised learning. All samples, rules, and acceptance targets below illustrate the method. No code was run, no model was trained, and no empirical results were obtained.

Example record A is the initial description of ticket T01: "The equipment will not start." B is a subsequent reply within T01: "It still will not start after a restart." C is a copy of A in another export file. D belongs to a separate ticket, T02: "The equipment will not power on." A and C are candidates for exact duplication; A and B belong to the same business event; A and D may simply be ordinary, similar expressions. They should not all be deleted merely because their wording is similar.

The task contract should also specify the data's time range, intended users, permitted fields, label definitions, expected distribution, and tasks for which the data is unsuitable. Suitability for ticket classification does not automatically make this dataset suitable for answering questions about repair procedures. Datasheets for Datasets advocates documenting composition, collection, processing, uses, and maintenance. Following that approach, this article includes the task description and processing records in the deliverable, rather than delivering only a data file.[1]

2. Check structure and provenance first, and quarantine unsuitable records

Check each file's encoding, parseability, field types, sample ID uniqueness, and permitted label values. Retain a source identifier, ticket ID, submission time, and label revision information for every record, but clearly identify which metadata is used only for governance and will not enter the model input. Ticket IDs themselves should not be treated as predictive features by default.

Non-empty text is not necessarily usable. A description containing only "See attachment" lacks the necessary information if the training workflow does not read attachments. Truncated text may lose a negation, while a description and label drawn from different tickets constitute an incorrect association. Sampling for review should cover source batches, categories, and length ranges, rather than only the beginning of a file.

Provenance and permission to use the data require separate confirmation, and sensitive information should be removed or handled according to the task's needs. In this example, fields such as phone numbers can be excluded while fault-related keywords are retained. Removing equipment model identifiers or error codes by mistake during de-identification changes the information available for the task. Record the rules and review affected samples. Quarantine records whose provenance or permission cannot be confirmed; technical readability is not a substitute for a decision on whether they may be used.[1]

Corrections should produce a new record status or a new version, with the reasons for exclusion preserved. Do not overwrite the only original copy. Quarantining problem records alone may reduce coverage of a particular category, so recalculate the distribution after the structural checks.

3. Deduplicate at three levels without treating similarity as a deletion command

The first level is exact duplication. Preserve the original text, then apply rules fixed in advance to line breaks and insignificant whitespace. Generate fingerprints from the fields relevant to the task and review records with matching fingerprints. Normalization must not arbitrarily remove numbers, negations, or equipment model identifiers. In the example, A and C can be grouped together while retaining a mapping to their sources. If their labels conflict, review the conflict first rather than arbitrarily keeping one record.

The second level is near duplication. Character-fragment overlap, MinHash, or similar methods can generate candidates for human review to determine whether they result from copying, rewriting, or repeated exports. Research on deduplicating language-model data uses checks for near duplicates and repeated substrings, and discusses overlap between training and validation data.[2] That research primarily concerns language models trained on web text. Its reported gains and matching thresholds cannot be applied directly to Chinese support-ticket classification.

The third level is event linkage. A and B need not contain duplicate wording, but both belong to ticket T01. If the task uses only the initial submission, B should be excluded from its input samples. If a separate task uses multiple turns within a ticket, the records belonging to the same ticket should be handled as a related group. Text deduplication alone cannot establish this business boundary.

Calibrate similarity thresholds using candidate pairs judged by humans, and examine false merges and missed duplicates separately. A short phrase such as "will not start" may occur in many independent events; excessive deletion changes its real frequency. Whether to retain such repeated expressions within the training set depends on the task distribution and the sampling objective. When checking similar samples across sets, distinguish shared templates from leaked answers. Not every common phrase constitutes leakage.

4. Design independent splits around the generalization you want to evaluate

Establish duplicate and event relationships before assigning records to sets. Any processing step that learns parameters from the data should then be fitted after splitting, using only the training set. The training set is for learning, the validation set for tuning and selecting an approach, and the test set for evaluation after that approach has been frozen. Limited sample size is not a reason to combine these responsibilities.

If the objective is to predict newly submitted tickets in the future, splitting by submission time is appropriate, with related events that span the boundary handled accordingly. If the objective also includes serving previously unseen customers, perform an additional check using customer groups. Temporal separation and customer separation answer different questions. Two evaluation slices can be established; neither one demonstrates the generalization measured by the other.

The scikit-learn documentation explains that when related samples are grouped, samples from the same group should not appear in both the corresponding training and test sets. Random cross-validation should not simply be applied to time-series data either.[3] At a minimum, this example requires that a ticket never spans different sets. If customers repeatedly report problems with the same equipment, assess whether equipment or fault incidents should define higher-level related groups.

Investigate leakage by checking each input field along the timeline. Resolution outcomes, the department ultimately responsible, and time to resolution are unavailable at initial submission and should be excluded from the input. The final category can be a label, but cannot be a feature. Next, check whether cleaning, vocabulary construction, normalization, feature selection, or other processes used test data to learn parameters. The official documentation explicitly states that this kind of leakage produces overly optimistic evaluations.[4]

Once the test set is frozen, do not repeatedly change prompts, rules, or models based on its results. If the team has been continually optimizing around test errors, treat that set as development material and establish a new, independent test set. For pretrained models with opaque data provenance, even a clean local split cannot prove that the underlying model has never encountered relevant public material. State this limitation in the delivery documentation.

5. Define denominators, handling rules, and assumptions for every acceptance metric

The first recommended group of checks covers structure and authorization: the number of parseable records divided by the number submitted; the number with complete required inputs divided by the number of candidate samples; and the number whose provenance and intended use can be confirmed divided by the number proposed for inclusion. In this example, every record formally included must be parseable, have usable required inputs, and have a confirmed basis for use. These requirements follow from the training workflow and delivery responsibilities. The original batch may contain bad records, but the number quarantined must be reported. Do not present the pass rate after quarantine as the quality of the original batch.

The second group checks separation: the number of exact-duplicate samples spanning sets, the number of ticket groups spanning sets, and the number of confirmed near duplicates spanning sets. This example makes zero overlap in the first two measures a release condition because the objective is to evaluate new tickets; export copies or overlap within the same event undermine that objective. Each near-duplicate candidate requires handling or a documented, justified reason for retention. These zero counts address only the detectable problems that have been defined. They do not guarantee the absence of unknown leakage.

The third group checks labels and coverage: the number and proportion of labels that disagree with review conclusions in a stratified sample, and sample counts for each category, time period, and major source. If two reviewers disagree on the boundary of "Equipment faults," resolve the disagreement and revise the guidelines before checking the related samples. Agreement between reviewers does not prove that labels are absolutely correct; review still needs sufficient business evidence.

Agree on tolerance for label errors in advance, based on the consequences of misclassification and the cost of correction, rather than prescribing one universal percentage. When a critical category has few samples, a single accuracy figure is unreliable. Report results by category and the corresponding sample sizes, and collect additional data where necessary. Increasing the sample size can reduce some statistical uncertainty, but cannot correct systematic labeling bias.

The fourth group checks traceability: whether every delivered shard has a checksum, whether the same data and split can be restored, and whether exclusions have documented reasons. Passing dataset acceptance means only that the dataset meets the agreed input and evaluation requirements. Whether the model achieves the business objective still requires subsequent training and independent evaluation; field completeness cannot establish that outcome.

6. Deliver a recoverable version, not just a filename

The delivery manifest should include the raw snapshot identifier, cleaning rules and program version, label guideline version, random seed or deterministic split rules, sample-to-set mapping, exclusion list, acceptance report, and content checksums for every shard. In the example, version v1 can record the consolidation of A and C and the reason for excluding B. A later revision to A's label should produce v2, with the affected splits and evaluations documented.

DVC's .dvc files record information such as data paths and checksums, and can be tracked with Git.[5] DVC is optional; teams can achieve similar objectives with controlled object storage and manifests. Checksums help determine whether content has changed. They cannot prove that labels are correct, permissions are sufficient, or data quality is acceptable. A manifest alone, without the retained data bytes, cannot restore a version either.

Handle common failures according to their impact. If an event spans sets, split again by group and rerun the evaluation. If a field is unavailable at prediction time, remove it and retrain; the old results must no longer serve as evidence of capability. If the label guidelines change, review affected samples and release a new version. If the old data cannot be restored, mark that version's results as non-reproducible rather than continuing delivery under a different filename.

Acceptance scope must change when the data's scale or use changes. Multi-turn conversations require checks for relationships within a conversation; generative tasks require review of inputs and target answers; continuously updated data requires a locked evaluation snapshot. The ticket rules in this article do not directly cover those tasks, but the common principle remains: every release decision should explain what was checked, how it was checked, how failures are handled, and what has not yet been established.

References

  1. [1] Timnit Gebru et al.: Datasheets for Datasets
  2. [2] Katherine Lee et al.: Deduplicating Training Data Makes Language Models Better
  3. [3] scikit-learn: Cross-validation: evaluating estimator performance
  4. [4] scikit-learn: Common pitfalls and recommended practices (Data leakage)
  5. [5] DVC: .dvc Files
Back to insights
鲁ICP备2024109755号-2
Drag to move. Right-click, touch and hold, or press Shift+F10 to choose a corner.