Enterprises discussing synthetic training data often start with volume and price per record. Better opening questions are: on which real inputs does the existing model fail, where will a generator obtain the knowledge needed to correct that failure, and what independent evidence will demonstrate improvement? This article argues for treating synthetic data as an engineering investment in a defined gap. Increasing the number of samples is neither evidence of stronger capability nor a sufficient procurement acceptance criterion.
The debate is shifting toward training workflows
Synthetic training data here means samples constructed through rules, simulators, or generative models for a model to learn from. This article focuses on expanding datasets for enterprise text and structured-data tasks; its conclusions do not automatically extend to all visual simulation or foundation-model pretraining. For data teams, model owners, and procurement decision-makers, the practical shift is that generation can be scaled, while the bottleneck moves toward adding useful information, controlling the distribution encountered during training, and demonstrating operational value.
A 2024 Nature paper demonstrated degradation in recursive training on model-generated data, including the gradual loss of less common content from the original distribution. Its findings concern particular recursive processes and accumulating errors, not the claim that every one-off fine-tuning exercise with synthetic data must fail.[1] The ICML 2025 paper Collapse or Thrive compared replacement, accumulation, and fixed-size subsampling workflows. Retaining real data and training on accumulated data produced different outcomes from repeatedly replacing earlier data with synthetic generations; a fixed training budget also changed the degradation observed.[2]
A preprint released on 16 September 2026, using a Fisher-Rao perspective, further studies mixtures of real and synthetic data. Its stability conditions rely on assumptions including finite categorical distributions and bounded perturbations.[3] This is theoretical evidence, not a universal recipe for enterprise fine-tuning. Our engineering inference is to shift the decision from a single synthetic-data percentage toward the real information retained, its actual participation in training, and the external evidence constraining the process.
Identify whether the gap concerns expression, coverage, or business facts
The first type of gap concerns expression. Business rules are already clear, but users describe the same issue with abbreviations, colloquial language, misspellings, or different word orders. Variants of verified source examples may broaden coverage of input forms, provided they preserve the facts that determine the answer. Changing “returns are not supported” to “returns are supported” changes label semantics; it cannot pass as ordinary paraphrasing.
The second type concerns combinations of known rules. A purchasing request, for example, may require checks of amount, department, and approval status. Boundary combinations can be constructed from rules, labeled with deterministic logic, and then rendered in natural language by a model. The information comes from those rules, not from the generator independently knowing the business. Ambiguous rules need clarification from the business owner first.
The third type concerns real mechanisms that have not been observed: actual failures of new equipment, the behavior of an unfamiliar customer group, or an unconfirmed cause of an anomaly. Generating many plausible cases does not supply the missing factual basis. Collection, annotation, or controlled trials should take priority. Synthetic samples may help formulate hypotheses or exercise interfaces, but they cannot demonstrate that these situations have been validated in reality.
A constructed scenario: supplementing purchasing-request triage
The following is a constructed example, not an IDENIFE project, and no training experiment was run. Suppose an enterprise classifies purchasing requests into three actions: “complete information; proceed to approval,” “additional information required,” and “outside the current workflow; refer to a person.” Historical examples are dominated by routine requests. Development-set analysis shows that the model often misses absent units and requests involving multiple currencies without an exchange-rate date.
Start with a task matrix covering request type, amount units, currency relationships, and required-field completeness. Domain staff confirm the appropriate action for each combination and remove combinations that cannot occur in the business. The generator then turns valid combinations into varied wording, retaining the rule identifier and expected action. Do not fill the dataset with impossible combinations merely to increase the proportion of difficult cases.
Cross-check structured fields against the narrative within each case. If a field says that the exchange-rate date is missing while the narrative gives a specific date, return the record for correction. Also distinguish provenance: real events, rewrites of real events, and rule-constructed cases should be recorded separately. An unverified generated answer does not become a business fact simply because another model agrees with it.
The methods in our article on training dataset acceptance can support checks of separation and traceability. The further question here concerns investment: even when sample format, labels, and versions are acceptable, do these samples address the intended gap, or merely help the model imitate the generator's preferred language?
Keeping real data in storage is not enough
A purchased or internally built solution should explain how real data participates in every training round. A stored file does not establish that a model actually sees its contents. With fixed-size random sampling, an expanding synthetic pool may reduce the chance that rare real cases are selected. Training on all accumulated data instead increases training and governance costs. The accumulation and fixed-budget strategies studied in the research are not interchangeable.[2]
We recommend tracking three quantities: the number of distinct available cases from each source, actual sampling counts by source during training, and exposure for critical business slices. One case rewritten in a hundred ways should not be reported as a hundred independent business observations. Where sample lengths differ, also record training tokens contributed by each source, so record-count ratios do not conceal actual weight.
This does not mean retaining every historical record in training indefinitely. Outdated rules, incorrect labels, and data no longer permitted for use need to leave the relevant training scope, with governance records retained. Decisions about which real samples remain should reflect the current task and permitted uses. Sampling allocations should be compared on development data and then independently validated, rather than copied directly from a paper.
Accept business improvement, not generation pass rates
Checks during generation can reduce obvious errors. NVIDIA NeMo Data Designer's Validators documentation describes structured results from validation of target columns, with code checks, custom functions, and remote validation among the available methods.[4] This demonstrates that checks can be embedded in production. Passing field validation, syntax checks, or scoring does not establish that the subsequently trained model meets the business objective.
For the constructed scenario, we recommend three comparison groups using the same base model and explicit budgets: existing real training data only; the same data supplemented with gap-focused synthetic examples; and a comparable allocation of human effort to annotating additional real cases. Fix evaluation rules during development, and record training steps, tokens, generation charges, and review hours. If the second group uses more compute, the outcome describes the entire intervention; the difference cannot all be attributed to data provenance.
Final acceptance should use held-out real cases that were not involved in generation seeds, prompt adjustments, or model selection. A synthetic challenge set can test declared boundaries, but cannot alone demonstrate generalization to actual operations. When rare real cases are deliberately oversampled, document that sampling and separately report results for the overall distribution and priority slices. The increased share of difficult cases in an evaluation set is not their rate of occurrence in production.
Connect metrics to the consequences of actions. For example, the number incorrectly passed to approval divided by all cases that actually required more information or human handling measures missed interventions. The number incorrectly asked for more information divided by all genuinely complete cases measures unnecessary burden. Report sample sizes and uncertainty for each slice. Agree tolerances beforehand based on error costs and human handling capacity; this article offers no universal threshold masquerading as an industry standard.
Count the full cost and define when to stop scaling
The full cost of a synthetic-data approach includes seed preparation, rule clarification, generation, filtering, human review, training, evaluation, and correcting errors after deployment. A price per thousand generated records covers only part of that work. Compare the total cost of meeting the same acceptance objective, or compare critical error rates under the same overall budget. If neither approach meets the objective, low cost does not make either acceptable.
In the constructed scenario, if expression variants improve routine triage but do not reduce missed interventions on cross-currency requests, revisit rules and input evidence instead of adding more wording. If real-data slices deteriorate while synthetic-set scores rise, pause expansion and examine generation style, sampling weights, and labeling bias. Where expert review dominates cost, directly annotating a smaller set of real cases may be more economical.
An expansion decision needs at least four understandable pieces of evidence: the specific failing slice and its business cost; the source of information introduced by synthetic samples; retention and sampling records for real data in training; and independent acceptance results with a cost comparison against alternatives. Scale has a justification only when this evidence supports further investment. What an enterprise needs to buy or build is a process for verifiable task improvement.
References
- [1] Shumailov et al. — AI models collapse when trained on recursively generated data (Nature, 2024)
- [2] Kazdan et al. — Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating World (ICML, 2025)
- [3] Marchi et al. — Preventing Model Collapse: A Fisher-Rao Perspective on the Dynamics of Training with Synthetic Data (arXiv v1, 16 September 2026)
- [4] NVIDIA NeMo Data Designer — Validators (Latest documentation, accessed 8 October 2026)
