Validating LLM Structured Outputs Before Database Entry: Four Layers Beyond JSON Schema

Valid JSON can still turn unknown values into zeros, confuse monetary units or copy a field from the wrong document. Using a constructed quotation-extraction task, this article distinguishes response completion, structural constraints, business consistency and source evidence. It develops exception routing, controlled repair and acceptance metrics for traceable, reviewable candidate records.

Asking a model to extract fixed fields from contracts, quotations or tickets can reduce manual preparation. Yet a database readily accepts records whose fields and types are correct while their content is false. The central principle here is to treat structured output as candidate data and let independent validation and business rules authorize entry. This article concerns document extraction, not models directly making payments or approvals. Examples, thresholds and statistics are illustrative; no experiment was run.

Layer one: confirm completion before parsing data

Check whether the request completed normally, was refused or was cut off by an output limit before parsing the result. HTTP success is not extraction success. Anthropic’s structured-output documentation explicitly notes that a refusal can return HTTP 200 and take precedence over schema constraints; it also lists output-token-limit exceptions.[1] Response fields are interface-specific, so implement branches for the actual platform.

Save response status, task identifier, source-file revision and raw return, then route to validation, bounded retry or human handling. Partial streamed content may already resemble a complete object, but should not enter the authoritative business table before completion. Waiting prevents a partial record from becoming the final result.

Where raw output contains sensitive enterprise information, define retention and access by purpose rather than keeping everything indefinitely for debugging. Operational logs can retain locating identifiers and error codes while full materials stay in a controlled area. Traceability and data minimization need to be designed together.

Layer two: define structure and preserve unknowns

JSON Schema specifies data structures and constraints. Its documentation explains that properties does not automatically make fields mandatory, additional fields may be allowed by default, and absence differs from null.[2] Decide separately whether a field must exist, whether it may be null, its type and allowed values, and how to handle unexpected fields.

For a quotation, specify document identifier, supplier, currency, amount, unit, tax status and evidence location. Retain the amount, source unit and normalized result separately. A supplier name is not a verified master-data match: the model must not invent an official supplier ID. Controlled rules or people should confirm the match.

“Freight not specified” is not zero freight, and an unreadable tax rate must not become a customary rate. Define states such as extracted, not provided, conflicting and unreadable, allowing the corresponding value to be null. Downstream code can then distinguish no information from an actual zero. Business owners should define a bounded state set rather than proliferating free-text labels.

Model-side constraints support the schema subset that the service actually implements. Server-side validators also have versions and settings. Pin the dialect and validation-library version and retain boundary examples. Do not assume a parameter named strict has the same meaning across products.

Layer three: check business consistency with deterministic rules

After type checks, validate monetary units, date relationships, duplicate items and detail totals. When quantity and unit price are explicit, calculate the line amount in code and compare it with the stated subtotal. Mark uncertainty when discounts, taxes or rounding are unspecified rather than forcing totals to balance. Arithmetic consistency establishes compatibility among fields, not that they came from the correct document.

Validators can also coerce data implicitly. Pydantic documents automatic conversion in its default mode and tighter conversion in strict mode, while JSON inputs such as dates may follow different rules.[3] Test the real input path after enabling strictness; a test using Python objects does not establish identical JSON-interface behavior.

Model-level validators can implement cross-field relationships; Pydantic provides custom validation at both field and model levels.[4] Record source extraction, rule-based normalization and business verification separately. Retain normalization-rule revisions instead of overwriting the source with a tidy-looking number.

For rounding-sensitive amounts, use decimal or minor-currency-unit representations appropriate to the business, with explicit currency and rounding rules. Tolerance should follow the agreed calculation method, not be relaxed merely to pass more examples. Cross-currency comparisons also need an exchange-rate source and time; the language model must not supply a plausible-looking rate.

Layer four: validate support from the source

Candidate records should carry file identifier, revision, page or stable location, quoted source text and extraction method. First verify that the location and text exist, then assess whether the text supports the field’s meaning. “Budget ceiling of 100,000” may genuinely appear, yet cannot establish “quoted price of 100,000.”

Optical character recognition can itself be wrong. Critical numbers need a route back to the original image, especially decimal points, minus signs, units and tables spanning pages. A second model reading the same incorrect OCR text does not automatically provide independent evidence. Give a reviewer information capable of exposing the initial error or have a person inspect the original.

The system can select evidence locations from known chunks or annotations rather than relying entirely on model-generated page numbers. Returned quotations still need checking, and document updates must not leave old locations pointing into new content. Keeping a semantically unverified clause in a candidate or review state is correct system behavior.

Constructed example: even a correct amount may not be ready for entry

Suppose a quotation states “equipment price: RMB 120,000, tax included; freight to be agreed,” expressed in the Chinese source as 12 units of ten thousand yuan. No tax rate is supplied. The task is to prepare a procurement-system candidate. Explicit unit conversion supports RMB 120,000; tax rate remains not provided, and freight remains pending. Neither should be guessed as 13% or zero. All details are constructed.

If the model returns amount 12, currency RMB, tax rate 13% and freight zero, the JSON can be entirely valid and all types can satisfy constraints. Business validation reveals missing unit information or incomplete conversion; evidence validation reveals the unsupported tax rate and conflicting freight status. More mandatory numeric fields without an unknown state can encourage missing facts to be packaged as complete records.

If the procurement system requires agreed freight on an authoritative order, route the whole candidate to review with an explicit missing-item list. If staged storage is supported, retain the equipment quotation and unresolved state, but do not present it as the full procurement total. The correct behavior follows the destination record’s business meaning, not merely whether the database permits nulls.

Now suppose another page also contains an older quotation. Choosing the larger amount or the value nearest the page bottom is insufficient. Establish version, effectiveness conditions and business confirmation. When unresolved, flag the conflict and cite both passages rather than manufacturing one definite amount.

Limit what the model can change during repair

After validation fails, return the specific error and relevant source passage, asking the model to re-extract affected fields. Do not repeatedly ask it to “make the record pass,” or allow verified line items to change merely to balance a total. Recheck affected dependencies after every repair, retaining differences and attempt counts.

Separate retryable parsing or format problems, unit and format issues correctable by explicit rules, and factual problems requiring materials or human confirmation. Ten further requests cannot reveal a tax rate absent from the source. Set retry limits and an escalation route according to cost, timeliness and consequences; this article prescribes no universal count.

An independent service should perform authoritative writes; validation itself should not alter business records. Use task identifiers and document revisions to identify repeated imports so retries do not produce duplicate rows. Recheck destination permissions and the pending data revision immediately before entry. Passing structural and business checks does not grant write authorization.

Accept the workflow with explanatory metrics, not parsing rate alone

Report complete-response rate, structural pass rate, field-evidence correctness and whole-record acceptance separately. Use all valid submitted tasks as the denominator for complete responses. Structural pass rate may use completed responses as its denominator, but label this condition and also report the proportion of all tasks. Whole-record acceptance requires all agreed critical fields to pass; do not average away a severe field error.

Measure field-evidence correctness through labeled references or human review and state whether counting fields or records. Correctly marking an absent field as unknown can be correct handling; if the task requires a complete document, the record may nevertheless be ineligible for automatic entry. Also report automatic-release proportion, review workload and erroneous automatic releases.

Constructed calculation: of 100 inputs, 95 produce complete responses, 90 pass structure checks and 78 meet every entry requirement. Relative to all tasks, the rates are 95%, 90% and 78%. The conditional structural pass rate among completed responses is about 94.7%. Reporting only that higher number as overall accuracy would be misleading. These are arithmetic examples only.

Include clear documents, missing fields, zero values, different monetary units, conflicting revisions, cross-page tables, blurred scans and truncated responses. Break down errors by material type and check whether the review queue can meet business deadlines. Lower automatic release accompanied by growing review backlog still requires scope adjustment or more review capacity.

The deliverables are a structural contract, independent rules, evidence locations, exception states and repeatable examples. They add implementation work but stop errors at identifiable gates. Structured output is ready to participate in a business data pipeline only when format, meaning and source support are jointly established.

References

  1. [1] Anthropic: Structured outputs(living documentation; accessed October 2, 2026)
  2. [2] JSON Schema: Understanding JSON Schema — Objects(accessed October 2, 2026)
  3. [3] Pydantic: Strict Mode(living documentation; accessed October 2, 2026)
  4. [4] Pydantic: Validators(living documentation; accessed October 2, 2026)
Back to insights
鲁ICP备2024109755号-2
Drag to move. Right-click, touch and hold, or press Shift+F10 to choose a corner.