Why can an anomaly detector perform well on a public dataset yet overwhelm inspectors on a production line? Research scores describe discrimination under particular evaluation conditions, while production must decide what happens to every part. We propose accepting systems against missed defects, false alarms, cycle time and review capacity at a fixed threshold. The method targets visual screening and does not treat a heatmap or a natural-language explanation as proof that a product meets its quality requirements.
Establish the gap between model output and the factory's decision
Industrial visual anomaly detection generally seeks samples that depart from normal appearance. PatchCore, for example, uses a representative memory bank of local features from normal images for anomaly detection and localization.[1] This provides a candidate technical approach, not evidence that every material, defect or inspection station is suited to it.
An anomaly score is not inherently the probability that a product is nonconforming, nor a measure of defect severity. A new batch's texture may be unusual but acceptable; an out-of-tolerance dimension may not be apparent in an ordinary two-dimensional image. Quality specialists should first define defect classes, acceptable variation, rejection criteria and issues the imaging setup cannot assess. Then decide whether the model performs screening, localization or final release.
Our engineering recommendation separates three decisions: whether the image is usable, whether the appearance warrants review, and whether the product meets business quality requirements. Blur, occlusion or a missing required view should lead to recapture or manual handling. A low model score must not automatically classify the product as normal. A multimodal model's explanation can assist review but should not replace defect evidence and quality rules.
Use parts as evaluation units and operating conditions as test boundaries
If a part has four images and production sends it to review whenever any image triggers, calculate part-level results after combining all four views. Per-image false-positive rates are not part-level rates, and views of the same part should not be distributed between development and acceptance sets. Freeze the view-combination rule before comparing models so the evaluation unit matches the business action.
Test material should cover the planned part numbers, production batches, shifts, cameras, lighting and equipment conditions. MVTec AD 2 explicitly includes test lighting conditions that may not occur in training.[2] This demonstrates that evaluation can deliberately test operating changes; using that dataset does not establish coverage of a particular factory's distribution.
We recommend separate collections of normal production samples, confirmed defects and changed operating conditions. Normal samples estimate false alarms and routine review volume. Defect samples estimate misses by defect type. Changed conditions test boundaries. Deliberately enriching a collection with defects helps examine rare cases, but its defect prevalence cannot be used to estimate alert precision on the real line.
Labels should retain part identifiers, acquisition time, operating conditions, final decisions and the basis for review. Mark disputed samples as awaiting adjudication rather than quietly deleting difficult cases. When dimensions, internal defects or strength require other measurement methods, use the relevant instruments or inspection process as evidence and state which portion visual inspection can cover.
AUROC compares discrimination; release decisions require a fixed threshold
A receiver operating characteristic, or ROC, curve describes how true-positive and false-positive rates vary across thresholds; AUROC is its area.[3] Because it aggregates across thresholds, models with similar AUROC can behave differently in the low-false-positive region a business actually permits. A high AUROC also does not imply an equally high detection rate at the selected threshold.
Report three rates tied directly to disposition. The miss rate is defective parts passed by the model divided by confirmed defective parts. The false-positive rate is conforming parts sent to review divided by confirmed conforming parts. Alert precision is genuine defective parts sent to review divided by all parts sent to review. Their denominators differ and are not interchangeable. Report timeouts, missing images and missing results separately, and include them in full-process disposition statistics.
Choose the threshold and image postprocessing rules on development material, then freeze them for independent acceptance material. If thresholds are adjusted after inspecting acceptance errors, that material has become part of development and additional independent validation evidence is needed. Normal samples alone can estimate false alarms at a threshold; without defect evidence they cannot establish its miss rate.
Threshold selection should satisfy both the limit on missed critical defects and the available review capacity. If no candidate threshold satisfies both, improve imaging, the model or the station workflow, or narrow the automated scope. Raising the threshold to suppress false alarms may increase misses. Lowering it to detect more defects may create a review backlog. These limits must be determined on site; this article supplies no universal industry threshold.
Constructed example: a low false-positive rate can still fill the review station
The following is entirely constructed; no model was run and no measurements were obtained. Assume 10,000 parts are inspected daily and true defect prevalence is 0.2%: twenty defective and 9,980 conforming parts. Further assume that a fixed threshold gives a 95% detection rate and a 1% false-positive rate, and that these conditions apply to the production distribution.
Using expected counts, the model detects nineteen defective parts, misses one, and sends approximately one hundred conforming parts to review each day. Only about 16% of roughly 119 alerts represent genuine defects. Apparently strong scores can therefore coexist with alerts dominated by conforming products. The 16% is arithmetic under the stated assumptions, not the performance of a particular model.
At an average of two minutes per reviewed part, those alerts require approximately 238 minutes of inspection. If only 180 minutes are reserved, work accumulates. Concentrated alerts can also violate the deadline for a part to move through the station even when total daily review hours are sufficient. Measure peak arrivals, the distribution of review durations and inspectors' actual availability.
In a separate assumption that holds prevalence and detection rate constant, reducing the false-positive rate to 0.1% produces approximately twenty-nine daily alerts and fifty-eight minutes of review. This illustrates why the low-false-positive region matters; it does not imply that changing a real threshold will preserve detection. Model teams must disclose the complete trade-off rather than showcase its most favorable metric.
Zero observed misses still require sample size and uncertainty
Detecting every defect in a small acceptance sample does not establish zero future misses. NIST's statistical documentation notes that common normal-approximation intervals may be inadequate when failure counts or sample sizes are small, and describes an exact binomial approach.[4] For critical defects, report counts and an interval alongside the percentage.
The following calculation is derived here from a binomial model. Assume defective parts are independent, drawn from the same target distribution, and have a fixed true miss probability p. The probability of zero misses in n defective parts is (1-p)^n. Setting that probability to 0.05 gives a one-sided 95% upper confidence bound of 1-0.05^(1/n) when no misses are observed. Detecting all sixty defects still yields an upper bound of about 4.87%; detecting all three hundred yields about 0.99%. These are neither acceptance standards nor measured results.
The calculation cannot repair missing coverage. Hundreds of photographs of one defective part are not hundreds of independent defective parts. Testing scratches alone does not justify applying the bound to cracks. Samples from one abnormal production event may be correlated. Explain the evidence in terms of independent parts, events and operating conditions, using a grouped statistical design where necessary.
When defects are rare, accumulate evidence through retained samples, defective parts confirmed by quality specialists, and subsequent shadow operation. Artificially induced or synthetic defects can support development and stress checks, but they cannot replace all on-site acceptance material before their similarity to real defects is established. Retain manual inspection or another detection method for critical categories with insufficient evidence.
Investigate imaging and distribution before replacing the model
If false alarms rise on one shift, replay the original images, crops, resized inputs and anomaly maps for the same parts. Check reflections, focus, exposure, fixture movement and background changes. If resizing has removed a tiny defect, a larger model cannot be assumed to recover information absent from its input. Higher resolution, multiple views and local tiling are options, but add acquisition, inference and result-combination costs, so cycle time must be reassessed.
When defect definitions are stable and sufficient labels exist, compare supervised classification or segmentation. When normal appearance is relatively stable but defect types are open-ended, consider anomaly detection trained on normal material. When the central issue is dimensions or geometric tolerances, assess metrology, rules or three-dimensional methods. Observable information and task requirements should guide the choice; a newer method name does not establish better suitability.
False alarms may also reflect incomplete coverage of normal materials. Adding confirmed normal samples can be more direct than replacing the underlying model, but disputed defects must not be absorbed as normal. Changes to a memory bank, model, threshold or preprocessing should create a new version, reviewed against a fixed regression set and samples from the relevant conditions.
For tasks that guide human inspectors, also verify whether the anomaly region covers the target defect and points to the correct location. Pixel-level scores cannot replace part-level miss rates. Correctly alerting on a part does not establish adequate localization. Define localization criteria around the review action—for example, whether an inspector can find the region requiring confirmation—rather than merely displaying a vivid heatmap.
Use shadow operation to test the complete disposition chain
We recommend running the model alongside the existing inspection process first. Record which parts it would flag, the final human judgments and where time is spent. Measure from the part's acquisition trigger to the disposition system receiving a valid result, including transmission, preprocessing, queueing, inference and communication. Model inference time alone does not represent production-line cycle time.
Sample parts that did not trigger an alert; otherwise only alert quality is visible and misses remain undiscovered. The sampling plan should cover high-risk conditions and random production traffic, with sampling probabilities and origins recorded. A review sample biased toward difficult cases cannot directly estimate the line-wide miss rate. Shadow observation does not replace existing inspection for defect categories still lacking evidence.
Also exercise camera outages, result timeouts, part-identifier mismatches and review-station congestion. Predetermined responses may include recapture, isolation, manual takeover or stopping the line under site rules. “No result” must not be treated as “no anomaly.” Automatic release or rejection additionally requires confirming that the result corresponds to the correct physical part.
Deliver evidence that makes acceptance repeatable
We recommend an evidence package containing acquisition conditions and part coverage, labeling and dispute-resolution records, model and preprocessing versions, thresholds and view-combination rules, per-part decisions, misses and false alarms broken down by defect and operating condition, statistical intervals, cycle time and review capacity, and failure-handling exercises. Quality, production and technical owners should jointly confirm the release scope.
After launch, alert rates and review backlogs can help reveal changes, but stable alert rates alone do not establish stable miss rates; continued sampling remains necessary. Changes to lighting, lenses, materials, processes or models require revalidation appropriate to their impact. This article develops a method from public research and official documentation and does not claim that IDENIFE has performed these experiments. Online sources were checked on October 1, 2026.
References
- [1] Roth et al.: Towards total recall in industrial anomaly detection—PatchCore research overview, Amazon Science (2022)
- [2] MVTec: Official MVTec AD 2 dataset description
- [3] scikit-learn: Metrics and scoring—ROC, confusion matrices and classification metrics
- [4] NIST Dataplot: Exact Binomial Confidence Limits
