Manufacturing 7 weeks to production, 2 lines Computer VisionMLOps & Deployment

Cutting defect escape 38% with inline computer-vision inspection

How a tier-one automotive components manufacturer replaced sampled manual inspection with a 100%-coverage vision system running at line speed on existing hardware.

Client: Global Manufacturing Co. (name withheld under NDA)

Results

  • −38% Defect escape rate to customer
  • 100% Parts inspected (was 4% sampled)
  • 41ms Inference latency per part
  • $2.1M Annualised scrap & warranty avoidance
  • 0.4% False-reject rate at target recall
  • 3.4x First-year ROI

Overview

Global Manufacturing Co. produces precision-machined aluminium components for tier-one automotive customers across two plants. Quality control relied on trained inspectors pulling a 4% sample from each lot under a bright-light station. The process was competent and well-documented — and structurally incapable of catching the defect class that was costing the business most: fine surface porosity that only becomes visible at particular angles of incident light.

The business case was not “replace inspectors.” It was “see every part.” A 4% sample cannot bound the escape rate on a defect that occurs in clusters, and the customer-side cost of an escape — a line stop at the OEM, plus the warranty tail — dwarfed the inspection cost by two orders of magnitude.

The problem

Three constraints shaped everything that followed.

The defect was hard to see on purpose. Porosity, chatter marks and incipient tool-wear scoring present as low-contrast texture differences, not shapes. Off-the-shelf inspection appliances we benchmarked were tuned for presence/absence and dimensional checks; on the client’s own golden sample set they topped out around 71% recall, with a false-reject rate high enough to make operators switch them off.

Line speed was non-negotiable. Two lines running 240 parts per minute, parts in continuous motion, with a reject gate 1.2 metres downstream of the imaging station. That is a hard 180ms budget from trigger to gate decision, including acquisition and PLC round-trip.

The plant network was closed. No outbound internet from the production VLAN, no cloud inference, no exceptions. Anything that needed an API call to a hosted model was disqualified before the conversation started.

And the labels did not exist. The client had six years of inspection records, but they were lot-level pass/fail counts in a quality database — not images, and not localised. We were starting from zero annotated examples of the defect we most needed to catch.

Approach

We ran this as a six-week fixed-scope engagement on one line, with a second line added in week seven after validation passed.

Week 1 — Discover. We instrumented the existing inspection station and measured the actual baseline, which is almost never the baseline a company believes it has. Two inspectors independently graded 800 parts; their agreement with each other was 86%. That number became the honest ceiling for “human-equivalent” and reframed the target: the system did not need to be perfect, it needed to beat a measurable 86% consistency at 25x the coverage.

Week 2 — Architect. Lighting came before modelling. Working with the plant’s controls engineer, we built a four-angle strobed darkfield rig that made porosity visibly high-contrast in at least one of the four exposures. This is the single highest-leverage decision in the project: it converted a hard texture-discrimination problem into a nearly ordinary detection problem, and it is why the model architecture ended up unremarkable. We specified Jetson AGX Orin nodes at the edge, sized the INT8 latency budget, and designed the evaluation harness before writing training code.

Weeks 3–4 — Build. Data collection ran continuously from the end of week two: every part imaged, every inspector decision captured alongside it. SAM 2 generated candidate masks that inspectors corrected rather than drew, which cut annotation time per part from roughly 90 seconds to 22. By the end of week four we had 14,200 labelled parts including 1,870 defect instances across seven classes. Training was a two-stage pipeline — YOLOv8-seg to localise candidate regions, then a per-region EfficientNet-B2 classifier — chosen specifically because it lets quality engineers review why a part was rejected, region by region, rather than arguing with a single opaque score.

Week 5 — Validate. The system ran in shadow mode for nine days: full inference, full logging, reject gate disabled. Every disagreement between model and inspector was adjudicated by the quality manager against the physical part. That adjudication set is the real deliverable of this phase — it is what moved the operating threshold from a guess to a decision, and it exposed two defect classes where the model was systematically over-sensitive to coolant residue. We added 400 targeted examples and retrained.

Weeks 6–7 — Deploy and scale. Live on line one with the reject gate armed, operator override always available, and a one-click rollback to sampled manual inspection. Line two followed with the same model and a 300-part site-specific calibration set.

Solution

The production system images every part through four strobed exposures as it moves, runs localisation and classification on an edge node, and returns a gate decision to the PLC over OPC UA — mean 41ms, p99 68ms, comfortably inside the 180ms budget.

Three design choices mattered more than the model:

Per-region explanations, not scores. The operator HMI shows the part image with each flagged region outlined and labelled by defect class and confidence. Quality engineers trust it because they can check it. Adoption problems in vision inspection are almost always explanation problems.

A threshold that belongs to the plant, not to us. Operating point is a plant-configurable parameter with the recall/false-reject tradeoff curve rendered directly in the HMI. Quality owns that dial, as they should.

Drift detection on inputs, not just outputs. The system monitors image-statistic distributions per exposure channel and alerts on shift — which caught a degrading strobe LED in month three, before it affected classification accuracy at all.

Results

Measured over the first four months of production running, against the instrumented week-one baseline:

  • Defect escape rate to customer down 38%, from 340 PPM to 211 PPM across both lines.
  • Coverage from 4% sampled to 100% inspected, at full line speed with no throughput loss.
  • False-reject rate of 0.4% at the plant’s chosen operating point — below the 1.0% threshold quality set as the condition for arming the gate.
  • 41ms mean inference latency, p99 68ms.
  • $2.1M annualised avoidance in scrap, rework and warranty reserve, validated by the client’s finance team against the prior-year actuals.
  • 3.4x first-year ROI on total engagement and hardware cost.

Inspectors were not displaced. Two moved to adjudicating flagged parts and maintaining the annotation set — higher-skill work that keeps the model current and keeps domain knowledge inside the plant.

Technology stack

See the stack table above. Everything runs on hardware the client owns, inside the production VLAN, with no outbound network dependency. Source code, trained weights, the annotation toolchain and the Terraform and Balena configuration were delivered into the client’s repositories at close.

Let's scope your AI opportunity

A 45-minute conversation is usually enough to tell whether there is a system worth building — and what it would take. No obligation, no deck.