Using Claude to catch a hidden printing defect before it ships.
Some of Stupell's product artwork was sourced from seamless, mirror-tiled patterns. When a single product image got cropped out of that tile without trimming to the boundary, mirrored fragments of the design bled in from the edges — a defect that's easy to miss at a glance. We replaced a brittle computer-vision detector with a Claude vision classifier tuned specifically to catch it.
Client
Stupell Industries
Defect
Mirrored / bled-in tile artifacts
Role
AI Engineer

Some of Stupell's source artwork was authored as a seamless, mirror-tiled repeating pattern — common for wallpaper, fabric, and watercolor-style designs. When a single product canvas was cropped out of that tile without trimming all the way to the tile boundary, small mirrored or rotated fragments of the same design bled in from one or more borders. The main subject still looked untouched in the center, so the defect was easy to miss on a quick look — especially in busy, multi-element compositions (rows of trees, fields of stars, painterly textures) where a mirrored fragment just looks like "another one of the same kind of element" instead of a duplicate.
The classifier runs as a minimal Flask API, callable directly from Stupell's existing production pipeline rather than as a separate standalone service. Input can be either a flattened product image or a print-ready PDF; PDFs go through a low-memory extraction path built on PyMuPDF and poppler so large, high-resolution files don't blow up memory before the image ever reaches the model.
The core of the system isn't a trained model — it's a detailed, taxonomy-driven system prompt given to Claude's vision capability, encoding exactly what the defect looks like (mirror-fold symmetry, corner kaleidoscope patterns) and what to explicitly ignore (barcodes, labels, legitimate repeating materials like wood grain or brick). The API returns a simple true/false flag with a confidence score, which is all the upstream pipeline needs to gate a file before print.
- 01A Claude vision classifier tuned to the specific 'repeat-tile bleed' printing defect, not a general anomaly detector
- 02A taxonomy-driven prompt that distinguishes the defect from barcodes, labels, and legitimate repeating materials like wood grain or brick
- 03High-recall calibration so ambiguous cases get flagged for human review instead of silently passing
- 04A minimal localhost API returning just a true/false flag and confidence score, callable from the existing production pipeline
- 05Support for both flattened product images and print-ready PDFs, with a low-memory extraction path for large files
Claude API + Computer Vision + OpenCV + Flask + PyMuPDF + Print QA Classifier + High-Recall Calibration
Built with the Claude API (vision), Flask, OpenCV, and PyMuPDF/poppler for PDF extraction.
How we built it
Start with classical computer vision
The first version used OpenCV: estimate the background color, detect barcode/label furniture, then flag any foreground blob that touched the canvas border away from the main subject.
Find where it breaks
On a real high-resolution production file, fine texture edges (leaf veins, petal edges) chained into one canvas-spanning blob, silently swallowing the barcode regions and missing real defects — the heuristic needed constant re-tuning per edge case.
Replace it with a Claude vision classifier
We wrote a detailed system prompt encoding the actual defect taxonomy — mirror-fold symmetry, corner kaleidoscope patterns — and what to explicitly ignore, like barcodes or legitimate repeating materials such as wood grain or brick.
Calibrate for high recall
A missed defect costs far more than a false alarm, so the prompt resolves genuinely ambiguous cases toward flagging them for review rather than passing silently.
What actually got in the way
Classical computer vision couldn't generalize
The original OpenCV heuristic — estimate the background, detect barcode/label furniture, flag any foreground blob touching the border — broke on real production files. Fine texture edges (leaf veins, petal edges) chained into one canvas-spanning blob, silently swallowing barcode regions and missing the actual defect.
Calibrating for high recall without drowning the team in false alarms
A missed defect costs far more than a false alarm, so the prompt had to resolve genuinely ambiguous cases toward flagging them — without flagging so aggressively that the QA team stopped trusting the tool and started ignoring its output.
Supporting large print-ready PDFs without exhausting memory
Some source files are large, high-resolution, print-ready PDFs. Extracting a usable image for the vision model needed a genuinely low-memory path rather than loading the full file into memory on every check.
The classifier now runs as a QA gate ahead of print, catching mirrored and bled-in artwork that's genuinely easy to miss at a glance — without the constant threshold-tuning the original computer-vision approach needed every time a new edge case showed up.
“Any issues that have occurred during our hundreds of hours of work together have been addressed and improved on.”
Todd Stupell
Owner, Stupell Industries
Verified review
Clutch5.0A generic anomaly detector is usually the wrong tool. Tuning to the specific defect taxonomy — and explicitly naming what to ignore, like barcodes or legitimate repeating materials — is what separated a reliable classifier from a noisy one.
Recall over precision is a design decision, not a default. Deciding that a missed defect costs more than a false alarm has to be made explicitly and encoded in the prompt, not left to whatever the model does by default.
Minimal infrastructure was the right call here — a small API returning a true/false flag and confidence score was far easier to drop into the existing production pipeline than standing up a heavier standalone service.
Have a similar product idea?
Share the workflow, product idea, or knowledge problem you're solving, and we'll help define the smartest first build.