Reading a crumpled thermal receipt photographed under a single overhead light, where the illumination gradient defeats both global and standard adaptive thresholding.
Coursework built for a university image-processing unit (December 2024). It is a study of one hard image, not a general-purpose OCR library — see Status before you clone it.
Thermal receipt paper photographed by hand has two properties that fight OCR:
- A radial illumination gradient. The centre of the shot is bright, the corners fall off. A single global threshold either blows out the middle or loses the edges.
- Low contrast between ink and paper, which shrinks as the paper ages, so the gradient dominates the signal you actually want.
The sample here is a 316×386 photograph of a Whole Foods receipt, skewed about 8° and lit from above.
OpenCV's adaptiveThreshold handles a smooth gradient well. It struggles
here because the useful window size is not constant across the image: near the
bright centre you want a small window (fine local detail, plenty of signal),
and near the dark corners you want a large one (more pixels to establish what
"paper" means).
Two thresholding strategies were tried against the same preprocessing chain.
The image is tiled into blocks whose size shrinks with distance from the centre, matching the tiling to the illumination gradient rather than assuming a uniform one:
dist = np.sqrt((y - center[0]) ** 2 + (x - center[1]) ** 2)
block_size = max(
VARIABLE_BLOCK_MIN_SIZE,
int(VARIABLE_BLOCK_BASE_SIZE - (dist / max(h, w)) * VARIABLE_BLOCK_DISTANCE_FACTOR),
)Because the size depends on both y and x, blocks near the middle of a row
come out taller than the ones at its ends, and a row cannot simply advance by
"the block size" — there isn't one. Each row is therefore laid out in two
passes: the first walks the columns to fix their widths and takes the row's
height to be the tallest block in it, the second emits every block in the row
at that shared height. Widths still follow the size law across the row, heights
are constant within a row, and the tiling covers every pixel exactly once.
benchmark.py asserts that on every run.
Each block is then thresholded against its own mean, floored so that a block containing only paper cannot hallucinate ink:
block_threshold = max(BLOCK_THRESHOLD_BASE, block_mean - BLOCK_THRESHOLD_OFFSET)Blocks are written to disk, stitched back into a single image in (y, x) order,
and handed to Tesseract. Three tiling laws are implemented and selectable via
BLOCK_GENERATION_METHOD: linear (fixed size, the control), variable
(linear falloff, above) and cubic (falloff by distance³, which keeps blocks
large across a wider central region and then drops off sharply).
A second strategy, per pixel rather than per block. Sample four neighbours at ±25px on the cardinal axes; if all four are bright and their mean exceeds a white cutoff, the pixel is deep inside paper, so raise its threshold by 42:
if neighbor_pixels_avg > white_threshold_value:
new_threshold = threshold_value + 42This gives cleaner results in open whitespace but is a pure-Python double loop over a 10×-upscaled image, so it is slow by orders of magnitude.
Shared by both, each stage individually toggleable and each writing a numbered
intermediate to p5/ so the chain can be inspected step by step:
| Step | Operation | Notes |
|---|---|---|
| 0 | Upscale | 4× INTER_CUBIC, then non-local-means denoise |
| 1 | Super-resolution | EDSR ×4 — implemented, disabled by default |
| 2 | Greyscale | |
| 3 | Contrast | convertScaleAbs, α=1.6 β=−125 |
| 4 | Sharpen | 3×3 Laplacian-style kernel |
| 5 | Rotate | −7.8°, hand-measured deskew |
| 6 | Crop | 9% vertical / 17% horizontal margin |
Tesseract runs with a receipt-specific character blacklist (all lowercase letters are excluded — receipt print is uppercase, so lowercase output is always an error) and two configurable OEM/PSM sets for A/B comparison.
groundtruth.txt is a hand transcript of the sample receipt, so runs can be
scored instead of eyeballed. Every price on it was checked against the printed
total: the fourteen line items sum to 101.33, which is what BAL says.
benchmark.py runs all three tiling laws through the same preprocessing chain
and the same Tesseract configuration and reports character error rate — the
Levenshtein distance from the transcript over the transcript's length.
law blocks edits CER tiling stitched sha256
---------------------------------------------------------------
variable 223 62 14.3% exact ff592c68a9e45671
cubic 130 65 14.9% exact a45f0efd44179371
linear 117 82 18.9% exact 4131edadb2af7d5e
Read that carefully. Distance-weighted tiling beats the fixed-size control by
4.6 points, which is the claim the project was built to make and it holds.
variable beats cubic by 0.6 points — three edits out of 435 — on a single
image, which is not a result. Two laws that tile the same photograph 223 and
130 ways landing within three characters of each other says the falloff
exponent barely matters once the tiling is distance-weighted at all.
Separating them needs more receipts, not more tuning.
Both sides of the comparison are normalised first: uppercased, folded to
A–Z 0–9 . , : ; - / ( ), whitespace collapsed. Characters the Tesseract
blacklist forbids — the * bullet column among them — are dropped from the
transcript as well as from the output, so the pipeline is not charged for
never emitting a character it was configured to suppress. / is scored but
also blacklisted, which costs every law the slash in 85/15; that is one
character and it is the same one character for all three.
The best run is still legible rather than clean:
WHOLE FOOOS MARKET - WLSTAORT.CT O6SRD
3€ BACON LS MP 4 99 F
365 BACON LS NP 4.99
366 BACCN LS NP 4.99 F
365 DACON (S NP 4.99 F
Four scans of the same line item decode four different ways. Prices survive
better than SKUs; decimal points are dropped more often than digits are
misread. Good enough to show the tiling idea works, not good enough to trust.
(outputfile.txt holds an earlier run, kept for comparison.)
Before the row-advance fix, generate_blocks_variable advanced y by the
block size computed at the last column of the row, while blocks nearer the
centre of that row were larger — so every one of the 21 rows overlapped the
next, by 8 to 27px. stitch_blocks then iterated os.listdir, so which row
won inside the overlap band depended on filesystem ordering.
That was not a cosmetic problem. Stitching the same 273 blocks in ascending
versus descending (y, x) order gave 17.7% and 18.6% CER on identical
input — the number moved by 0.9 points depending on how the filesystem
happened to enumerate a directory. The cubic law overlapped by 2–6px and
scored 18.9%/19.1% the same way.
Fixing the row height brought variable to 14.3% and cubic to 14.9%, and
both are now byte-identical across runs. The old overlap bands were thin and
landed mid-character, which is a large part of why the same line item used to
decode four different ways.
Coursework, cleaned up but still coursework. Published because the distance-weighted tiling idea is worth reading, not because the code is ready to use. Concretely:
- No test suite, no CI, no packaging.
benchmark.pychecks the tiling is an exact partition and scores the output, which is the closest thing to a test here. Paths inarchive/are hardcoded relative to the script's own location, so those scripts only run from insidearchive/. - One sample image. Every constant — the −7.8° rotation, the crop margins,
α/β, the ±25px probe distance — was tuned by hand against that single image
and will not transfer. It is also why
variablevscubicis unsettled. - The transcript is one person's reading of a blurred photograph. The four
SKU columns read
365on a brand-name product priced identically four times, and the line-item prices sum exactly to the printedBAL, so the digits are well constrained — but the logo lines (WHOLE/FOODS/MARKET) are stylised artwork being scored as text, and that is a judgement call. - Configuration is module-level globals, not a config file or CLI flags.
Ordered by what would most improve the result, not by effort.
1 — Make it general
- Replace the hand-measured −7.8° with automatic deskew (Hough or minimum-area rectangle) and the fixed crop with contour-based receipt detection.
- Estimate the illumination gradient from the image itself rather than assuming it is radial and centred — a low-pass background model would let the tiling law follow the actual light, and would handle off-centre lighting.
- Collect a handful more receipts. One image cannot distinguish a good method
from a lucky constant, and the 0.6-point gap between
variableandcubicis exactly the kind of thing more images would settle.
2 — Make it usable
- Package as a CLI with a config file instead of module-level globals.
- Vectorise the
practical-7.pycross-probe with array shifts — same result, without the Python loop. - Add a licence.
Done — reproducibility (exact-partition tiling, order-independent
stitching, pinned dependencies) and measurability (ground-truth transcript,
CER benchmark across the three laws). Generated output and the 37 MB model no
longer ship in the tree. Note that they are gone from HEAD, not from history:
a clone still pays for them until the history is rewritten.
practical-5.py the main pipeline — start here
benchmark.py score the three tiling laws by character error rate
groundtruth.txt hand transcript of the sample receipt
fetch_model.py download EDSR_x4.pb (37MB) on demand
archive/practical-1.py image → Excel via Tesseract, baseline
archive/practical-2.py the same, refactored
archive/practical-3.py morphological table/cell detection
archive/practical-4.py adaptive threshold + morphology baseline
archive/practical-6.py contrast/upscale experiments
archive/practical-7.py cross-probe per-pixel thresholding
tsconfig the raw tesseract CLI invocation
img.png the sample receipt
outputfile.txt extracted text from an earlier run
The numbering is chronological — the practicals were submitted in sequence, so the archive doubles as a record of what was tried and discarded.
p5/, bench-out/ and archive/p6/*.png are generated and gitignored; each
run rewrites them from scratch.
Requires Python 3 and a system Tesseract install (brew install tesseract).
The numbers above came from Tesseract 5.4.1 / leptonica 1.85.0 with the pinned
Python versions in requirements.txt; character error rates move with the
Tesseract version, so report yours alongside any number you quote.
pip install -r requirements.txt
python practical-5.pyIntermediates land in p5/, extracted text goes to stdout.
To score the tiling laws instead:
python benchmark.pyStitched images and per-law OCR output land in bench-out/<law>/. The
stitched sha256 column should match the table above on any machine; if it
does not, something in the chain is not reproducible and that is worth knowing.
The super-resolution step is disabled by default. To enable it, fetch the model
first — it is verified against a pinned SHA-256 — then set
SUPER_RESOLUTION_ENABLED = True:
python fetch_model.py