Skip to content

Repository files navigation

Uneven-lighting OCR

Reading a crumpled thermal receipt photographed under a single overhead light, where the illumination gradient defeats both global and standard adaptive thresholding.

Coursework built for a university image-processing unit (December 2024). It is a study of one hard image, not a general-purpose OCR library — see Status before you clone it.


The problem

Thermal receipt paper photographed by hand has two properties that fight OCR:

  1. A radial illumination gradient. The centre of the shot is bright, the corners fall off. A single global threshold either blows out the middle or loses the edges.
  2. Low contrast between ink and paper, which shrinks as the paper ages, so the gradient dominates the signal you actually want.

The sample here is a 316×386 photograph of a Whole Foods receipt, skewed about 8° and lit from above.

OpenCV's adaptiveThreshold handles a smooth gradient well. It struggles here because the useful window size is not constant across the image: near the bright centre you want a small window (fine local detail, plenty of signal), and near the dark corners you want a large one (more pixels to establish what "paper" means).

The approach

Two thresholding strategies were tried against the same preprocessing chain.

Distance-weighted block thresholding (practical-5.py)

The image is tiled into blocks whose size shrinks with distance from the centre, matching the tiling to the illumination gradient rather than assuming a uniform one:

dist = np.sqrt((y - center[0]) ** 2 + (x - center[1]) ** 2)
block_size = max(
    VARIABLE_BLOCK_MIN_SIZE,
    int(VARIABLE_BLOCK_BASE_SIZE - (dist / max(h, w)) * VARIABLE_BLOCK_DISTANCE_FACTOR),
)

Because the size depends on both y and x, blocks near the middle of a row come out taller than the ones at its ends, and a row cannot simply advance by "the block size" — there isn't one. Each row is therefore laid out in two passes: the first walks the columns to fix their widths and takes the row's height to be the tallest block in it, the second emits every block in the row at that shared height. Widths still follow the size law across the row, heights are constant within a row, and the tiling covers every pixel exactly once. benchmark.py asserts that on every run.

Each block is then thresholded against its own mean, floored so that a block containing only paper cannot hallucinate ink:

block_threshold = max(BLOCK_THRESHOLD_BASE, block_mean - BLOCK_THRESHOLD_OFFSET)

Blocks are written to disk, stitched back into a single image in (y, x) order, and handed to Tesseract. Three tiling laws are implemented and selectable via BLOCK_GENERATION_METHOD: linear (fixed size, the control), variable (linear falloff, above) and cubic (falloff by distance³, which keeps blocks large across a wider central region and then drops off sharply).

Cross-probe per-pixel thresholding (archive/practical-7.py)

A second strategy, per pixel rather than per block. Sample four neighbours at ±25px on the cardinal axes; if all four are bright and their mean exceeds a white cutoff, the pixel is deep inside paper, so raise its threshold by 42:

if neighbor_pixels_avg > white_threshold_value:
    new_threshold = threshold_value + 42

This gives cleaner results in open whitespace but is a pure-Python double loop over a 10×-upscaled image, so it is slow by orders of magnitude.

Preprocessing chain

Shared by both, each stage individually toggleable and each writing a numbered intermediate to p5/ so the chain can be inspected step by step:

Step Operation Notes
0 Upscale 4× INTER_CUBIC, then non-local-means denoise
1 Super-resolution EDSR ×4 — implemented, disabled by default
2 Greyscale
3 Contrast convertScaleAbs, α=1.6 β=−125
4 Sharpen 3×3 Laplacian-style kernel
5 Rotate −7.8°, hand-measured deskew
6 Crop 9% vertical / 17% horizontal margin

Tesseract runs with a receipt-specific character blacklist (all lowercase letters are excluded — receipt print is uppercase, so lowercase output is always an error) and two configurable OEM/PSM sets for A/B comparison.

Results

groundtruth.txt is a hand transcript of the sample receipt, so runs can be scored instead of eyeballed. Every price on it was checked against the printed total: the fourteen line items sum to 101.33, which is what BAL says.

benchmark.py runs all three tiling laws through the same preprocessing chain and the same Tesseract configuration and reports character error rate — the Levenshtein distance from the transcript over the transcript's length.

law         blocks  edits      CER  tiling     stitched sha256
---------------------------------------------------------------
variable       223     62   14.3%  exact      ff592c68a9e45671
cubic          130     65   14.9%  exact      a45f0efd44179371
linear         117     82   18.9%  exact      4131edadb2af7d5e

Read that carefully. Distance-weighted tiling beats the fixed-size control by 4.6 points, which is the claim the project was built to make and it holds. variable beats cubic by 0.6 points — three edits out of 435 — on a single image, which is not a result. Two laws that tile the same photograph 223 and 130 ways landing within three characters of each other says the falloff exponent barely matters once the tiling is distance-weighted at all. Separating them needs more receipts, not more tuning.

Both sides of the comparison are normalised first: uppercased, folded to A–Z 0–9 . , : ; - / ( ), whitespace collapsed. Characters the Tesseract blacklist forbids — the * bullet column among them — are dropped from the transcript as well as from the output, so the pipeline is not charged for never emitting a character it was configured to suppress. / is scored but also blacklisted, which costs every law the slash in 85/15; that is one character and it is the same one character for all three.

The best run is still legible rather than clean:

WHOLE FOOOS MARKET - WLSTAORT.CT O6SRD
3€ BACON LS MP 4 99 F
365 BACON LS NP 4.99
366  BACCN LS NP 4.99 F
365  DACON (S NP 4.99 F

Four scans of the same line item decode four different ways. Prices survive better than SKUs; decimal points are dropped more often than digits are misread. Good enough to show the tiling idea works, not good enough to trust. (outputfile.txt holds an earlier run, kept for comparison.)

What the reproducibility fix was worth

Before the row-advance fix, generate_blocks_variable advanced y by the block size computed at the last column of the row, while blocks nearer the centre of that row were larger — so every one of the 21 rows overlapped the next, by 8 to 27px. stitch_blocks then iterated os.listdir, so which row won inside the overlap band depended on filesystem ordering.

That was not a cosmetic problem. Stitching the same 273 blocks in ascending versus descending (y, x) order gave 17.7% and 18.6% CER on identical input — the number moved by 0.9 points depending on how the filesystem happened to enumerate a directory. The cubic law overlapped by 2–6px and scored 18.9%/19.1% the same way.

Fixing the row height brought variable to 14.3% and cubic to 14.9%, and both are now byte-identical across runs. The old overlap bands were thin and landed mid-character, which is a large part of why the same line item used to decode four different ways.

Status

Coursework, cleaned up but still coursework. Published because the distance-weighted tiling idea is worth reading, not because the code is ready to use. Concretely:

  • No test suite, no CI, no packaging. benchmark.py checks the tiling is an exact partition and scores the output, which is the closest thing to a test here. Paths in archive/ are hardcoded relative to the script's own location, so those scripts only run from inside archive/.
  • One sample image. Every constant — the −7.8° rotation, the crop margins, α/β, the ±25px probe distance — was tuned by hand against that single image and will not transfer. It is also why variable vs cubic is unsettled.
  • The transcript is one person's reading of a blurred photograph. The four SKU columns read 365 on a brand-name product priced identically four times, and the line-item prices sum exactly to the printed BAL, so the digits are well constrained — but the logo lines (WHOLE / FOODS / MARKET) are stylised artwork being scored as text, and that is a judgement call.
  • Configuration is module-level globals, not a config file or CLI flags.

Roadmap

Ordered by what would most improve the result, not by effort.

1 — Make it general

  • Replace the hand-measured −7.8° with automatic deskew (Hough or minimum-area rectangle) and the fixed crop with contour-based receipt detection.
  • Estimate the illumination gradient from the image itself rather than assuming it is radial and centred — a low-pass background model would let the tiling law follow the actual light, and would handle off-centre lighting.
  • Collect a handful more receipts. One image cannot distinguish a good method from a lucky constant, and the 0.6-point gap between variable and cubic is exactly the kind of thing more images would settle.

2 — Make it usable

  • Package as a CLI with a config file instead of module-level globals.
  • Vectorise the practical-7.py cross-probe with array shifts — same result, without the Python loop.
  • Add a licence.

Done — reproducibility (exact-partition tiling, order-independent stitching, pinned dependencies) and measurability (ground-truth transcript, CER benchmark across the three laws). Generated output and the 37 MB model no longer ship in the tree. Note that they are gone from HEAD, not from history: a clone still pays for them until the history is rewritten.

Repository layout

practical-5.py           the main pipeline — start here
benchmark.py             score the three tiling laws by character error rate
groundtruth.txt          hand transcript of the sample receipt
fetch_model.py           download EDSR_x4.pb (37MB) on demand
archive/practical-1.py   image → Excel via Tesseract, baseline
archive/practical-2.py   the same, refactored
archive/practical-3.py   morphological table/cell detection
archive/practical-4.py   adaptive threshold + morphology baseline
archive/practical-6.py   contrast/upscale experiments
archive/practical-7.py   cross-probe per-pixel thresholding
tsconfig                 the raw tesseract CLI invocation
img.png                  the sample receipt
outputfile.txt           extracted text from an earlier run

The numbering is chronological — the practicals were submitted in sequence, so the archive doubles as a record of what was tried and discarded.

p5/, bench-out/ and archive/p6/*.png are generated and gitignored; each run rewrites them from scratch.

Running it

Requires Python 3 and a system Tesseract install (brew install tesseract). The numbers above came from Tesseract 5.4.1 / leptonica 1.85.0 with the pinned Python versions in requirements.txt; character error rates move with the Tesseract version, so report yours alongside any number you quote.

pip install -r requirements.txt
python practical-5.py

Intermediates land in p5/, extracted text goes to stdout.

To score the tiling laws instead:

python benchmark.py

Stitched images and per-law OCR output land in bench-out/<law>/. The stitched sha256 column should match the table above on any machine; if it does not, something in the chain is not reproducible and that is worth knowing.

The super-resolution step is disabled by default. To enable it, fetch the model first — it is verified against a pinned SHA-256 — then set SUPER_RESOLUTION_ENABLED = True:

python fetch_model.py

About

OCR experiment using open cv in order to cover un-even lighting and low resolution

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages