Scanned documents (OCR)
What the rule requires
A scanned page is just a photograph of text — a screen reader finds nothing to read. Before any other accessibility work can happen, the document needs a real text layer.
How Proof clears it
Detected at prescan: image-only documents are identified the moment they arrive.
The page is rebuilt as structured content — real headings, paragraphs, reading order, and tab order, with the scan artifacted underneath. It looks identical, and reads like a real document.
Chains straight into the rest: the structured output feeds tagging and reading-order repair, and every rule re-verifies it on the same document.
What to expect
OCR output is structured, accessible-by-default — not just an invisible text layer. Heading recovery from scans is still heuristic, so review flagged pages rather than trusting every level; the rule engine re-verifies everything afterward.