Scan your guide free

What the scan checks

The scan runs 33 checks, in four groups. Each one marks something software will miss or misread, and says where the check comes from: a published standard, a known failure in PDF text extraction, or our own rule.

This list is generated from the scanner itself, scan version 5.

Text a machine can't see

11 checks

Words that exist only inside images

AI reading the text never sees them.

Source
Tesseract OCR of rendered image regions; words as recognized, engine version in the report

Large images with no text over them

Anything written inside them is invisible to text extraction.

Our rule
GuideReady heuristic; not taken from a published standard

Pages with no extractable text

Scanned or flattened pages read as blank.

Source
Known PDF text-extraction failure mode, as handled or documented by pdfminer.six and pdfplumber

Characters that extract as codes

Words come out garbled, such as (cid:12).

Source
ISO 14289-1 §7.21.7 (veraPDF 7.21.7-1)

Icons, bullets, or ligatures that extract as meaningless codes

Some drop letters from words.

Source
Known PDF text-extraction failure mode, as handled or documented by pdfminer.six and pdfplumber

Ligature characters

Words like 'final' may not match searches.

Source
Known PDF text-extraction failure mode, as handled or documented by pdfminer.six and pdfplumber

Text drawn twice in the same place

Common tools extract it doubled.

Source
Known PDF text-extraction failure mode, as handled or documented by pdfminer.six and pdfplumber

Words split across line breaks

They extract as two fragments.

Source
Known PDF text-extraction failure mode, as handled or documented by pdfminer.six and pdfplumber

Rotated text

It extracts out of reading order.

Source
Known PDF text-extraction failure mode, as handled or documented by pdfminer.six and pdfplumber

Very small text

Qualifiers often hide there.

Our rule
GuideReady heuristic; not taken from a published standard

Text outside the visible page

Machines read it; people never see it.

Source
Known PDF text-extraction failure mode, as handled or documented by pdfminer.six and pdfplumber

Tables

5 checks

Table headers with no text

The column labels are probably images.

Our rule
GuideReady heuristic; not taken from a published standard

Row labels a standard extractor drops

Numbers lose what they count.

Source
Known PDF text-extraction failure mode, as handled or documented by pdfminer.six and pdfplumber

Merged or missing table cells

Values can attach to the wrong column.

Source
Known PDF text-extraction failure mode, as handled or documented by pdfminer.six and pdfplumber

Checkmarks or icons that won't read as symbols

A benefit matrix loses its marks.

Source
Known PDF text-extraction failure mode, as handled or documented by pdfminer.six and pdfplumber

Table rows separated from their headers

AI search tools retrieve rows without labels.

Source
GuideReady retrieval chunker (800-token chunks, 100-token overlap); rows retrieved without their header

Rules that invite different readings

10 checks

Rules that point elsewhere for their details

AI given only this guide can't follow them.

Our rule
GuideReady heuristic; not taken from a published standard

Links out of the guide

The rule may live on another page.

Our rule
GuideReady heuristic; not taken from a published standard

Discretion or right-to-change language

AI tends to state a definite answer anyway.

Our rule
GuideReady heuristic; not taken from a published standard

Exception words (unless, except, notwithstanding)

Conditions are easy to drop.

Source
Federal Plain Language Guidelines, 'Avoid double negatives and exceptions to exceptions', p. 54

Rules qualified by 'provided that'

The condition is easy to drop.

Our rule
GuideReady heuristic; not taken from a published standard

Vague phrases (as appropriate, timely)

Open to more than one reading.

Source
ARM weak phrases (Wilson 1997, §e)

Open-ended lists (etc., and/or)

The rule's scope is unclear.

Our rule
GuideReady heuristic; not taken from a published standard

Deadlines in days, not saying business or calendar

A deadline can differ by days.

Our rule
GuideReady heuristic; not taken from a published standard

'Within' a period with no start event

It's unclear when the clock starts.

Our rule
GuideReady heuristic; not taken from a published standard

Footnote markers

The qualifying text sits elsewhere.

Our rule
GuideReady heuristic; not taken from a published standard

File setup

7 checks

Not tagged for accessibility

No logical structure for tools to follow.

Source
ISO 14289-1 §6.2 (veraPDF 6.2-1)

No document language set

Accessibility tools handle it worse.

Source
ISO 14289-1 §7.2 (veraPDF 7.2-29, 7.2-34)

No document title in the metadata

Tools show a file name instead.

Source
ISO 14289-1 §7.1 (veraPDF 7.1-9)

Title not set to display

Viewers show the file name.

Source
ISO 14289-1 §7.1 (veraPDF 7.1-10)

Copying text is restricted

Some tools refuse to read it.

Source
ISO 32000-1:2008 §7.6.3.2, Table 22, bit 5

Accessibility extraction is blocked

Screen readers and parsers are refused.

Source
ISO 14289-1 §7.16 (veraPDF 7.16-1)

No version or effective date

Readers can't tell which rules are current.

Our rule
GuideReady heuristic; not taken from a published standard

Each finding is a fact about how the file reads to software, not a judgment of the program. Thresholds, such as what counts as very small text, are our own settings.

See which of these are in your guide.

Scan your guide free