Original research · Document accessibility
What one city's document library actually looks like
Every published document of a mid-sized American city, counted rather than sampled: roughly 2,700 documents and more than 60,000 sheets, every sheet opened and examined. Not an estimate. Not a vendor's claim about its own tooling. A full corpus pass, with the method written down and the two places we were wrong shown as corrections, not quietly fixed.
- Measured
- Judged
- Not assessed
The city is not named. We have written to them, and until they choose to be named, publishing an identifiable audit of someone before speaking with them would be discourteous — the same reasoning our own accessibility statement runs on. Every figure below is rounded for the same reason: precise enough to be useful, not precise enough to identify who it describes. Nothing here is a conformance claim, and nothing here is rounded in a direction that flatters us.
This is a measurement, not a pitch. If you came here from a search rather than from us, the whole page is below — no form, no email address, nothing gated.
What was measured
Counted by opening every document and every sheet, on a machine, against a stated rule — reproducible by anyone who read the same files.
-
8%
of sheets are pictures of text, not text a screen reader can read
-
1 in 5
documents holds at least one such sheet
-
~50%
of budget documents contain a scanned page — the highest of any type
-
900,000+
words recovered from those sheets by on-device recognition
That 8% is a floor, not a ceiling. A sheet only counts as a picture when it carries an image and at most two characters of real text — the allowance exists because a scanned page often carries a stamped page number in a real font on top of it, and treating that as a text layer would hide the sheets being looked for. Some scanned sheets with a longer typed header are not counted by this rule, and every one of them is still a picture a screen reader cannot read.
The pattern behind the budget figure is not random: the pages that get scanned are overwhelmingly the ones that carry a signature — resolutions, certifications, executed agreements. Two document types in the library carried no scanned page at all, across several hundred documents between them — a real signal, not a small-sample artifact.
Roughly two in five documents — and roughly two in five pages — come from a recurring publication: the same board's minutes, the same monthly report, published from the same template every time. That matters more than it sounds: a template is one thing to fix, not hundreds. Nineteen instances of one budget book tagging correctly is worth far more, engineering-hour for engineering-hour, than the same hour spent on a document nobody will ever publish again.
Roughly six in ten pages sit inside a file whose pages are not all alike — a single document can open with a typed resolution and close with scanned survey maps. A breakdown of pages by document type can only honestly describe what a file's first page looks like, not every page inside it, once this many files mix content the way this library's do.
Checked against a machine-checkable proxy for the accessibility standard the law actually requires, not one of the documents passed cleanly — but read that plainly rather than starkly. The proxy standard is stricter than the law in places, and roughly four in ten documents have no underlying structure at all: for those, the honest failure is not "tagged wrong," it is "not tagged," which is a different and larger job to fix. Only a small share of the library — about one in sixty documents — could be corrected by fixing metadata alone; nearly everything else needs structural work.
Two corrections
A research page that shows its own corrections is more trustworthy than one that quietly edits an old number. Both of these were wrong in a way that would have reached a customer had either been the final figure.
The first count of scanned pages was a sample, and the sample understated it by more than half. Probing a handful of pages per document and labelling the whole file from that put the scanned share at roughly 3%. Opening and reading every single sheet in the corpus — not a sample — put the true figure at 8%, the number reported above. A 500-page document with scanned exhibits in the middle is not well described by three pages of it.
A table breaking pages down by document type was published, then found to describe only each file's first page. Once it became clear how many files mix content, as described above, the original table's numbers were quietly wrong for the same reason: a file typed by its first page can carry pages of an entirely different kind further in. The version above is the corrected one, and the earlier one should not be quoted.
What a person judged
Software found problems and sorted documents by type. A person checked a sample of both, by hand, and this is what that checking found.
The document-typing software eventually reached roughly four in five correct on the documents a person could check it against — but that accuracy figure covers only about a quarter of the library's pages. Most of the rest were typed from a web address or a title alone and never checked against a person's judgment at all; a further one in six pages received no type at all, marked unknown rather than guessed. A small hand-checked sample of the untested, web-address-typed majority found the software agreeing with its own product vendor's naming convention rather than with what the document's own words said it was — a meeting notice typed as an agenda because of where it was hosted, when its own title block named it something else entirely. That is not a rate; it is one honest look at work nobody had checked before, and it found the same fault it had already found once.
Recognition software that reads the picture pages reported itself, on average, more than 97% confident in what it recovered. Checking a handful of its output by hand found two documented misreadings, both at full, unqualified confidence: a date read wrong by one letter — June misread as Jule — and a short two-word phrase on the city's own letterhead read wrong in both words. High confidence is what a wrong reading looks like from the inside — a machine grading its own work is not the same thing as a person confirming it, which is exactly why the technical standard this library was measured against treats this kind of error as something only a person can catch.
A first description of the library's largest files called some of them "several documents bound together." Checking the actual claim against one of the largest files found that description wrong: the file's pages genuinely are not alike, but that does not make it more than one document — it is one document whose sections cover very different ground. The measurement (the pages differ) was right; the description of what that meant was not, and it was corrected once someone actually opened the file rather than trusting the pattern.
What was not assessed
The technical standard this library was checked against defines dozens of ways a document can fail that no software can decide at all — the standard's own authority, not a limit of our tools. None of the following was assessed on any document in this library, by us or by anyone, and none of it is implied by anything above:
- Whether text described as alternative text for an image actually describes that image, rather than merely existing.
- Whether the order a screen reader would encounter a document's content in makes sense read aloud, rather than merely being technically valid.
- Whether a heading is tagged at the level it is because of its meaning, rather than because of how large the text looked.
- Whether content marked decorative genuinely is decorative, rather than real information hidden from a screen reader.
- Whether the language a document declares itself written in is the language it is actually written in.
- Colour contrast anywhere in this library — a requirement of the accessibility standard the law actually requires, with no automated check in the standard used to measure the rest of this page at all.
A small number of documents in this library have been listened to, start to end, with a screen reader. That count is small relative to the size of the library, and it is stated here as exactly that — a beginning, not a validation of everything above. The structural figures on this page describe what software and a person reading tags and structure can see; they are not a substitute for what a person using a screen reader actually experiences, and the two are reported separately on purpose.
Six documents in the library could not even be opened for a machine check at all — five would not parse, one raised an unexpected error. They are recorded as unassessed, not as passing and not as failing: a file that could not be read has not been found wanting.
This is one city, on one common government website platform. A second agency's library has not been measured, and until one has, none of this should be read as a description of municipal document libraries generally — only of this one, in full.
Who did this, and why
Recordstake counts a public agency's published documents, tests them against the accessibility standard the ADA's Title II rule points to, and reports what was measured, what a person judged, and what nobody has assessed at all — never blurring the three. This page is that same method, run once over an entire library and published without a sale attached. Read what the free scan and the paid Inventory do, or read the same kind of audit, of our own site.
Found something wrong on this page, or want the underlying method for your own library? Email contact@recordstake.com.