Skip to content
Ending soon50% off every plan — ends 31 OctoberStart free
Free deep sweep · no signup

Every PDF on your site, brought into the light.

Each contact on the scope is a document we found on your domain. We report what it weighs, how long a real person waits in the dark for it, and whether a screen reader could read a word. Nothing is uploaded. Nothing is converted. We just look.

robots.txt respected · a minute or two · the file never leaves your server

awaiting a domain
contacts
—
examined
—
total mass
—
heaviest
—
silent to readers
—
pages surveyed
—
Instruments

Four readings per document.

Everything below is measured, never estimated. If a sweep stops early we tell you how far it got.

01 · Legibility

Can a screen reader find its way in?

We read each file's structure tree — the thing that tells assistive technology which text is a heading and what order to read a page in. Without it, a document is an undifferentiated wall of text, and tables become unreadable.

02 · Mass

How long people stare at nothing.

A PDF shows nothing while it downloads: no progress, no first paragraph. We report each file's size as the wait it causes on a normal mobile connection, because seconds are what a reader actually experiences.

03 · Orbit

Which PDFs are acting as landing pages.

A PDF listed in your sitemap can rank and take traffic — into a file with no header, no navigation and no way back to your site. We flag the ones you have submitted to search engines.

04 · Dark matter

Everything you'll never learn about them.

A downloaded file leaves your analytics the moment it is saved. Nothing tells you whether it was read, shared, or abandoned on page two.

Flight plan

A polite little expedition.

Four steps, a minute or two, only where your robots.txt says we may go.

  1. 01

    We ask permission

    Your robots.txt tells us where the sitemaps live and where you'd rather we didn't. We obey it, and we slow down when it asks us to.

  2. 02

    We follow the sitemaps out

    The fastest way to see what is genuinely published, nested indexes included.

  3. 03

    We drift through a few pages

    Your domain only, a handful at a time, for about a minute. Your server will not notice us.

  4. 04

    We coax each PDF into telling us about itself

    A range request lifts a few kilobytes from each end of the file — enough for version, size, language and structure. The document itself stays exactly where it is.

We report what we measured, and nothing else. If a scan stops early, the report says how many pages it covered rather than estimating a total we cannot stand behind.

Ground control

Questions before launch.

Do you download or store my PDFs?+

No. We read a few kilobytes from each file using HTTP range requests — enough to see its header, its structure and its size — and never hold the document itself. Nothing is uploaded, stored or converted, and we do not keep copies of anything we read.

How accurate is the accessibility check?+

Where we can read a document's catalog, the answer is definitive: either it has a structure tree or it does not. Some PDFs store that information in a way we cannot reach from a partial read, and those are reported as "couldn't tell" rather than counted as failures. We would rather under-report than tell you a document is inaccessible when we could not actually check.

Why does it say it only checked some of my pages?+

Every scan is capped at two minutes and under a thousand requests, so we are polite to your server and quick for you. On a large site that means we see a sample, and the report says exactly how many pages we looked at. We never estimate a site-wide total from a sample — a guess printed next to measured facts makes the facts worth less.

It found no PDFs, but I know we have some. What happened?+

Three likely reasons, and the report tells you which. If it says the site is behind a bot challenge, a firewall is asking every automated client to prove it is a browser — running it again will not help, but allowing LivingPageAudit/1.0 will. If it says requests were rate-limited, waiting a few minutes and trying again usually gets further. Otherwise your PDFs probably sit deeper in the site than a one-minute scan reaches, or behind a search form or login, which no crawler can see. A scan that stops early always says so; we never report "no PDFs" as though it were a clean result.

What does your crawler identify itself as?+

LivingPageAudit/1.0, with a link back to this page in the user agent, so anyone reading their server logs can see exactly who called and why. We do not disguise ourselves as a browser. That means some sites will block us — we would rather be blocked and tell you so than get results by pretending to be someone else.

Will this slow down or overload my website?+

No. We obey your robots.txt, identify ourselves in the user agent, fetch at most a few pages at a time, and stop after roughly sixty seconds. If your server pushes back with a 429, we drop to one request at a time, honour the Retry-After it sends, and give up rather than keep hammering. A scan is a fraction of what an ordinary search engine crawl does.

Am I legally required to fix these?+

We cannot tell you, and we would rather say so than guess. Many countries set accessibility requirements for documents published to the public, and which ones apply to you depends on where you operate, what sector you are in, how large you are and who your readers are — none of which a crawler can see from the outside. This tool reports the state of your files against WCAG, the technical standard most of those rules point at. What that means for your obligations is a question for your own advisers.

Do I need an account?+

No. Enter a URL and read the report. There is no signup, no email required, and the whole report is visible to anyone who runs it.

Bring them into the light.

Living Page turns a PDF into a page-turning reader that opens instantly in the browser, keeps your branding, and tells you how far people actually read.

Start free