012 min
Why a clean scan does not finish an accessibility evaluation
In 2024 the digital team of the British Columbia government checked its own website. Most of the site passed an automated test with a Lighthouse score of 100%. Lighthouse is a free tool built into the Chrome browser. It runs a fixed set of accessibility checks and reports one score. The team did not treat that score as the answer. It also audited 39 pages by hand against 56 success criteria from WCAG 2.2, and the audit found 13 that the site did not yet meet. A success criterion is one testable requirement in the standard, such as a minimum contrast between text and background.
The hand audit showed two kinds of failure. Some depended on the context of a page, and others were usability problems. No score from a program would have listed either kind.
That gap is why accessibility evaluation is a practice and not a single tool run. A score answers a narrow question: did the checks that a program knows how to run pass? An evaluation asks a wider one. Which requirements of the standard does this product fail, and can people with disabilities really use it? The standard behind most evaluations is WCAG, and the goal behind it is accessibility. The first question is cheap to answer. The second needs several methods, and the sections below follow them in the order a team usually meets them.
How common failures are
The scale of the problem shows why the cheap question is not enough. In 2025 WebAIM, an accessibility organization, ran automated checks on the home pages of the top one million websites. WebAIM detected WCAG 2 failures on 94.8% of those home pages. It averaged 51 detected errors per page, and low-contrast text was the most common, found on 79.1% of pages.
A high score is not a conformance claim
A score that looks good is easy to misread as "this product conforms". The British Columbia example shows why that reading is wrong. A Lighthouse score is built only from the checks Lighthouse runs, so it says nothing about the criteria those checks do not cover. WebAIM makes the same point from the other side. A page with no detected errors is not thereby accessible or conformant. A passing score means "no failures found by this tool", and the evaluation exists to find the failures the tool cannot see.
022 min
What automated checks find in an accessibility evaluation
Automated checks are the first layer, so it helps to know exactly what they do. A checker loads a page, reads its code and its rendered result, and tests both against rules that a program can decide. Is there an image with no text alternative? Is a form field missing its label? Is the contrast between text and background lower than the ratio the standard sets? Common examples are WAVE, axe and Lighthouse. The W3C keeps a long list of tools that can be filtered by what they test, which standards they cover and where they run, from a browser extension to a build pipeline. It also publishes shared test rules, called ACT Rules, so that different tools can reach the same verdict on the same check.
The strength of these tools is speed. A scan covers every page in minutes, costs almost nothing and can run on each code change. A failure that a developer introduces on Tuesday can be caught on Tuesday.
What a scan misses
Vendors rarely say how much a scan leaves out, so in 2017 the UK Government Digital Service measured it. Its accessibility team built a page with 143 known barriers and ran 10 free tools on it. The best result counted only errors and warnings, and it belonged to Tenon, which found 37% of the barriers. The weakest of the ten, Google Developer Tools, found 17%. The more telling number is the barriers that no tool reported: 42 of the 143, or 29%. They included long passages in italics, tables with empty cells and links marked by colour alone.
Why tools cannot judge meaning
The cause is built into the method. A program can check that an image has an alt attribute, but it cannot judge whether the words describe the image. It can see that a link has a different colour, but not whether a reader who cannot tell the colours apart still knows it is a link. Anything that needs a judgment about meaning, order or context is left for a person.
Using several tools together
The team also found that running several tools together caught more barriers. The cost was higher effort for the teams, because the tools report in different ways and need separate set-up. So combining tools raises coverage, but it does not remove the 29%.
033 min
Expert review in accessibility evaluation
The W3C states that no tool alone can decide whether a site meets accessibility standards, and that knowledgeable human evaluation is required. So the second layer is a trained person using the product directly.
A good place to begin is the W3C's Easy Checks. They are thirteen quick looks, among them image text alternatives, page title, headings, colour contrast, a skip link, visible keyboard focus, page language, zoom, captions, transcripts, audio description and form labels. They are meant as a first look.
A full expert review goes further and covers what no tool can settle:
- Keyboard only. Move through every control with the Tab key. Check that the order makes sense, that the focus mark is always visible, and that nothing needs a mouse.
- Screen reader. Listen to the page with NVDA, JAWS or VoiceOver. Check that names, headings and changes of state are announced.
- Zoom and text size. Enlarge the page to 200% and 400%. Check that nothing is cut off or overlaps.
- Colour. Measure text contrast against the standard's ratios, 4.5:1 for normal text and 3:1 for large text, and check that colour is never the only carrier of meaning.
- Meaning. Read the alt text, link text, labels and error messages, and decide whether each says the right thing.
The evaluator needs to know the standard, accessible design and code, assistive technology, and how people with disabilities use the web. That is why the W3C describes full conformance evaluation as work for experienced people.
What a pass on the Easy Checks means
A page that passes all thirteen can still have serious barriers, because the checks cover only a few issues and were designed to be quick. Treat a pass as permission to move to the next layer, and a failure as a fix to make before spending expert time.
Judging by criteria or by barriers
Experts also differ in what they write down. The usual method, conformance review, goes through the standard's criteria and marks each as pass or fail. In 2008 Giorgio Brajnik compared it with another method, the barrier walkthrough. In a barrier walkthrough the evaluator lists the barriers that a given group of users would meet and rates how serious each one is. Brajnik found significant differences in correctness between the two. The choice matters in practice. A list of failed criteria tells a developer which rule was broken, while a list of barriers tells them what a user could not do.
When the evaluator is new to assistive technology
One caution applies to the reviewer. Using a screen reader well takes training. An evaluator who has used one for a day can miss a control that works, or report a failure that comes only from how they used the tool. The W3C's testing notes warn that this creates false beliefs about what users experience. Confirm a suspected failure on a second browser and screen reader pair, or ask an experienced user, before it goes into a report.
042 min
Testing accessibility with disabled users
The third layer puts people with disabilities in front of the product with real tasks. The W3C explains what this adds: sessions with users find usability problems that a conformance evaluation does not find. Say a booking page passes every check on its form, yet a screen reader user hears an error message only after leaving the field and cannot tell which entry caused it. No single criterion counts that, and a session shows it in minutes. This is the same kind of work as usability testing, with participants who use assistive technology.
Several habits make the sessions useful:
- Run an expert review first. The W3C advises removing the significant barriers before sessions, so the participants' time goes to what only they can show.
- Keep sessions small and frequent. Informal sessions throughout development work better than one large study at the end.
- Let people use their own setup. They know their own tools, and a borrowed machine tests the setup instead of the page.
- Pay participants and include different levels of experience. A novice and an expert with the same screen reader meet different problems.
- Watch what people do, not only what they say. When something fails, work out whether the cause is the page, the user's setup, the browser or the assistive technology.
Why user testing is not the whole evaluation
A few participants cannot stand for everyone. A blind person who uses one screen reader does not speak for a person with low vision or a motor impairment, and a session covers only the tasks the team chose. The W3C says that involving users with disabilities has many benefits but cannot, on its own, decide whether a website is accessible. The sessions work as a check on the expert review and the scans, never as a replacement for them.
052 min
The WCAG-EM procedure for accessibility evaluation
Scans, expert passes and sessions give findings. To turn them into a statement about a whole product, evaluators use WCAG-EM, the W3C's evaluation methodology for WCAG. It applies to websites, mobile apps and kiosks, and it has five steps:
- Define the scope. Decide which parts of the product are in, which WCAG version and conformance level (A, AA or AAA) is the target, and which browsers and assistive technologies must be supported.
- Explore the product. Learn its types of page, its technologies and the functions people depend on, such as search and sign-in.
- Select a representative sample. Choose a structured set that covers each type of page and function, and include complete processes, such as a whole checkout from basket to confirmation. Then add a random set.
- Evaluate the sample. Check every sampled page against each success criterion in the chosen level, using the tools and manual checks described above.
- Report the findings. State the scope, the result for each criterion, and for each failure where it is, how to reproduce it and how to fix it.
The W3C provides an open-source report tool that records the results and produces the report.
Why part of the sample is random
A person picks the structured sample, so it reflects what that person thought mattered. The random set is a check on that choice. The method adds randomly chosen pages equal to 10% of the structured sample, so 8 more for a sample of 80. If the random pages show the same kinds of failure as the chosen ones, the sample is a fair picture of the product. If they show new kinds, the structured sample missed part of the product, and the evaluator goes back to step 3 and widens it.
Why the sample includes whole processes
Some failures appear only across steps. Say a sample holds the first page of a payment form but not the confirmation page. An error message that the screen reader never announces on the last step would go unseen. That is why WCAG-EM asks for complete processes, from the first step to the last, and not only single pages.
What a sample cannot prove
Even a good sample is only a sample. A conformance claim for a whole website cannot rest on an evaluation of a chosen subset of its pages, because unfound failures may remain on pages nobody opened. In practice a report says which pages and processes were evaluated and what was found. It does not say that the site conforms. A full claim needs every page evaluated, or a build process that guarantees each requirement.
062 min
Putting accessibility evaluation into a project
Teams rarely run these methods once. Picture a product team building a checkout flow. They run automated checks on every code change, so the cheap failures never reach review. Before each release candidate, a designer or engineer who knows the standard does the keyboard, screen reader, zoom and contrast passes by hand on every step of checkout. Then the team invites a few disabled customers for sessions on the real task of buying something. Before launch, and again at regular intervals, an experienced evaluator runs a WCAG-EM evaluation, so the team can say publicly what was checked and what was found.
| Method | Finds well | Misses | Best moment |
|---|---|---|---|
| Automated scan | Missing labels, empty links, many contrast failures, on every page | Whether text is meaningful, order, context | Every code change |
| Expert review | Keyboard problems, screen reader problems, zoom failures, unclear wording | Pages outside the sample; depends on skill | Before each release |
| Sessions with users | Real task failures and usability problems | Few people and few tasks | Throughout, after expert review |
| WCAG-EM evaluation | A statement for each criterion on a defined sample | Pages not sampled | Before launch, then regularly |
Start early. The W3C's testing notes point out that fixing problems found late takes more work than doing the job right at the start. Evaluation finds what is broken, and inclusive design aims to stop it breaking, so the two work best together.
When the findings come in, rank them by what they stop a person from doing. A failure that blocks a keyboard user from paying comes before a skipped heading level. Look for causes shared across pages, such as one button component with no visible focus mark, and fix it once. The decision to watch for is where each method sits. Put the cheap method where it runs most often, and put the costly ones where they find what the cheap one cannot.
071 min
Reading an accessibility evaluation from a vendor
Often a team does not evaluate a product itself but buys it. The vendor then hands over an Accessibility Conformance Report, or ACR, completed on the VPAT template that the Information Technology Industry Council publishes. It lists, criterion by criterion, whether the product supports the standard.
A report is a claim, not proof
The accessibility consultant Adrian Roselli notes that the existence of such a report does not show that the product is accessible, or that the report is accurate. So the buyer evaluates the report. Ask how the vendor tested. A report that relies only on automated tools, or on the vendor's own product knowledge, is weak. Look for recent browser, operating system and assistive technology versions that match your users. Be cautious if testing used one browser or one screen reader, or never mentions voice control. Then run your own checks on the parts of the product your people use most.
081 min
Accessibility evaluation vs related practices
Several practices overlap with accessibility evaluation, and the names are used loosely. "Audit", "assessment" and "testing" are often used for the same work. The useful distinction is the question each practice answers.
| Practice | Question it answers | Result |
|---|---|---|
| Not in the library yetAccessibility evaluation | ||
| DesignHeuristic evaluation | ||
| ResearchUsability testing |
A heuristic evaluation can include accessibility rules, but it does not check a standard's criteria one by one. Usability testing becomes part of an accessibility evaluation when the participants are people with disabilities, which is the third layer above. Accessibility is the quality of the product, and the evaluation is how that quality gets measured.
?4 questions
Questions people ask
Which standard should an accessibility evaluation use?
Who should carry out an accessibility evaluation?
How often should a product be evaluated again?
Does the same method work for mobile apps?
§11 sources
Sources
W3C Web Accessibility Initiative, "Website Accessibility Conformance Evaluation Methodology (WCAG-EM)", W3C Group Note.
W3C Web Accessibility Initiative, "Evaluating Web Accessibility Overview".
W3C Web Accessibility Initiative, "Easy Checks - A First Review of Web Accessibility".
W3C Web Accessibility Initiative, "Involving Users in Evaluating Web Accessibility".
Show all 11 sourcesShow fewer sources
W3C Wiki, "Accessibility testing".
WebAIM, "The WebAIM Million: 2025 report on the accessibility of the top 1,000,000 home pages".
WebAIM, "Web Accessibility Evaluation Guide".
Mehmet Duran, "What we found when we tested tools on the world's least-accessible webpage", UK Government Digital Service, 24 February 2017.
Government of British Columbia digital team, "Improving accessibility on digital.gov.bc.ca", 13 September 2024.
Giorgio Brajnik, "A comparative test of web accessibility evaluation methods", 2008.
Adrian Roselli, "How I Evaluate an ACR (VPAT)", January 2026.




