About Blog Contact Links Vault
Latest
Home / QA Frameworks, Tools, & Debugging / How to Run WCAG Accessibility Testing Without Letting the Tool Redesign Your Site
QA Frameworks, Tools, & Debugging
14 min read · October 5, 2026 · 18 views

How to Run WCAG Accessibility Testing Without Letting the Tool Redesign Your Site

Accessibility checkers flag things that are technically fixable and practically a bad idea. Here is the order I would run WCAG testing in, from an axe core scan to keyboard, NVDA and contrast checks, and how to tell which flags are real.

Share:

Accessibility checkers flag things that are technically fixable and practically a bad idea, and WCAG testing mostly comes down to knowing which is which. If you have ever run a scan, watched a wall of errors appear, and wondered whether you were supposed to change the site or ignore the tool, this post is for you. Every checker hands you a list, and nothing on that list tells you which items are real.

I run WAVE and Lighthouse on my own sites, and my gripe is the same with both of them. If you are not looking for something specific, the output is noise, and if you know what you are looking at, WAVE is a better tool than DevTools for this job. DevTools hands me warnings about font sizes and deprecated features the site may or may not even be using, while WAVE shows more errors than that, and a few of them look like they make no sense until you work out what the rule is protecting.

This post is the order I would run WCAG accessibility testing in: the standard first, then an automated scan, then the keyboard, a screen reader, contrast, and finally the bug report. Each step is one check you can actually perform, and each one tells you something the others cannot.

What WCAG Actually Asks a Tester to Check

WCAG stands for the Web Content Accessibility Guidelines, the W3C standard that most accessibility requirements point to. It is organized around four principles, perceivable, operable, understandable and robust, and every success criterion under them carries a conformance level of A, AA or AAA. When a client or a contract says WCAG compliance, they almost always mean level AA, which is the line this post tests against.

The version matters because WCAG 2.2 builds on 2.1 instead of replacing it. A site that meets WCAG 2.1 AA still has gaps against WCAG 2.2, which added criteria such as focus not being hidden behind sticky content, a minimum size for tappable targets, and accessible authentication. If you are choosing a target, test against WCAG 2.2 AA, which keeps you ahead of any requirement written for WCAG 2.1 AA. You can read every criterion in the official WCAG quick reference.

As a tester you do not memorize the full list. You map the criteria to checks you can perform: text alternatives for images, contrast, keyboard operation, visible focus, labels on form fields, a sensible heading and landmark structure, error messages that explain themselves, and text that survives zoom. Everything below is one of those checks, and together they are what WCAG testing looks like in practice.

Automated Accessibility Testing: What It Catches and Where It Stops

Automated accessibility testing tools can only flag what a rule can decide. A rule can see that an image has no alt attribute, a form field has no label, a button has no name, two elements share an ID, or text fails a contrast ratio. It cannot decide whether the alt text describes the picture, whether the tab order makes sense, or whether the page still works when you can only hear it. The Playwright accessibility documentation says as much, noting that many accessibility problems only turn up in manual testing, and that is the right way to treat any accessibility checker you run.

Lighthouse Accessibility Testing Has the Same Blind Spots

Lighthouse accessibility testing runs on axe core underneath, so a score of 100 means every rule it can check passed, not that the page is accessible. I have written about that gap before. In my PageSpeed and Lighthouse debugging post I put it as QA doesn’t stop when a tool says pass, and my agentic browsing audit post ended up in the same place from the other direction, that no score or audit ships with the judgment built in. An accessibility score is a claim to investigate, and the investigation is the work.

WAVE Accessibility Checks Are Noisy Until You Know What You Are Looking At

My honest take on both WAVE and DevTools is that they are useless if you are not looking for something specific, and better than almost anything else once you are. The difference is what the noise is about. DevTools complains about font sizes and deprecated features as a code quality matter, while WAVE is asking whether a person can actually read and use the page, which is why its list is longer and why some of it needs interpreting.

Font size is the example I keep coming back to. A small font really is a readability problem, which is exactly why an accessibility tool cares about it. But if you enlarge the text only to make the flag disappear, you end up shipping a different site than the one you designed, and you have fixed the tool’s complaint instead of the reader’s problem. The right move is to look at the text the flag points to, decide whether a person would struggle to read it, and then change the design on purpose instead of by reflex.

That is the rule I hold myself to with every accessibility checker. A flag is a question, not an order, and WAVE puts a lot of questions on one screen: errors, contrast errors, alerts, features, structural elements and ARIA. Read the alerts and the structure view as well as the red errors, because a page can show zero errors and still have a heading outline that makes no sense.

Axe Core in Playwright and Cypress for Automated WCAG Testing

Axe core is the open source rules engine from Deque, and most of the accessibility testing tools you will meet are wrappers around it. Lighthouse uses it, Deque publishes a Playwright package for it, and Cypress has a community plugin. Once you understand how axe reports violations you understand all of them, so learn it once and use it where your tests already live.

In Playwright you install the package, import AxeBuilder, load the page, and call analyze. The result has a violations array, and an empty array is your pass condition. The scan only sees the page as it is at that moment, so if the thing you care about is a menu or a modal, you have to open it and wait for it to appear before you call analyze, otherwise axe never sees the elements you wanted checked. If you are new to the framework itself, my Playwright guide for QA testing covers the basics first.

npm install --save-dev @axe-core/playwright
import { test, expect } from '@playwright/test';
import AxeBuilder from '@axe-core/playwright';

test('homepage has no detectable WCAG A or AA violations', async ({ page }) => {
  await page.goto('https://your-site.com/');

  const results = await new AxeBuilder({ page })
    .withTags(['wcag2a', 'wcag2aa', 'wcag21a', 'wcag21aa'])
    .analyze();

  expect(results.violations).toEqual([]);
});

By default axe runs a wide set of rules, including best practice rules that no WCAG criterion requires. The withTags call narrows the scan to rules tied to WCAG success criteria, which is what you want for WCAG compliance reporting. The tags above cover WCAG 2.0 and 2.1 at levels A and AA, and newer axe core releases also carry a tag for WCAG 2.2 AA, so check the tag list in the axe documentation for the version you install.

Every real site has known issues, and axe gives you two ways to park them. The exclude call removes an element and everything inside it from the scan, and the disableRules call switches off a rule entirely using the rule id from a violation. Both are temporary parking and not fixes, and exclude is the more dangerous of the two because it skips every rule on those elements, not only the one you already know about.

In Cypress the plugin is cypress axe, which needs axe core and Cypress installed beside it. You import it once in your support file, visit a page, inject axe into it, and then call checkA11y wherever you want a scan. Because checkA11y runs against the page at the moment you call it, you can click something first and scan the result, which is how you catch problems that only exist after an interaction. My Cypress automation guide covers the setup around it.

npm install --save-dev axe-core cypress cypress-axe
// cypress/support/e2e.js
import 'cypress-axe'

// in your spec file
describe('homepage accessibility', () => {
  beforeEach(() => {
    cy.visit('https://your-site.com/')
    cy.injectAxe()
  })

  it('has no detectable WCAG A or AA violations', () => {
    cy.checkA11y(null, {
      runOnly: { type: 'tag', values: ['wcag2a', 'wcag2aa'] }
    })
  })
})

If you are adding this to a site that already has a pile of violations, the includedImpacts option lets you fail only on critical issues at first, and skipFailures logs violations without failing the build. Treat both as a ramp and not a destination. A suite that never fails on accessibility is just a log file, and nobody reads those.

Manual Accessibility Testing: Keyboard Accessibility First

Keyboard accessibility is the cheapest manual accessibility testing you can do, because it needs nothing installed. Put the mouse aside and use Tab and Shift Tab to move through the page, Enter and Space to activate things, Escape to close things, and the arrow keys inside menus and tabs. You are checking that everything interactive can be reached, that focus is visible at every step, that the order follows what you see on the screen, and that you never get stuck somewhere you cannot leave.

These checks map to concrete WCAG criteria: keyboard operability, no keyboard trap, focus order and visible focus, and WCAG 2.2 adds focus not being hidden behind other content. That last one is where a sticky header or a cookie banner catches people, because focus lands on a link that is sitting underneath the header. Modals deserve their own pass, since focus should move into the dialog, stay inside until it closes, and return to the control that opened it.

A scanner misses most of this because it checks the markup, not the experience of moving through the page. If you already run timeboxed sessions like the structure in my exploratory testing session post, an accessibility pass fits inside one as a charter. It is also where you find the problems that matter most, because a keyboard trap on a checkout form is not a style issue, it is a person who cannot buy anything.

Screen Reader Testing With NVDA

I need to be straight about this section. I have used NVDA, but I have not used it to test a project, so what follows is how it works and what to listen for, not a field report. My gripe with it is the same one I have with WAVE. If you are not listening for something specific, a screen reader is just a voice reading your page at you, and if you know what you are listening for, it tells you things DevTools never will.

NVDA is a free screen reader for Windows from NV Access, and it works with Chrome and Firefox. Start it, open your page, and use the browse mode shortcuts to move by structure instead of reading from the top: H for the next heading, D for the next landmark, F for the next form field, B for the next button, and K for the next link. The Control key stops it talking, and the Speech Viewer in the Tools menu shows what it says as text, which makes it much easier to take notes and attach them to a bug.

Listen for the failures a scanner only partly sees. A button that announces as just button has no name, a form field that does not announce its label is a form nobody can fill in without seeing it, and a heading list that makes no sense out of context means the page structure is decoration. Also listen for what happens after an interaction, because a message that appears on screen and is never announced is invisible to anyone who is not looking at the screen.

NVDA covers Windows, and phones have their own screen readers, VoiceOver on iOS and TalkBack on Android. Screen reader testing on a phone is a separate pass, and my post on testing a website on mobile is the reminder that an emulator saying everything is fine is not a result you can trust.

Use a Color Contrast Checker, Then Look at the Page

A color contrast checker gives you a ratio, and the WCAG AA requirement is 4.5 to 1 for normal text and 3 to 1 for large text, with 3 to 1 for interface components such as input borders and icons. DevTools shows the ratio in the color picker, WAVE reports contrast errors on the page, and the WebAIM contrast checker lets you try colors before you commit to them. Use whichever you like, but check text over images and hover and focus states by hand, since scanners often cannot determine the background there.

When a contrast check fails, change the color and leave the layout alone. This ties straight back to the WAVE font size point, because the fix that satisfies the tool and keeps the design is usually a darker shade of the same color, not bigger and heavier everything. After that, zoom the page to 200 percent and look at it, since text resize and reflow are the kind of criteria a rule cannot judge for you.

That visual check is the same judgment problem I wrote about in visual and cross browser regression, where the tool flags a difference and a person has to decide whether it is real. Accessibility testing has the same shape. The tool finds candidates, and you decide which ones are defects.

How to Log Accessibility Bugs So They Get Fixed

An accessibility bug report needs everything any bug report needs, plus three more things. Include the WCAG criterion it breaks, the tool or assistive technology you found it with, and the selector for the exact element. Axe gives you the rule id, the impact level and a target selector for each affected node, and that selector is worth more to a developer than a screenshot.

Then call severity honestly. A keyboard trap on a checkout form blocks someone completely, and a decorative image with a missing alt attribute does not, even though a scanner will list them side by side. I use the split between severity and priority from my severity vs priority rubric here, and the rest of the format follows my post on bug reports developers actually act on.

Where ADA and Section 508 Fit

I am a QA engineer and not a lawyer, so treat this as orientation and not legal advice. The ADA does not come with a technical standard written for websites, which is why WCAG AA is what people actually test against, and the Department of Justice rule for state and local government websites points to WCAG 2.1 AA. Section 508 covers US federal agencies and points to WCAG 2.0 AA, which is why requests for 508 compliance and WCAG compliance get used almost interchangeably in testing work. No widget you paste into a site does this testing for you, so the checks in this post are still what you run.

The order is the point. Read the standard, run an automated scan with axe core, do the keyboard pass, the NVDA pass and the contrast check, and write up what you find against the WCAG criterion it breaks. The tools are noisy until you know what you are looking for, and then they are better than DevTools at this job. Treat each flag as a question, and answer it with the page in front of you.

Share this article:
Jaren Cudilla
QA Overlord

A QA engineer who works in Playwright and Cypress and runs WAVE and Lighthouse against his own sites. He treats every accessibility flag as a claim to check against the page, not a verdict to accept.

Leave a Comment

What is How to Run WCAG Accessibility Testing Without Letting the Tool Redesign Your Site?

Accessibility checkers flag things that are technically fixable and practically a bad idea, and WCAG testing mostly comes down to knowing which is which.