Video camera recording two people speaking into microphones with text about video accessibility

Your Website Passed the Accessibility Audit. Your Videos Probably Didn’t.

Most businesses that take web accessibility seriously get the same things right. Alt text on images. Keyboard navigation. Color contrast. Form labels a screen reader can actually announce.

Then there’s a product video on the homepage with no captions, and a training library where nothing has a transcript.

Video is where accessibility work quietly falls apart. Not because teams don’t care, but because fixing video has historically meant hiring a vendor, waiting days, and paying by the minute of footage. So it gets deferred. And it stays deferred until an audit, a legal complaint, or a customer forces the issue.

That’s the gap Accessify Video is built to close.

What the standards actually require

It’s worth being specific here, because “make your videos accessible” is vague enough to ignore.

WCAG — the international standard for web accessibility — sets out separate rules for video and audio. These are the ones that apply to a typical pre-recorded business video.

Level A, the baseline:

  • 2.2 Captions. Captions that stay in sync with the audio, for every pre-recorded video with sound.
  • 2.3 Audio Description or Media Alternative. Either a spoken description of the visuals, or a full written alternative that covers them.

Level AA, the standard most organizations are held to:

  • 2.5 Audio Description. A spoken description track for every pre-recorded video. The written-alternative option from Level A no longer counts.

That last point matters more than most teams realize. Level AA is the standard named in most legal settlements, buyer requirements, and public sector rules. And at Level AA, captions alone don’t get you there.

If your video shows anything that isn’t said out loud — a chart, a figure on screen, a demonstration, text on a slide — you need audio description.

Captions serve people who can’t hear the audio. Audio description serves people who can’t see the screen. They are not substitutes for one another, and audio description is the one almost everyone skips.

(Live video has its own separate rule, 1.2.4. Everything below concerns pre-recorded video.)

Why automatic captions don’t count

Every major video platform will generate captions for free. It’s reasonable to assume that handles the caption requirement.

It doesn’t, and the reason is accuracy. Speech recognition struggles with exactly the things business video is full of: product names, industry terms, acronyms, accented speech, people talking over each other, and anything recorded in a room with an echo. It also tends to skip punctuation and speaker names, which matters enormously when more than one person is talking.

People who rely on captions have a blunt word for the results — “craptions” — and the standard is unforgiving on this point. Captions have to be accurate to count as captions. A rough approximation of what someone said is not an accessible alternative to what they said.

So the useful question about any captioning tool isn’t whether it produces something automatically. It’s what happens after it does.

What good audio description actually sounds like

Writing audio description is a real craft, and it’s worth knowing what you’re aiming for before you judge any tool that does it.

Description has to fit into the natural gaps in the dialogue, so it’s limited by time. You often have three or four seconds. It has to be factual rather than interpretive — describe what’s shown, not what it means. And it has to be ruthlessly selective, because you cannot narrate everything, and the things worth narrating are the ones a sighted viewer would need in order to follow along.

Here’s what that looks like in practice. Imagine a four-second gap after a narrator says “…and that brings up your dashboard.”

First pass: “A dashboard appears showing a bar chart with five colored bars of increasing height, labelled January through May, above a summary panel containing three statistics.”

Edited: “A bar chart appears. Sales rise from January to May.”

The first version is accurate and unusable — it needs about fifteen seconds and you have four. The second fits, and it carries the one thing a viewer actually needs.

That edit took ten seconds to make and would have taken several minutes to write from nothing. Reviewing a draft is a fundamentally easier job than starting from an empty timeline. That’s the shift a good tool gives you.

[TO CONFIRM] Replace the illustration above with a real generated description from Accessify Video and its edited version.

What Accessify Video does

Accessify Video handles the full set of video requirements from a single upload, then hands you the controls.

Upload the video. No preparation, no separating out the audio, no converting formats first.

Get a transcript and captions in minutes. Both are produced automatically, so the wait is closer to a coffee break than a vendor’s turnaround time.

Get audio descriptions for the visuals. This is the part that’s normally a specialist service. The tool works out what’s happening on screen and writes descriptions for it.

Get the spoken audio to go with them. A written description still has to become something people can hear. Accessify Video produces the audio, so you’re not sending a script to a voice artist and waiting on a recording session.

Run the whole thing in one click. If you have volume to get through, the full process can run end to end without stopping.

Edit anything before you finish. The transcript, the captions, and the descriptions are all editable. Fix a mangled product name, cut a description that’s too long for the gap it sits in, add speaker names, correct a timing offset.

That last one decides whether any of this is genuinely usable. Automation gets you most of the way in a few minutes; a person closes the gap in a few more. What you want is exactly that split — not a sealed box you can’t correct, and not an empty editor that makes you start from scratch.

What you get at the end

[TO CONFIRM] This section needs the real answer before publishing. Readers ask it second, right after “does it work”, and an unanswered version of it costs more sign-ups than pricing does. It should cover:

  • Which files come out (caption files, transcript, description script, description audio)
  • What formats they’re in
  • How you get them onto your video, wherever it’s hosted

Who this is for

Teams with a backlog. If you have a library of training videos, recorded webinars, or product demos that have never been captioned, the volume is the obstacle. Automatic generation with human review is the only realistic way through a backlog that size.

Anyone fixing problems after an audit or a complaint. Video is often the biggest single item in an audit report, and the slowest to clear. Compressing that timeline changes what you can promise.

Organizations that have to prove accessibility to a buyer. If you’re producing a formal accessibility report for a customer or a government body — an ACR or a VPAT — the video rules get checked directly. Audio description is where these reviews usually find the gap.

Anyone publishing video regularly. Fixing video after the fact always costs more than doing it at publication. A tool fast enough to sit inside your normal workflow is what makes accessibility routine instead of a project.

The part that isn’t about compliance

Compliance is why most businesses start looking at this, which is why this post leads with the standards. But it undersells the case badly.

A study by Verizon Media and Publicis Media, which surveyed 5,616 US consumers, found that 80% of people are more likely to watch a video all the way through when captions are available. The same research found 69% watch video with the sound off when they’re in public.

And roughly 80% of the people using those captions have no hearing impairment at all. They’re commuting, sitting in an open-plan office without headphones, watching in a second language, or simply reading faster than they listen.

Transcripts do something similar on the search side. A page with a published transcript is searchable, quotable, and readable by search engines in a way that a bare video file is not.

This is the pattern across all accessibility work: something built for a specific group turns out to serve a much wider audience. Video is one of the clearest examples there is.

Try it

Accessify Video is live at accessifyvideo.com. Upload a video and see what comes back.

[TO CONFIRM] Say what happens when someone signs up — free tier, trial length, whether a card is needed. “See what comes back” asks people to click without telling them what it costs them.

If you’d rather have the work done for you, or you need a formal audit, an accessibility report for a buyer, or human-verified descriptions for a large library, that’s what our audio description services and web accessibility audit and support teams handle. Book a meeting with our accessibility team and we’ll work out which route fits.

Frequently asked questions

Do videos legally need audio description?

At WCAG Level AA, yes, for pre-recorded video. Criterion 1.2.5 requires a spoken description track. Level AA is the standard named in most legal settlements and public sector requirements, so in practice it’s the bar most organizations are held to.

Do YouTube’s automatic captions meet WCAG?

Generally not. Captions have to be accurate to count, and speech recognition reliably struggles with product names, acronyms, accented speech, and people talking over each other. Automatic captions are a useful starting draft, not a finished one.

What’s the difference between captions and audio description?

Captions are text of what’s said, for people who can’t hear the audio. Audio description is a spoken account of what’s shown, for people who can’t see the screen. Meeting Level AA needs both.

Do I need audio description if my video has no important visuals?

If everything shown is also said out loud, there’s nothing left to describe. That’s rare in practice — charts, on-screen text, and demonstrations all count as visual information that isn’t spoken.

Is a transcript the same as captions?

No. Captions are timed to the video and appear on screen as it plays. A transcript is the full text in one block, usually published alongside the video. Level A lets a written alternative stand in for description, but Level AA does not.

What about live video?

Live video falls under a separate rule, 1.2.4, which requires real-time captions. Accessify Video covers pre-recorded video.

Share: