
Ask most people what AI video interview software does, and you will get an answer about facial expressions. The software watches your face, reads your micro-expressions, and decides whether you seem trustworthy.
That was a real product category, and it is largely gone. The major vendors abandoned visual analysis years ago, and the reason is more interesting than the controversy that surrounded it. What replaced it is a pipeline that looks simpler and concentrates its risk somewhere almost nobody is looking.
The question worth asking is not whether the software watches your face. It is what happens to your words on the way to becoming a number.
Exposure is now mainstream. Greenhouse's research across 2,950 job seekers found 63% have faced an AI interview, up 13 percentage points in six months. At the market leader, cumulative video interviews went from 33 million in late 2022 to more than 70 million by October 2025.
Employer understanding has not kept pace with what the products do. One 2024 survey of business leaders had 60% reporting their tools assess tone, language and body language. That claim directly contradicts what the major vendors say they analyze. That gap is worth sitting with: a substantial share of buyers do not know what they bought.
The six stages below describe what a mainstream product actually does between a candidate pressing record and a recruiter seeing a result.

The candidate opens a browser, runs a camera and microphone check, and works through practice questions that are recorded and discarded. Each real question then runs a preparation countdown followed by a response timer, with the number of retakes set by the employer rather than the candidate.
When the response timer expires, recording stops and the file uploads. Nothing is analyzed live. This is an upload-then-process pipeline, not a real-time one, which is why results appear minutes or hours later rather than immediately.
Two design choices here are set by you, not the vendor, and both shape who does well. A short preparation window rewards fluency over consideration. Zero retakes rewards composure over content. Neither is inherently wrong, and neither should be left at the default. Agencies submitting recorded profiles to clients face the same settings question from the other side, which the staffing workflow guide covers in more detail.
Takeaway: review your prep time and retake settings as selection criteria, because that is what they are.
The recording is converted to text by an automatic speech recognition system, usually a third-party one. This is the stage where the pipeline's real risk sits, and it is the stage buyers ask about least.
Headline accuracy figures look excellent, with leading systems now reporting word error rates in the low single digits on benchmark audio. Subgroup accuracy is a different picture. The canonical study, published in PNAS, found average word error rates of 35% for Black speakers against 19% for white speakers across five commercial systems, with over 20% of samples from Black speakers having at least half their words misrecognized. A 2026 peer-reviewed study on a different corpus found the same pattern, narrowed but not closed, at 20% against 15%.
Accent effects are separate and additive. Analysis of Whisper across 696 speakers found significantly higher error rates for non-American English and for speakers of tonal languages. And transcription does not only mishear. A FAccT study found around 1% of transcriptions contained entirely fabricated phrases, with 38% of those hallucinations carrying explicit harms.
This is also why a bias audit of the scoring model alone is incomplete. If the transcript is systematically worse for some speakers, the scoring model can be perfectly even-handed and the pipeline still is not. That is one reason an audit needs to run stage by stage rather than end to end.

Takeaway: ask your vendor for word error rate broken down by speaker group. A single headline figure can't tell you what you need to know.
Visual analysis is gone from the major vendors, and the reason is the most useful fact in this whole subject. HireVue announced in January 2021 that it would not use any visual analysis in its pre-hire algorithms going forward. Fortune's reporting put a number on why: nonverbal data contributed about 0.25% to predictive accuracy in most cases, rising to 4% for customer-facing roles.
Facial analysis was not dropped only because it was contentious. It was dropped because it barely did anything. Tone and pause analysis followed later the same year, and the current position is transcript-only. The company's own page states plainly that its video assessments do not use facial analysis, video, or audio data to evaluate candidates.
Two caveats keep this honest. Smaller vendors still actively market voice analytics measuring pitch, pace and emotional cues, so "the industry has moved on" is true of the large enterprise players and false as a blanket statement. And in the EU, inferring emotions in a workplace context is prohibited outright under the AI Act, applicable since February 2025.
Takeaway: get your vendor's answer on visual and vocal analysis in writing. It is a one-line question and the answers still differ across the market.
Here the products diverge sharply, and the differences matter more than any feature comparison.
Some score with a model trained to reproduce the ratings of trained human evaluators, reporting convergent validity against those raters in the region of .55 to .74. Some train against actual job outcomes such as sales performance or turnover, with reported predictive validity between .25 and .49. A third approach uses a large language model reasoning against a written rubric. A fourth maps transcript language onto a personality inventory.
These fail in different ways, which is why the distinction is not academic. An outcome-trained model inherits the biases of your past hiring decisions. A rater-trained model inherits the biases of the raters and can never exceed the human judgment it was validated against. An LLM grading against a rubric is only as good as the rubric and is the least studied of the three.
One independent critique is worth knowing. Reviewing a major vendor's explainability statement, the Center for Democracy & Technology noted that its competency models were trained on more than 30,000 aggregated video interviews, suggesting models not tailored to specific employers, occupations or industries.
Takeaway: ask which of these four your tool does. If the vendor cannot say, that is the answer to a different and more important question.
The direction of travel at the enterprise end is away from a machine-generated score and toward machine-generated evidence for a human score.
Current products increasingly produce transcripts, recaps, talk-time balance and competency tagging against behaviorally anchored rating scales, with a large language model generating explanations rather than primary scores. One vendor names the specific model it uses for this and states that human teams retain control over every outcome.
This is a meaningful improvement and it introduces a subtler problem. A reviewer reading an AI-generated summary that says a candidate demonstrated limited evidence of stakeholder management is not forming an independent judgment. They are ratifying one, faster. Anchoring is the failure mode of this design, and it does not show up in any audit of the scoring model.
Takeaway: have reviewers score before they read the AI summary, at least on a sample. If the scores move, you have measured your anchoring effect.
The accessibility failures in this pipeline are concentrated in stage two and they are severe. A peer-reviewed study of four commercial speech-to-text systems found mean word error rates of 52.6% for d/Deaf and hard-of-hearing speakers against 5.0% for controls, rising to 85.9% where speech intelligibility was low. At those rates a scoring model is not evaluating a weak answer. It is evaluating noise.
This is now a live legal question. The ACLU of Colorado filed a complaint with state and federal agencies on behalf of a Deaf, Indigenous employee alleging that a video interview platform lacked working captions and that she was rejected with feedback about her communication style. These are allegations, not findings, and the vendor has called them without merit.
The obligations are clearer than the case law. California's rules, in force since October 2025, require employers to consider reasonable accommodations for candidates assessed by automated systems and specifically scrutinize systems examining tone of voice or physical characteristics. Illinois has required since 2020 that employers using AI to analyze video interviews notify applicants, explain how the AI works and what characteristics it evaluates, obtain consent, and delete recordings within 30 days of request.
Takeaway: publish an alternative route before someone has to ask for one. An accommodation a candidate must request mid-process is one most will not request at all.
Export ten completed interviews from your own tool and read the transcripts against the recordings.
Nothing else you do this quarter will be as informative. You will find out how often the transcript diverges from what the candidate said, whether accented or fast speech degrades it, and whether the words your scoring model read were the words spoken. This is a two-hour exercise requiring no vendor cooperation and no budget.
Then pick the ten from the widest range of speakers you can assemble. The average transcript will be fine. The question this pipeline turns on is not what happens on average. It is what happens at the edges, and whether anyone has looked.
Discover fresh insights, trends, and tips on tech talent and offshore development. Stay informed with our latest updates
