Bridge Labs Logo

BRIDGE LABS

Solution

Bridge HR

Follow

LinkedinInstagramFacebookX

Resources

Insights

Use Cases

SMBsBPOsStaffing AgenciesStartupsEnterprise

© 2026 Bridge Labs. All rights reserved. | Privacy Policy | Terms and Conditions

How AI Video Interview Software Works: The 6 Stages

Regine CyrilleRegine CyrilleTechnology26 Aug 2026
Blog Cover_How AI Video Interview Software Works_ The 6 Stages

How AI Video Interview Software Works: The 6 Stages Explained

Ask most people what AI video interview software does, and you will get an answer about facial expressions. The software watches your face, reads your micro-expressions, and decides whether you seem trustworthy.

That was a real product category, and it is largely gone. The major vendors abandoned visual analysis years ago, and the reason is more interesting than the controversy that surrounded it. What replaced it is a pipeline that looks simpler and concentrates its risk somewhere almost nobody is looking.

The question worth asking is not whether the software watches your face. It is what happens to your words on the way to becoming a number.

What is actually running in 2026?

Exposure is now mainstream. Greenhouse's research across 2,950 job seekers found 63% have faced an AI interview, up 13 percentage points in six months. At the market leader, cumulative video interviews went from 33 million in late 2022 to more than 70 million by October 2025.

Employer understanding has not kept pace with what the products do. One 2024 survey of business leaders had 60% reporting their tools assess tone, language and body language. That claim directly contradicts what the major vendors say they analyze. That gap is worth sitting with: a substantial share of buyers do not know what they bought.

The six stages below describe what a mainstream product actually does between a candidate pressing record and a recruiter seeing a result.

AI Video Interview Pipeline: From Recording to Score


1. Capture

The candidate opens a browser, runs a camera and microphone check, and works through practice questions that are recorded and discarded. Each real question then runs a preparation countdown followed by a response timer, with the number of retakes set by the employer rather than the candidate.

When the response timer expires, recording stops and the file uploads. Nothing is analyzed live. This is an upload-then-process pipeline, not a real-time one, which is why results appear minutes or hours later rather than immediately.

Two design choices here are set by you, not the vendor, and both shape who does well. A short preparation window rewards fluency over consideration. Zero retakes rewards composure over content. Neither is inherently wrong, and neither should be left at the default. Agencies submitting recorded profiles to clients face the same settings question from the other side, which the staffing workflow guide covers in more detail.

Takeaway: review your prep time and retake settings as selection criteria, because that is what they are.

2. Transcription

The recording is converted to text by an automatic speech recognition system, usually a third-party one. This is the stage where the pipeline's real risk sits, and it is the stage buyers ask about least.

Headline accuracy figures look excellent, with leading systems now reporting word error rates in the low single digits on benchmark audio. Subgroup accuracy is a different picture. The canonical study, published in PNAS, found average word error rates of 35% for Black speakers against 19% for white speakers across five commercial systems, with over 20% of samples from Black speakers having at least half their words misrecognized. A 2026 peer-reviewed study on a different corpus found the same pattern, narrowed but not closed, at 20% against 15%.

Accent effects are separate and additive. Analysis of Whisper across 696 speakers found significantly higher error rates for non-American English and for speakers of tonal languages. And transcription does not only mishear. A FAccT study found around 1% of transcriptions contained entirely fabricated phrases, with 38% of those hallucinations carrying explicit harms.

This is also why a bias audit of the scoring model alone is incomplete. If the transcript is systematically worse for some speakers, the scoring model can be perfectly even-handed and the pipeline still is not. That is one reason an audit needs to run stage by stage rather than end to end.

When Averages Hide the Gap


Takeaway: ask your vendor for word error rate broken down by speaker group. A single headline figure can't tell you what you need to know.

3. What is deliberately not analyzed

Visual analysis is gone from the major vendors, and the reason is the most useful fact in this whole subject. HireVue announced in January 2021 that it would not use any visual analysis in its pre-hire algorithms going forward. Fortune's reporting put a number on why: nonverbal data contributed about 0.25% to predictive accuracy in most cases, rising to 4% for customer-facing roles.

Facial analysis was not dropped only because it was contentious. It was dropped because it barely did anything. Tone and pause analysis followed later the same year, and the current position is transcript-only. The company's own page states plainly that its video assessments do not use facial analysis, video, or audio data to evaluate candidates.

Two caveats keep this honest. Smaller vendors still actively market voice analytics measuring pitch, pace and emotional cues, so "the industry has moved on" is true of the large enterprise players and false as a blanket statement. And in the EU, inferring emotions in a workplace context is prohibited outright under the AI Act, applicable since February 2025.

Takeaway: get your vendor's answer on visual and vocal analysis in writing. It is a one-line question and the answers still differ across the market.

4. Scoring

Here the products diverge sharply, and the differences matter more than any feature comparison.

Some score with a model trained to reproduce the ratings of trained human evaluators, reporting convergent validity against those raters in the region of .55 to .74. Some train against actual job outcomes such as sales performance or turnover, with reported predictive validity between .25 and .49. A third approach uses a large language model reasoning against a written rubric. A fourth maps transcript language onto a personality inventory.

These fail in different ways, which is why the distinction is not academic. An outcome-trained model inherits the biases of your past hiring decisions. A rater-trained model inherits the biases of the raters and can never exceed the human judgment it was validated against. An LLM grading against a rubric is only as good as the rubric and is the least studied of the three.

One independent critique is worth knowing. Reviewing a major vendor's explainability statement, the Center for Democracy & Technology noted that its competency models were trained on more than 30,000 aggregated video interviews, suggesting models not tailored to specific employers, occupations or industries.

Takeaway: ask which of these four your tool does. If the vendor cannot say, that is the answer to a different and more important question.

5. Surfacing to the reviewer

The direction of travel at the enterprise end is away from a machine-generated score and toward machine-generated evidence for a human score.

Current products increasingly produce transcripts, recaps, talk-time balance and competency tagging against behaviorally anchored rating scales, with a large language model generating explanations rather than primary scores. One vendor names the specific model it uses for this and states that human teams retain control over every outcome.

This is a meaningful improvement and it introduces a subtler problem. A reviewer reading an AI-generated summary that says a candidate demonstrated limited evidence of stakeholder management is not forming an independent judgment. They are ratifying one, faster. Anchoring is the failure mode of this design, and it does not show up in any audit of the scoring model.

Takeaway: have reviewers score before they read the AI summary, at least on a sample. If the scores move, you have measured your anchoring effect.

6. Accommodation and the human decision

The accessibility failures in this pipeline are concentrated in stage two and they are severe. A peer-reviewed study of four commercial speech-to-text systems found mean word error rates of 52.6% for d/Deaf and hard-of-hearing speakers against 5.0% for controls, rising to 85.9% where speech intelligibility was low. At those rates a scoring model is not evaluating a weak answer. It is evaluating noise.

This is now a live legal question. The ACLU of Colorado filed a complaint with state and federal agencies on behalf of a Deaf, Indigenous employee alleging that a video interview platform lacked working captions and that she was rejected with feedback about her communication style. These are allegations, not findings, and the vendor has called them without merit.

The obligations are clearer than the case law. California's rules, in force since October 2025, require employers to consider reasonable accommodations for candidates assessed by automated systems and specifically scrutinize systems examining tone of voice or physical characteristics. Illinois has required since 2020 that employers using AI to analyze video interviews notify applicants, explain how the AI works and what characteristics it evaluates, obtain consent, and delete recordings within 30 days of request.

Takeaway: publish an alternative route before someone has to ask for one. An accommodation a candidate must request mid-process is one most will not request at all.

Where to Start

Export ten completed interviews from your own tool and read the transcripts against the recordings.

Nothing else you do this quarter will be as informative. You will find out how often the transcript diverges from what the candidate said, whether accented or fast speech degrades it, and whether the words your scoring model read were the words spoken. This is a two-hour exercise requiring no vendor cooperation and no budget.

Then pick the ten from the widest range of speakers you can assemble. The average transcript will be fine. The question this pipeline turns on is not what happens on average. It is what happens at the edges, and whether anyone has looked.

Latest From Our Blog

Discover fresh insights, trends, and tips on tech talent and offshore development. Stay informed with our latest updates

Subscribe to our Newsletter

Get the best insights on remote work, hiring, and engineering management in your inbox.

Global AI/ML Developers Cost Comparison: Finding Value in a Competitive Market

Global AI/ML Developers Cost Comparison: Finding Value in a Competitive Market

Workforce14 Apr 2025
Africa's Growing Digital Infrastructure

Africa's Growing Digital Infrastructure

Technology04 Oct 2023
AI Screening in Recruitment: 5 Real Use Cases, Documented Risks, and What Independent Research Shows

AI Screening in Recruitment: 5 Real Use Cases, Documented Risks, and What Independent Research Shows

Human Resource18 May 2026
Bridge Labs & DevMatch: Streamlining Technical Assessments for African Software Developers

Bridge Labs & DevMatch: Streamlining Technical Assessments for African Software Developers

Human Resource25 Apr 2024
Best Communication Tools for Remote Developers 2025 | Team Efficiency Guide

Best Communication Tools for Remote Developers 2025 | Team Efficiency Guide

Human Resource02 Jun 2025
3 Considerations When Hiring Remote Software Developers

3 Considerations When Hiring Remote Software Developers

Human Resource29 Mar 2024

Screen It, Practice It, Score It.

FAQs

  1. Does AI video interview software analyze facial expressions?
    The major enterprise vendors stopped years ago, having found that visual data added almost nothing to predictive accuracy. Smaller vendors still sell voice and expression analysis, and in the EU inferring emotions at work is now prohibited. Ask your specific vendor rather than assuming either way.
  2. If it only reads the transcript, is it fairer?
    It removes one source of bias and concentrates the remainder in the transcription step. Speech recognition error rates vary substantially by dialect, accent and speech disability, so a transcript-only pipeline is not neutral. It has relocated its risk to a stage most buyers never evaluate.
  3. Can candidates be scored entirely by software with no human review?
    Technically yes, and it is the format candidates abandon most often. It also carries the highest legal exposure and the weakest defense when a rejected applicant asks how a decision was reached. Keep a human in the decision, and be able to name them.
  4. What should we ask a vendor before buying?
    Four things: which stage produces the score, whether the transcript or the audio and video is analyzed, the word error rate broken down by speaker group, and what the accommodation path is for a candidate who cannot complete a timed recorded interview. The 12-point buyer's checklist covers the rest of the procurement question.

Related Articles

  • Do Candidates Trust AI Interviews? 6 Findings From 2026 Data
  • AI Screening Software for Recruiters: A 12-Point Buyer's Checklist
  • How to Run an AI Hiring Bias Audit: A 6-Step Process
  • One-Way Interviews: The Candidate's Guide to Standing Out on Camera
  • How On-Demand Interviews Cut Time-to-Hire Without Sacrificing Quality
AI Video Interview SoftwareAI Video SoftwareExplainedHowThe
Hiring process