AI Interviews

AI Candidate Evaluation: How AI Actually Scores an Interview

A score of 72 looks like a measurement. It is a model's guess, compressed. Here is what goes into it, what it misses, and how much weight it can carry.

AI Interviews — How AI actually scores an interview

HireVue stopped scoring candidates' faces in 2020. Its then-CEO Kevin Parker told Fortune in January 2021 that non-verbal data contributed roughly 0.25% to the model's predictive power in most cases, and about 4% even for customer-facing roles — not worth the bias risk. That single decision tells you most of what you need to know about AI candidate evaluation. An AI candidate evaluation is a score generated from what a candidate said, mostly derived from a machine transcript of their answers and any structured response data the interview captured. Everything else is decoration, and some of it is dangerous decoration.

The rest of this piece is about where the score is real, where it's noise, and what has to sit between the score and a rejection email.

What does an AI candidate evaluation actually measure?

Strip away the marketing and the pipeline is short. Audio goes into a speech-to-text model. Text comes out. A scoring model — usually a large language model working against a rubric, sometimes a trained classifier — reads that text against your configured criteria and emits per-question ratings plus a rollup. Structured items such as a code submission or a multiple-choice screen are scored separately and are far more reliable than anything derived from speech.

The academic evidence points the same way. Hickman and colleagues, publishing in the Journal of Applied Psychology in 2022, found that automated video interview models trained on interviewer ratings leaned overwhelmingly on verbal behaviour — what was said, and how much of it — rather than on pitch, loudness, facial expression or head pose. Models trained against observer ratings showed reasonable convergent validity; models trained against candidate self-reports were inconsistent.

So a well-built AI candidate evaluation can genuinely tell you: did the candidate answer the question asked; did they name specific systems, numbers, timeframes and outcomes; is their account of a project internally consistent; does their claimed depth survive a follow-up. Those are content judgments made against text, and they are the reliable core of how AI evaluates candidates today.

What does it infer weakly, and what can't it see at all?

Two categories get conflated constantly, and separating them is the whole job.

SignalHow the score derives itHow much weight it deserves
Relevance and specificity of an answerDirectly from transcript textHigh — this is the model's real competence
Structured responses, code, scored exercisesDeterministic scoringHigh
Domain depth on a named tool or processTranscript plus rubricMedium-high, if the rubric is written by someone who does the job
Communication style, clarity, fluencyProxied through transcript length, filler words, sentence structureLow — confounded by accent, nerves, first language
Confidence, enthusiasm, "culture fit"Inferred from tone or pacingEffectively none
Judgment under ambiguity, integrity, coachabilityNot observableZero

Confidence is the signal buyers most want and the one least available. A pause before answering reads as hesitation to a scoring model and as careful thinking to an experienced interviewer. Same acoustic event.

Then there is the category the machine cannot reach at all. It has no idea whether this person would be good to work with, whether they'd raise a hand when a project is going wrong, or whether the two-year gap on their CV was a caregiving break, a failed startup or a visa problem. It doesn't know that the candidate's last manager was the reason three people left. Context, in other words — not a limitation to be engineered away in the next model generation, but information that was never in the input.

How should an AI interview score map onto a human scorecard?

One rule holds up: the machine fills in cells, it does not sign the form.

A working interview scorecard for a mid-level engineering role might have six criteria. An AI interview can populate three of them with evidence — technical depth, problem decomposition, relevant experience — and each cell should carry the quote it came from, not just a number. The remaining criteria stay empty until a human interviews the person. The recruiter reading the scorecard sees a partially completed document, which is honest, rather than a single number out of 100, which is not.

Two practitioner details that only show up after a few hundred interviews. First, average the question scores at your peril. A candidate who scores 9, 9, 9 and 2 is usually more interesting than one who scores 7 across the board, because the 2 is often a question that was badly worded or outside their stated scope. Averaging hides it. Look at the distribution and the outlier. Second, senior candidates give shorter answers. They've explained their architecture decision a hundred times and they've compressed it. Any automated candidate scoring configuration that rewards answer length will systematically down-rank your most experienced applicants, and it will do it silently.

Where does AI interview scoring break?

These are the failure modes worth testing for before a system touches live candidates. Not hypotheticals — each has published evidence behind it.

Accent and dialect in speech-to-text. Koenecke and colleagues, in PNAS in 2020, tested five commercial speech recognition systems from Amazon, Apple, Google, IBM and Microsoft against 2,141 matched audio snippets from 73 Black and 42 white speakers. Average word error rate was 0.35 for Black speakers against 0.19 for white speakers. Error rates were nearly twice as large for Black speakers in every single system tested, and the researchers traced the gap primarily to the acoustic model rather than to vocabulary or grammar.

Non-native and regionally accented English. A 2024 evaluation of OpenAI's Whisper published in JASA Express Letters found error rates rising with lower English proficiency, and varying sharply by first language — lowest for Swedish and German speakers, highest for Vietnamese and Thai. Even among native varieties, British English produced higher error rates than American. Conversational speech scored worse than read speech, which is exactly the register an interview uses.

Pauses and disfluency. The "Careless Whisper" study presented at ACM FAccT in June 2024 by Koenecke, Choi, Mei, Schellmann and Sloane analysed over 13,000 clips from AphasiaBank and found around 1% of transcriptions contained entirely fabricated phrases. Hallucination was more likely for speakers with longer pauses between words. A candidate who thinks before speaking can have sentences invented and attributed to them.

Domain jargon. Indian hiring runs on vocabulary that general-purpose speech models handle badly: TDS, EPFO, SEBI, Ind AS, plus product and company names absent from training data. Sarvam AI reported in February 2026 that its Saaras V3 model reaches 19.31% word error rate across the ten most widely spoken Indian languages on the IndicVoices benchmark — an improvement, and still roughly one word in five. Ask any vendor which speech model they use and its measured error rate on your accent mix. If they can't answer, that is the answer.

Answer length at both extremes. Very short answers starve the model of text. Very long answers get truncated or diluted. Both drag scores toward the middle regardless of quality.

Audio quality. A candidate on a shared campus laptop in a Tier-2 city, or on mobile data in a co-working phone booth, is being scored partly on their bandwidth. This is the failure mode most likely to correlate with socioeconomic background, and the one least likely to appear in a vendor's bias audit.

See Xara AI interview a candidate live

Structured questions, adaptive follow-ups, a transcript and a scorecard your team can argue with. Book a 30-minute demo — no slides.

Book a demo →

What's the difference between a score that ranks and a score that filters?

This distinction decides whether an AI evaluation is a productivity tool or a legal exposure.

A ranking score changes the order in which a human reviews candidates. Everyone above a generous cutoff still gets read. If the model is wrong about someone, the cost is that a good candidate got read fourth instead of first.

A filtering score removes candidates without a human ever seeing them. If the model is wrong, that person is gone and nobody knows. The error is invisible by construction, so it never enters your feedback loop and never gets corrected.

Ranking degrades gracefully. Filtering fails silently. Given that speech recognition error rates differ by accent by a factor approaching two, a filter applied to a transcript is a filter applied unevenly across demographic groups, whatever the rubric says.

Why are auto-rejection thresholds the riskiest setting in the product?

Because they convert a statistical estimate into a final decision with no witness. It's the single worst configuration available in this category of tool, and it is usually one toggle away.

The enforcement record is unambiguous. In September 2023 the US Equal Employment Opportunity Commission settled with iTutorGroup for $365,000 after the company's application software was programmed to automatically reject female applicants aged 55 or older and male applicants aged 60 or older; more than 200 qualified applicants were rejected. That was a hard-coded rule rather than a learned model, which makes the point sharper: nobody in the process saw the rejections happen.

The more instructive case is still running. In Mobley v. Workday, Judge Rita Lin granted preliminary certification in May 2025 for a nationwide collective action under the Age Discrimination in Employment Act, covering applicants aged 40 and over who were denied recommendations through the platform since September 2020. The named plaintiff had applied more than 100 times. In a May 2026 discovery ruling, Magistrate Judge Laurel Beeler held that Workday's own bias-testing data was protected by attorney-client privilege because counsel had curated it — a reminder that running bias tests through your legal team may protect the results from disclosure but does nothing to protect the candidates.

What do regulators actually require of the human in the loop?

The picture shifted in 2026 and a lot of published guidance is now out of date.

Under the EU AI Act, recruitment and candidate-selection systems sit in Annex III as high-risk, and those obligations were originally due to apply from 2 August 2026. The Digital Omnibus on AI changed that: Parliament approved it on 16 June 2026 and the Council on 29 June 2026, deferring standalone Annex III obligations to 2 December 2027 and product-embedded systems to 2 August 2028. What did not move is Article 50 transparency — from 2 August 2026, providers must still tell people when they are interacting with an AI system.

The deferral is a timing change, not an amnesty. Article 14 of the Act, when it applies, requires that a named human overseeing a high-risk system can understand its capacities and limits, stay alert to automation bias — the tendency to over-trust the output — interpret what the system produces, and decide to disregard, override or reverse it. Build to that standard now; retrofitting oversight into a hiring process eighteen months from launch is far more expensive than designing it in.

New York City's Local Law 144, enforced since 5 July 2023, is more prescriptive. Employers using an automated employment decision tool must commission an independent bias audit within the past year covering selection rates and impact ratios by sex, race/ethnicity and intersectional categories, publish a summary, and notify candidates at least 10 business days in advance. Compliance in practice is another matter: a 2024 Cornell and Data & Society study checked 391 employer websites and found bias audit reports on 18 of them, and transparency notices on 13.

India has no equivalent binding rule yet. MeitY published the India AI Governance Guidelines in November 2025 as a principles-based framework rather than an enforcement regime, which means the discipline is currently voluntary and the reputational exposure is not.

A governance checklist for keeping a human accountable

  1. Name one person, by name, accountable for every rejection the system contributes to. Not a team.
  2. Configure the tool to rank, never to auto-reject. If the vendor won't let you disable threshold rejection, that tells you something.
  3. Require the transcript and the supporting quote alongside every score. A number without evidence is unreviewable.
  4. Test speech-to-text separately from scoring. Record twenty employees across the accents you actually hire, and read the transcripts yourself.
  5. Audit outcomes quarterly by gender, age band, region and education tier — selection rates and impact ratios, not just aggregate scores.
  6. Publish the criteria to candidates before the interview, and offer a human-reviewed alternative on request.
  7. Read the ten lowest-scored transcripts every week. If you can't tell why they scored low, your rubric isn't measuring what you think it is.
  8. Recalibrate annually against people you actually hired and who are twelve months in. Predictive validity that isn't checked is a claim, not a finding.
  9. Keep decision records for at least the limitation period in your jurisdiction. Mobley is being litigated over applications from 2020.
  10. Reserve the right to switch it off for a role, a market, or a week.

Vendors differ on how much of this they hand you. Some platforms — Xakal's Xara AI Interviews among them — expose per-question scores with the transcript attached rather than a single opaque number, which is the minimum you need to review a decision rather than merely ratify it. Worth checking on any tool you're evaluating, at thexakal.com or anywhere else.

AI interview accuracy is a real and measurable property. It is also narrower than the word "evaluation" suggests, and the gap between the two is where candidates get lost.