Interview Prep

How to Prepare for an AI Interview: What the System Is Actually Looking For

The invitation arrives with a link, a three-day window, and no interviewer's name anywhere on it. That absence is the tell — and it changes how you should prepare.

Interview Prep — What the system is looking for

The invitation arrives with a link, a window of three or four days, and no interviewer's name anywhere on it. That absence is the tell. An AI interview is a recorded, chat or voice interview in which software captures your answers and produces a structured evaluation of them, usually by transcribing what you say and scoring that text against competencies the employer chose before you applied. Knowing how to prepare for an AI interview comes down to understanding that narrower target. You are, in effect, writing an essay out loud, and the first reader is a transcript.

That is a colder proposition than a conversation with a person. It is also a more predictable one, and predictability is something you can work with.

What is an AI interview, and which formats will you meet?

Four formats cover almost everything currently in use.

The one-way recorded interview is the most common. Questions appear on screen one at a time, you get a short preparation window, and you record an answer under a timer. Nobody is watching live. Research by Dunlop and colleagues, published in the International Journal of Selection and Assessment in 2022 and drawn from 2,550,105 responses by 627,999 candidates across 12,105 interview templates at an Australian vendor, found most of these interviews contained four or five questions, with roughly 30 seconds to prepare and two minutes to answer. Employers overwhelmingly kept the platform's default settings rather than choosing their own.

The conversational agent is newer: a voice or chat bot that asks a question, listens, and asks a follow-up based on what you said. It feels more like a phone screen and is more forgiving of a rambling first sentence, because it can probe.

The third is a live human interview with an AI notetaker producing a transcript and summary. The fourth is a hybrid, where a coding or written task sits alongside recorded answers. The honest answer to what to expect in an AI interview is that the format matters less than the scoring, and the scoring is usually the same underneath.

What does an AI interview actually measure?

Most of what candidates fear is not being measured at all. HireVue, one of the larger vendors, states plainly in its published explainability material that its assessment "relies only on what is said by the candidate and does not use any video analysis or other audio characteristics," specifying that it does not assess facial expressions, body language, background or tone of voice. The company discontinued facial analysis in its screening models in 2021, as SHRM reported at the time. Its pipeline converts speech to text using Rev.ai, runs the text through a natural language model, and scores it against more than twenty competencies including problem-solving, adaptability, teamwork and drive for results, returning a top, middle or low tier for a human recruiter to act on.

The system generally scoresThe system generally cannot score
The content of your transcriptWhether the interviewer likes you
Whether you answered the question askedRapport, banter, a shared joke
Structure: situation, action, resultCharisma or presence in a room
Specificity — named tools, numbers, rolesYour appearance
Whether you finished the answerThe story you meant to tell but didn't say
Completion of every questionNervous energy, if the words are still there

The Center for Democracy and Technology has criticised vendors on exactly this point: published explainability statements rarely list the full set of competencies or explain how each is measured. So treat the right-hand column as reassurance rather than a guarantee, and assume the left-hand column is where the marks are.

Why is answering a machine different from answering a person?

A human interviewer fills in your gaps without noticing. They know from context what "that project" means, they read the shrug that stands in for "it was a mess", and they wait through your pause because they can see you thinking. A transcript preserves none of it.

Four habits close most of that gap.

Front-load the answer. Say the conclusion in your first sentence, then support it. A person will wait ninety seconds for your point; a scorer reading a truncated transcript may never reach it.

Name the situation explicitly. Not "at my last company" but "at a logistics startup in Pune, on a team of six, I owned the vendor onboarding flow." Concrete nouns are what a language model has to work with.

Kill pronouns with no antecedent. "We fixed it and it worked" is four words of nothing. Say who, say what, say which system.

Say the technical terms out loud instead of gesturing at them. If you rebuilt a pipeline, name the tools. If you used a framework, say its name once and then explain what you did with it. This is the single change that most often moves a score, and it feels unnatural at first because in conversation it sounds like showing off.

What answer structure fits a 90-second or two-minute limit?

A shape that survives the timer:

  1. One sentence of direct answer. Ten seconds. "The hardest deadline I've handled was a payment integration we had to ship in eleven days instead of the six weeks originally scoped."
  2. The situation, with specifics. Fifteen to twenty seconds. Company type, team size, your actual role, what was at stake.
  3. What you personally did. Sixty seconds, and the bulk of the marks. Three or four concrete actions in first person singular. "We" is fine for context and useless as evidence.
  4. The outcome, with a number. Fifteen seconds. Even an honest approximation beats "it went well."
  5. One line of reflection, if time allows. What you'd do differently.

Rehearse this out loud on a timer. Reading it silently teaches you nothing, because the constraint is speech rate, not thought. Most people speak 130 to 150 words a minute under pressure, which means a two-minute answer is roughly 280 words. That is shorter than it feels.

Prepare five or six stories, not twenty. A single well-specified story about a difficult stakeholder can answer conflict, influence, communication and judgement questions with different emphasis each time.

Practise before it counts

Xakal runs AI interviews that read what you actually say — not how you look. Try a practice round and see your transcript.

Try Xakal free →

How to prepare for an AI interview: the setup that decides your score

Audio matters more than everything else combined, because if the transcription is wrong the score is scoring the wrong words. In rough order of impact:

  1. Audio. Wired earphones with an inline microphone beat a laptop's built-in array. Bluetooth earbuds are the worst option of the three: they compress speech aggressively and often clip the first syllable after a silence, which is exactly where your answer starts. Record thirty seconds and play it back before you begin.
  2. Room noise. A ceiling fan, a road-facing window and a nearby air conditioner all raise error rates. A smaller room with soft furnishings transcribes better than a large empty one.
  3. Lighting and framing. Sit facing a window, not with your back to it, camera at eye level, head and shoulders in frame. Cosmetic for a transcript-only system, but it matters when a human watches the recording later, and one usually does.
  4. Connection and device. Wired or a strong Wi-Fi signal, laptop over phone, everything else closed. The upload happens after you finish — do not close the tab.
  5. Interruptions. Doorbell off, phone face down and silent, someone told not to walk in.

An hour of AI video interview practice with your own setup, recording answers and listening back, will surface more problems than any amount of reading. Most AI interview tips skip this because it's tedious. It's also the part that changes results.

What should you do about pauses, restarts and a disabled retake button?

Half of how to prepare for an AI interview is deciding in advance what you'll do when something goes wrong. Assume you get one take: the Dunlop study found preview and re-recording were only rarely permitted, so a retake button is a gift rather than a right.

If you stumble mid-answer, do not restart from the beginning. Restarting burns thirty seconds of a two-minute window and the transcript keeps the false start anyway. Say a clean bridge — "let me put that more precisely" — and carry on. A transcript with one self-correction reads as a person thinking. A transcript that ends mid-sentence because you ran out of time reads as an incomplete answer.

Short pauses cost you nothing. A four-second silence while you think appears in the transcript as a gap, and gaps are not scored. Filler is worse than silence, because "um, so, yeah, that's a good question" consumes ten of your ninety seconds and adds no content.

Finish every question, even a weak one. An unanswered question is usually a zero, and a mediocre answer is not.

Is it cheating to read from a script?

Yes, and it is more detectable than people assume. Reading produces a flat prosody, an unnatural word rate and eye movement that tracks left-to-right across a screen. Scripted answers arrive in polished, uniform-length paragraphs with none of the self-correction real speech contains, and vendors now flag exactly these patterns.

The version that gets people caught most often is the second device: a phone or tablet with an AI assistant feeding answers. Response latency gives it away, as does the gaze offset, and increasingly the substance does too — the answers are generic where the question was specific.

Bullet points on paper beside the camera are fine and always have been. Six words per story, glanced at, not read. If a platform states that notes are not permitted, believe it.

What if you have a speech difference, a strong accent, or English is not your first language?

The concern is legitimate and it is documented. Koenecke and colleagues, publishing in PNAS in 2020, tested five commercial speech recognition systems on matched samples and found an average word error rate of 0.35 for Black American speakers against 0.19 for white speakers — Apple's system showed 0.45 against 0.23. A 2025 arXiv study of automatic speech recognition on non-native English, using the L2-ARCTIC corpus covering Hindi, Chinese, Arabic, Korean, Spanish and Vietnamese first-language backgrounds, found modern systems performing far better on read speech than older benchmarks, but error rates rising on spontaneous speech, which is what an interview is.

What you can do: slow down slightly, and separate your words more than feels natural. Say acronyms and product names, then expand them once. Avoid trailing off at the end of sentences, where recognition errors cluster.

What you can ask for: an accommodation. The US Department of Justice's guidance on algorithms and disability discrimination in hiring warns explicitly that voice analysis tools can screen out people with speech impairments who are fully qualified, and requires employers to have clear procedures for requesting an alternative — and to ensure that asking does not damage your candidacy. Ask the recruiter for a live interview instead. It is a reasonable request, and in some jurisdictions a protected one.

What happens after you submit, and what rights do you have?

Often, less than you'd like. Greenhouse's May 2026 survey of 2,950 candidates across the US, UK, Germany, Australia and Ireland found 63% had faced an AI interview, up 13 percentage points in six months, and that 51% of those who completed one never received an outcome. Forty-one per cent said AI had made job searching more stressful. Thirty-eight per cent had abandoned a hiring process because of it. Only 19%, though, wanted less AI in hiring — most wanted the same or more, with better safeguards.

Your rights depend on where you are. In the EU, Article 50 of the AI Act has applied since 2 August 2026 and requires that you be told, clearly and at the outset, when you are interacting with an AI system, unless it is obvious; deployers of emotion recognition systems must also inform you. The heavier high-risk obligations covering recruitment tools were deferred by the Digital Omnibus to 2 December 2027, so the transparency duty is what you can rely on today. In Illinois, the Artificial Intelligence Video Interview Act has since 2020 required employers to explain how the AI works and which characteristics it evaluates, obtain your consent, limit who sees the video, and delete it within 30 days of your request — including copies held by vendors. HB 3773 extended notice duties across AI in employment decisions from 1 January 2026. In India, the Digital Personal Data Protection Rules were finalised in November 2025 with core obligations phasing in to May 2027.

Two requests are worth making regardless of jurisdiction: ask what the tool measured, and ask whether a human reviewed the output before the decision. In the Greenhouse data, 46% of candidates wanted the option of a human interview and 38% wanted human review before anything was final. Employers hear that request more often now, and a growing number honour it.

Most of how to prepare for an AI interview reduces to something unglamorous: five stories, a timer, decent earphones, and one recording of yourself you're willing to listen to. Serious AI interview preparation stops there. If you want a low-stakes run at the format, Xakal's platform at thexakal.com offers candidates a free Xara AI Interviews attempt once a day, which is enough to read your own transcript before it counts for something.