AI interviews create a procurement problem before they create a productivity gain
A July 20 discussion among HR professionals asked why video and AI interviews keep being approved when practitioners often oppose them. The thread drew 52 points and 34 comments. The most revealing responses were not about model architecture. They described decision power: executives see a dashboard and a faster time-to-fill promise, while recruiters' warnings about candidate experience, nuance, and implementation arrive later or carry less authority.
August 1 evidence update: the top practitioner comment, with 132 upvotes, said a dashboard and a claim to cut time-to-fill can send candidate-experience objections "straight to the trash." Another reply, with 87 upvotes, reduced the governance problem to eight words: "HR has no say." These are community observations, not measured employer outcomes, but they identify a review failure that a technical bias audit alone will not catch.
That is the operating gap this workflow closes. “AI interview” can mean a scheduling chatbot, a recorded one-way video tool, speech-to-text, an interview-notes assistant, a structured-question engine, an automated scorer, a ranking system, or an avatar that conducts a conversation. These functions do not share one risk level. Approving the product name without mapping the exact function allows a lower-risk convenience feature to smuggle in a higher-risk decision role.
The employer remains responsible for the hiring process it chooses. Vendor assurances do not answer whether a feature is job-related for the intended role, accessible to the applicant population, properly disclosed, validated for the actual decision, or configured in a way that satisfies applicable law and policy. A polished demo can show consistency while hiding exclusions, proxy features, unreliable transcription, silent model changes, or reviewer overreliance.
The right procurement question is not “Does it use AI?” It is: what data enters, what feature is derived, what output is produced, who sees it, how it changes the candidate's opportunity, which person owns the final decision, and what evidence supports every link in that chain?
If the vendor cannot explain the decision role, inputs, features, validation population, accessibility path, and change process, HR does not have a product to approve. It has an unbounded claim.
This guide is operational guidance, not legal advice. Requirements vary by location, employer, role, tool function, and how much the output influences a decision. Use qualified legal counsel and the appropriate privacy, accessibility, labor, security, procurement, and HR owners.
September update: approve the live conversation, not only the vendor
A newly published deployed study gives HR a sharper test than “the bot completed the script.” Across 428 turns, the AI interviewer used a deepening probe in only 4.9% of turns, while 28.7% of question-bearing turns bundled multiple questions despite an explicit one-question-at-a-time instruction. Researchers also observed information loss, premature termination, latency, and interruption. These are not cosmetic defects. Each one changes which evidence a candidate gets to provide.
The study involved 15 participants in a semi-structured research-interview setting, not a hiring validation study. Do not transfer its percentages directly to a vendor or job. Use the failure classes as release tests. A recruiting system has higher stakes, more varied candidates, more devices and languages, and a direct institutional signal: delegating the first human contact tells applicants what the employer values.
Current practitioner evidence shows that reaction can be strongly negative. In an August 31 r/recruiting thread, the linked AI interviewer was described as impersonal, unsettling, and something participants would refuse to use. The top comment said, “Horrific, did one, will never do another.” This is qualitative community evidence, not a representative survey. It still identifies a failure a model-accuracy benchmark cannot see: qualified candidates may abandon before HR collects any interview evidence.
There is counterevidence. Classet, which sells an AI phone interviewer, reports 10,167 opt-in post-interview ratings collected through August 25 and says 87.5% were positive. Treat this as useful first-party product evidence with clear limits. It covers one vendor's invited population and respondents who supplied feedback. It does not establish acceptance for every job, demographic group, disability, language, interface, employer, failed session, abandoned invitation, or AI decision role.
The release question is not whether candidates like AI interviews on average. It is whether this exact journey lets every covered candidate understand the process, provide job-relevant evidence, recover from failure, choose an equivalent path, and reach a person with authority.
Freeze one release manifest
Record the exact model and version, system prompt, interviewer outline, rubric, voice or avatar, turn-taking settings, latency budget, termination rules, retry behavior, transcription model, scoring thresholds, integrations, data locations, languages, jobs, jurisdictions, candidate groups, and human-review configuration. The approval applies to that manifest only. A model swap, new avatar, changed threshold, added scoring feature, or different job returns the release to HOLD until its regression evidence is accepted.
InvitationPlain-language AI disclosure, purpose, data, time, support, accommodation, and equivalent human or accessible route.
ConversationVersioned outline, one question at a time, job-related probes, interruption recovery, latency handling, and visible progress.
EvidenceAnswer, transcript, correction, rubric mapping, source segment, reviewer notes, and every failed or missing item remain traceable.
TakeoverA named human can pause deadlines, continue the interview, correct the record, and prevent automated disposition.
ReleaseQualified owners sign one manifest with monitoring, stop thresholds, rollback, expiry, and re-review triggers.
Test conversational evidence, not synthetic friendliness
Build a production-like case set from the actual job rubric. Include concise answers, long narratives, ambiguous examples, pauses, self-corrections, interruptions, requests to repeat, unsupported assumptions, low-bandwidth audio, accent and speech variation, assistive technology, background noise, lost connections, accommodation requests, and answers that satisfy a competency using unexpected vocabulary.
For each case, define the expected evidence before running the bot. Then compare the transcript and rubric record with the source conversation. A greeting and frequent acknowledgment can make the interface sound attentive while collecting shallow or incomplete evidence. Measure whether the bot asks the missing job-related probe, not how often it says “that makes sense.”
| Release measure | What to inspect | Hold or stop signal |
| Rubric coverage | Required job evidence elicited and traceable to exact answer segments | Missing criterion treated as negative evidence |
| Deepening probes | Follow-up asks for context, action, reasoning, result, or learning when required | Generic acknowledgment replaces a needed probe |
| Question load | One answerable question per turn unless the script justifies a bundle | Candidate answers one part and the system silently drops the rest |
| Information retention | Facts survive paraphrase, summary, handoff, and rubric extraction | Contradiction, omission, or invented detail affects evaluation |
| Termination | Session ends only after required coverage or explicit candidate choice | Silence, latency, or misunderstood answer closes the interview |
| Interruption recovery | Bot stops speaking, preserves the candidate's turn, and resumes from known state | Overtalk, repeated loop, or lost answer |
| Latency | End-of-speech to response, timeout behavior, and candidate-visible status | Dead air causes abandonment or duplicate answers |
| Transcript accuracy | Consequential names, numbers, negatives, qualifications, and corrections | Reviewer cannot inspect or correct the source record |
| Human takeover | Transfer keeps context, pauses deadlines, and prevents automated disposition | Support can apologize but cannot restore opportunity |
Do not invent one universal pass threshold from the research percentages. Set requirements from the job, decision role, legal advice, risk appetite, and baseline human process. Some outcomes are blockers rather than metrics: an unapproved automatic rejection, inaccessible path without an equivalent route, unsupported trait inference, lost consequential answer, disclosure failure, or takeover channel without decision authority should stop release regardless of the average score.
Make candidate control available before the clock starts
Disclosure must arrive early enough for the candidate to make a practical choice. Explain that an AI system will conduct or assist the interview, what it records, whether it transcribes or evaluates answers, who receives the output, how long data is retained, how to request an accommodation, and how to choose an equivalent human or accessible route. Test comprehension with candidates rather than measuring whether the notice link rendered.
Current rules illustrate why specificity matters. Ontario guidance says covered public job postings must state whether AI is used to screen, assess, or select applicants. New York City's Local Law 144 applies when a covered automated employment decision tool is used and requires a recent bias audit, public information, and notice. The exact applicability is a qualified legal question; neither rule should be reduced to a global product checkbox.
The EEOC's worker guidance gives a concrete disability failure: an interview chatbot asks whether an applicant can stand for three hours and ends the interview after a wheelchair user says no, even though the person could perform the job while seated. The operational lesson is broader than one jurisdiction. An alternate route must evaluate ability to perform the job with accommodation, preserve equivalent timing and consideration, and reach someone empowered to correct an invalid screen.
Give candidates a way to correct transcript errors before the record affects review. Preserve the original, the candidate correction, the reason, and the human resolution. A silent overwrite destroys evidence; refusing correction turns a transcription system into an unappealable decision source.
Bind the go-live decision to observable stop rules
Launch in evidence-support mode. The interviewer may collect answers and propose rubric notes, but it should not reject, rank-cut, or make the final hiring decision. Reviewers must see the relevant answer segment, inspect missing evidence, disagree with the suggested mapping, and document a reason. Sample cases without showing the AI recommendation when feasible to detect automation bias.
Monitor invitations, starts, completions, abandonment by step, technical failure, latency, interruptions, alternate-path requests, accommodation response, transcript correction, human takeover, complaints, reconsideration, reviewer overrides, missing rubric evidence, and version changes. Break these down by job, location, language, device, and relevant subgroup where lawful and statistically responsible. Do not treat zero complaints as zero harm; candidates may not know what happened or may avoid challenging a potential employer.
Stop immediatelyAutomatic rejection, undisclosed evaluation, unsupported trait inference, inaccessible participation without an equivalent route, data exposure, or unapproved version change.
Pause and investigateLost answers, premature termination, repeated interruption loops, material transcript error, abnormal abandonment, delayed takeover, subgroup difference, or complaint spike.
Repair opportunityReopen the candidate path, preserve evidence, offer a qualified human interview, correct the record, and prevent the failed session from disadvantaging the candidate.
Require renewalApproval expires on a fixed date and on any material change to model, prompt, voice/avatar, rubric, threshold, integration, job, jurisdiction, or decision role.
Update: define decision rights before asking whether the vendor works
A vendor review can collect excellent evidence and still fail if nobody knows who may stop the purchase. The accountable executive may own budget and business outcomes, but that should not turn every risk domain into an informal recommendation. Accessibility, employment-law, privacy, security, validation, and candidate-experience reviewers need defined authority, response dates, and an escalation path before the sales process creates momentum.
Use a decision-rights map that distinguishes sponsorship from evidence acceptance. The sponsor states the business problem and accepts operational tradeoffs within policy. Domain owners decide whether evidence in their field is sufficient. The HR process owner defines the actual decision role and candidate path. Procurement enforces conditions in the contract. A named stop owner can pause a pilot when pre-registered rules fire. The final signer approves only the exact use, version, jobs, jurisdictions, settings, and review period documented in the packet.
| Role |
Decision authority |
Evidence required |
| Executive sponsor |
Defines business problem, budget, and acceptable operational tradeoff |
Baseline cost, time, quality, candidate impact, and alternative options |
| HR process owner |
Defines workflow stage, decision role, reviewer workload, and candidate path |
Process map, job criteria, baseline outcomes, support and escalation data |
| Domain reviewers |
Accept or block evidence within legal, validation, accessibility, privacy, and security domains |
Dated artifacts, tests, limitations, remediation owner, and closure evidence |
| Procurement owner |
Binds approved limits, evidence duties, change notice, audit, exit, and deletion terms |
Contract schedule mapped to the approval packet |
| Pilot stop owner |
Pauses use when a pre-defined threshold or incident occurs |
Live metrics, complaint route, incident alert, configuration and version log |
| Final signer |
Approves one bounded use or records HOLD / REJECT |
Closed blockers, conditions, named owners, expiry, and renewal triggers |
This structure does not require consensus on every preference. It requires clarity about which issues are preferences and which are blocking evidence failures. An executive can decide that a slower implementation is acceptable. An executive should not convert missing disability access, undisclosed scoring, an unvalidated feature, or an uncontained data flow into “accepted risk” without the qualified owners and governance process required by the organization.
A July 30 field-experiment preprint on automated voice interviews adds a useful counterweight to practitioner skepticism. The authors report that AI voice interviews produced more structured and consistent questioning while remaining responsive, and that transcripts contained more hiring-relevant information. That is a benefit claim worth testing. It does not answer whether a particular vendor, job, population, language, deployment, scoring feature, notice, or accommodation path is valid. The decision-rights map ensures promising evidence enters a controlled evaluation instead of becoming an all-purpose approval.
Record disagreements in the packet. If candidate experience is poor but interview consistency improves, do not average the two into one composite score. Preserve both, identify the owner of each outcome, and state the tradeoff the final signer is accepting. If the pilot cannot provide an equivalent human path or explain how evidence affects decisions, stop even if the time-to-fill dashboard improves.
Step one: map the system before reviewing the vendor
Start with a one-page system boundary. Do not let the phrase “human in the loop” substitute for a diagram. Trace the candidate from invitation through final disposition and name every system, output, and decision owner.
Candidate invitation
|
v
Notice, consent, accommodation, alternate path
|
v
Capture layer
- text, audio, video, device and session data
|
v
Processing layer
- transcription, summarization, feature extraction
|
v
Evaluation layer
- question rubric, score, rank, recommendation, flag
|
v
Recruiter or hiring-manager review
- source evidence visible?
- disagreement allowed and recorded?
|
v
Employment decision and candidate communication
|
v
Retention, deletion, appeal, audit, and model-change loop
For each box, record whether it is performed by the employer, vendor, subprocessor, model provider, applicant tracking system, or another integration. Document where data is stored, the identity used to access it, whether it crosses borders or tenant boundaries, and whether it is reused for product improvement or model training.
Then assign a decision role to every output. Use a simple ladder:
| Role |
Example |
Default review posture |
| Administrative |
Scheduling, reminders, connection checks |
Privacy, accessibility, reliability, and disclosure review |
| Evidence support |
Transcript, structured notes, question coverage |
Accuracy sampling, source access, correction, retention controls |
| Recommendation |
Competency score, risk flag, suggested disposition |
Job-validation, subgroup, automation-bias, and legal review |
| Decision |
Automatic rejection, progression, rank cutoff |
Highest scrutiny; prohibit by default without explicit approval and strong evidence |
A transcript can still affect a decision if it becomes the only evidence a reviewer sees. A recommendation can become a de facto decision when recruiters process hundreds of applicants and rarely override the score. Review the real operating behavior, not the contract label.
Replace the demo with a vendor evidence register
Ask the vendor to support each material claim with a dated artifact. “Bias tested” is not evidence until the vendor provides the construct, population, sample, method, outcomes, subgroup results, limitations, and relevance to your intended jobs. “Explainable” is not evidence until a reviewer can connect the output to job-related source material.
| Evidence area |
What to request |
Reject or hold signal |
| System definition |
Architecture, data flow, all models, features, outputs, thresholds, and subprocessors |
“Proprietary AI” used to avoid describing the decision path |
| Job relevance |
Feature-to-job requirement mapping and criterion evidence |
Emotion, personality, honesty, gaze, appearance, accent, or vague culture-fit inference |
| Validation |
Comparable roles and populations, outcome definition, sample, comparator, errors, and limitations |
Vendor-wide accuracy number with no job or population match |
| Subgroup performance |
Selection, error, completion, accommodation, and appeal results by relevant group where lawful |
Fairness claim based only on removing protected fields |
| Accessibility |
Testing method, standards, disability participation, alternate path, and support response |
Accommodation requires disclosure to a generic chatbot or arrives after the deadline |
| Human review |
Source evidence, override workflow, training, reviewer metrics, and audit record |
Reviewer sees a score without the underlying answer or rubric |
| Change control |
Version history, customer notice, regression results, opt-out, rollback, and revalidation triggers |
Vendor may change models or thresholds without notice |
| Data controls |
Retention, deletion, backups, sharing, model training, residency, access logs, and incident response |
Indefinite retention or reuse outside the hiring purpose |
Validation must match the claim. If the tool predicts job performance, ask how job performance was defined and measured. If it claims to improve interview consistency, compare question delivery, follow-up logic, completion, and reviewer reliability. If it only transcribes, measure word error, speaker attribution, omissions, and correction rates across languages, accents, speech patterns, devices, and noisy environments relevant to the applicant pool.
Prefer observable constructs. A candidate's answer can be scored against a documented behavioral rubric tied to job requirements. Claims about enthusiasm, trustworthiness, emotional state, personality, or “executive presence” invite ambiguity and proxy risk, especially when inferred from face, voice, timing, or background. If the employer would not train human interviewers to make the inference reliably and job-relevantly, it should not buy a machine that claims to do so.
License and ownership matter when the vendor uses open-source components. Request a software bill of materials or component register, licenses, model terms, and a description of modifications. Open source can improve inspectability; it does not validate the employer's use or transfer responsibility away from the deployment team.
Build a jurisdiction matrix, not one global notice
Employment technology sits under overlapping discrimination, accessibility, privacy, automated-decision, video, biometric, labor, and consumer-protection rules. The matrix below highlights official public sources that should trigger qualified review; it is not a complete legal inventory.
| Reference point |
Operational signal |
Evidence to keep |
| U.S. federal employment law and ADA |
Selection procedures must be job-related and employers must provide reasonable accommodation where required |
Job analysis, validation, accommodation process, alternative assessment, reviewer training |
| Illinois AI Video Interview Act |
For covered interviews, notify, explain how AI works and general characteristics, obtain consent, limit sharing, and honor deletion requests within 30 days |
Notice version, consent, explanation, recipients, deletion request and completion log |
| New York City Local Law 144 |
Covered automated employment decision tools require a recent bias audit, public audit information, and notices |
Scope analysis, audit date and summary, notice timing, tool version, usage record |
| Colorado ADMT law and rulemaking |
New requirements for automated decision-making technology influencing consequential decisions take effect January 1, 2027 |
Readiness assessment, developer and deployer duties, data correction path, rulemaking watch |
| NIST AI RMF and DOL inclusive hiring framework |
Use continuous governance, mapping, measurement, and management with accessibility included in procurement and deployment |
Risk register, owners, measures, decisions, incidents, changes, and periodic review |
Illinois provides a concrete example of why generic consent is inadequate. The statute's text calls for notice that AI may analyze the video, information explaining how the AI works and the general types of characteristics evaluated, and consent before evaluation. It also limits sharing and requires employers to delete interviews and instruct recipients to delete copies, including backups, within 30 days of an applicant request.
New York City's official DCWP page states that covered AEDTs cannot be used without a bias audit from within one year, public information about the audit, and required notices. The threshold question is whether the tool and use fall within the law. Do not rely on the vendor's generic classification; document the employer's own scoped analysis using the actual configuration and decision role.
Colorado's Attorney General reports that the 2026 reenacted ADMT provisions take effect January 1, 2027 and include developer and deployer requirements for technology that materially influences consequential decisions. A vendor contract signed in 2026 may run into that effective period. Procurement should therefore include update rights, required evidence, rule-change cooperation, and termination rights now.
A notice does not cure a bad system. Consent does not establish job relevance. A bias audit does not prove accessibility. Human review does not excuse weak data controls. Treat legal duties as minimum gates inside a broader employment-quality and risk process.
Accessibility is a participation requirement, not a support ticket
The EEOC's AI and ADA resources direct employers to disability risks in software, algorithms, and AI used to assess applicants. The U.S. Department of Labor's AI and Inclusive Hiring Framework extends the practical point into procurement: employers should evaluate accessibility before acquisition and throughout use, not wait for an excluded applicant to report a failure.
Test the complete candidate path with disabled participants and assistive technology. Cover keyboard-only navigation, screen readers, captioning, audio-only participation, speech and hearing differences, camera limitations, color and contrast, timing, cognitive load, neurodiversity, low bandwidth, older devices, mobile use, and the process for requesting an accommodation.
The alternative path must be comparable. An applicant should not lose timing, visibility, recruiter attention, or scoring consistency because they use an accommodation. Avoid designs where the only way to request help is inside the tool that is failing, after a timer has started, or through a channel that requires disability disclosure to an unnecessary vendor.
Measure accessibility outcomes during the pilot: completion rates, time, technical failure, help requests, accommodation response time, alternate-path use, withdrawal, complaint, transcript correction, and reviewer treatment. A vendor accessibility conformance statement can inform testing; it is not a substitute for observing the intended workflow.
Accessible invitationState the format, technology, time, support contact, accommodation option, and non-AI alternative before the candidate starts.
No penaltyEnsure an accommodation or alternate path does not reduce consideration, create delay, or expose unnecessary information.
Evidence correctionLet candidates report transcript or identity errors before those errors influence evaluation.
Human supportProvide a reachable person with authority to pause deadlines, correct records, and move the candidate to an equivalent process.
Run a bounded pilot that cannot silently become production
Write the pilot charter before configuring the tool. Name the jobs, locations, candidate volume, languages, workflow stage, approved functions, prohibited functions, data fields, reviewers, duration, baseline, measures, complaint route, and stop authority. Give the pilot an expiry date. Continued use after expiry requires a signed decision based on results.
Begin in shadow or evidence-support mode. The tool may generate a transcript or proposed rubric notes, but human reviewers make decisions using the existing process. Compare the tool with the baseline without showing its score to decision-makers when feasible; this reduces automation bias while the team measures reliability.
| Measure family |
Useful pilot measures |
Bad shortcut |
| Validity |
Agreement with job-related rubric, error analysis, outcome relationship, reviewer reliability |
Vendor's global “accuracy” |
| Fairness |
Selection, error, completion, override, and appeal patterns by relevant subgroup where lawful |
Protected fields were removed |
| Accessibility |
Completion, failure, accommodation, alternate path, support response, correction |
Tool has an accessibility page |
| Candidate experience |
Disclosure comprehension, withdrawal, complaint, support, trust, and preference |
Interview link was opened |
| Human oversight |
Override rate, evidence opened, review time, reason quality, reversal |
A recruiter clicked approve |
| Operations |
Total time, rework, technical failure, support burden, deletion completion |
Minutes of interviewer calendar saved |
Predefine stop rules. Stop immediately for an unapproved automatic rejection, inaccessible participation without an equivalent path, undisclosed model or feature changes, exposure of candidate data, use of prohibited traits or proxies, missing deletion capability, or a serious security incident. Pause and review when subgroup differences, complaint rates, transcript errors, overrides, or reviewer reliance cross agreed thresholds.
Do not interpret lack of complaints as safety. Applicants may not know that AI influenced a decision, may not understand the appeal route, or may avoid challenging a potential employer. Include proactive sampling, candidate feedback, and independent quality review.
Common failure modes and the control that catches each one
| Failure mode |
What it looks like |
Control |
| Function creep |
Transcription feature later exposes a candidate score by default |
Approved-function register, configuration lock, change notice, regression review |
| Proxy scoring |
Voice, timing, vocabulary, or video quality stands in for capability |
Feature inventory, job mapping, subgroup error analysis, prohibited-feature policy |
| Automation bias |
Reviewers accept the recommendation because the queue is large |
Source evidence, blinded sampling, override analysis, workload caps, reviewer training |
| Accessibility dead end |
Candidate cannot complete video but support answers after the deadline |
Pre-start alternate path, human contact, deadline pause, equivalent evaluation |
| Silent model drift |
Vendor changes transcription, scoring, or thresholds without customer review |
Version notice, revalidation trigger, opt-out, rollback, contract remedy |
| Notice without understanding |
Long privacy policy says “AI may be used” but not how it affects the candidate |
Layered plain-language notice, comprehension test, purpose and output explanation |
| Retention sprawl |
Video, transcript, features, embeddings, scores, and backups follow different rules |
Data-element schedule, deletion verification, recipient instructions, audit logs |
| Appeal theater |
Candidate can email support but nobody can change the outcome |
Named owner, correction authority, response target, independent human reconsideration |
Issue one of four decisions
A useful review does not end with “risks noted.” It ends with a bounded decision, accountable owner, expiry, and conditions.
- Reject: the intended function is inappropriate, evidence is fundamentally weak, prohibited inferences are central, or the vendor cannot support accessible and controlled operation.
- Hold for evidence: the use might be acceptable, but validation, data, accessibility, security, legal, or change-control evidence is incomplete.
- Approve limited pilot: the pilot has defined roles, jurisdictions, data, reviewers, baseline, metrics, no automatic rejection, stop rules, and an expiry.
- Approve defined use: name the exact function, jobs, locations, version, settings, decision role, retention, owners, monitoring, revalidation triggers, and renewal date.
System boundary completeEvery model, input, feature, output, integration, subprocessor, and decision owner is recorded.
Job evidence acceptedQualified owners confirm the evaluated constructs and criteria are documented and job-related.
Accessible alternative readyThe alternative is available before the deadline and does not reduce candidate opportunity.
Notices testedCandidate materials explain the AI role, data, evaluation, support, alternative, correction, and deletion path.
Human review is realReviewers see evidence, can disagree, document reasons, and own the final employment decision.
Change and exit rights signedThe employer receives version notices, can pause updates, export records, verify deletion, and terminate unsafe use.
Connect this gate to the site's AI interview scheduling exception workflow, AI recruiting pilot evaluation skill, resume-screening risk guide, EU AI Act transparency workflow, and HR AI output release gate. The vendor gate approves a system; each live workflow still needs its own operating control.
Frequently asked questions
How should HR test a bot-led interview before release?
Use production-like cases tied to the real job rubric. Measure required evidence coverage, deepening probes, stacked questions, information loss, premature termination, latency, interruption recovery, transcript accuracy, disclosure comprehension, accessible fallback, human takeover, and reviewer evidence use. Bind approval to the exact model, prompt, voice/avatar, rubric, thresholds, integrations, jobs, and jurisdictions tested.
Should HR use AI to score video interviews?
Only with evidence proportionate to the decision impact: documented job relevance, validation for comparable jobs and populations, subgroup and accessibility testing, meaningful notice and alternatives, controlled data use, qualified review, and outcome monitoring. Many employers should start with lower-risk scheduling, transcription, or structured note support.
Is consent enough?
No. Consent may be one requirement for a covered use, but it does not establish fairness, validity, accessibility, security, or job relevance. In an employment context, candidates may also have limited practical freedom to refuse without a credible equivalent path.
Is a human in the loop enough?
No. Reviewers need source evidence, authority to disagree, training, manageable workload, and a documented standard. Measure overrides, reversals, evidence access, and decision reasons to detect rubber-stamp review.
What if the vendor will not disclose proprietary features?
Trade-secret limits do not force the employer to buy an unevaluable system. Define the minimum evidence needed to establish the intended use, risks, controls, validation, and contractual accountability. Hold or reject when that evidence is unavailable.
How often should approval be renewed?
Set a fixed review date and trigger immediate re-review when models, features, thresholds, training data, subprocessors, integrations, jurisdictions, candidate populations, jobs, or decision roles materially change, or when monitoring finds an incident or performance drift.
Sources and reference points
Current regulatory, product, research, and community pages were checked on September 2, 2026. Vendor feedback is labeled first-party evidence and community reports are qualitative signals. This list supports the workflow but is not a substitute for jurisdiction-specific advice.