Bias Audits for AI Hiring: A Practical Methodology
"Bias audit" is one of those phrases that sounds rigorous until you try to define it. Different regulators mean different things. Different vendors produce wildly different artefacts under the same label. And too many HR teams accept a one-page summary as evidence of fairness when it is really evidence of marketing.
This article sets out what a defensible bias audit of an AI-assisted interview system looks like in 2026 — drawing on the methodology established by New York City Local Law 144 of 2021, the NIST AI Risk Management Framework (AI RMF 1.0, January 2023), and the long-standing EEOC Uniform Guidelines on Employee Selection Procedures.
A bias audit is not a certification. It is a measurement exercise that produces numbers you can defend, with limitations you are willing to disclose.
What "bias" actually means in a hiring context
In selection science, bias has a precise meaning: a systematic difference in outcomes for one group versus another that is not explained by job-relevant differences. The classic operationalisation is the four-fifths rule from the EEOC Uniform Guidelines — if the selection rate for a protected group is less than 80% of the rate for the highest-selected group, there is prima facie evidence of adverse impact.
That single metric is necessary but not sufficient. A complete bias audit looks at:
- Adverse impact ratios across protected categories.
- Score distributions — whether the underlying scores cluster differently for different groups.
- Predictive validity by group — whether the score predicts on-the-job performance equally well across groups.
- Construct validity — whether the system is measuring what it claims to measure.
The first two are statistical. The second two require post-hire outcome data, which most vendors do not have at the time of an initial audit.
What NYC Local Law 144 actually requires
NYC Local Law 144 of 2021, which took enforcement effect on 5 July 2023, requires employers and employment agencies using "automated employment decision tools" (AEDTs) for candidates resident in or applying for jobs located in New York City to:
- Have the tool audited for bias by an independent auditor within the year prior to use.
- Publish a summary of the bias audit results.
- Notify candidates that an AEDT is being used and disclose the categories of data collected.
The Department of Consumer and Worker Protection's implementing rules specify the methodology. The audit must compute selection rates and impact ratios across sex categories, race/ethnicity categories, and intersectional categories of sex and race/ethnicity. It must use either historical applicant data or test data, with a clear disclosure of which.
For HR teams outside New York City, LL144 still matters: it is the most concrete, prescriptive regulatory standard in the world for AI hiring audits, and many vendors are aligning their internal methodologies to it as a de facto baseline.
What the NIST AI Risk Management Framework adds
The NIST AI RMF, finalised in January 2023, is broader and more principle-based than LL144. It organises AI risk management around four functions — Govern, Map, Measure, Manage — and emphasises that fairness is one of seven characteristics of "trustworthy AI" alongside validity, reliability, safety, security, accountability, and transparency.
The practical takeaway for HR is that the NIST framework treats a one-time bias audit as a single artefact within an ongoing risk management process. The audit answers "what is the current state?"; the framework asks "what is the process by which you maintain that state as the system evolves?"
A vendor that has done one audit and stopped is not aligned with NIST. A vendor that audits on a defined cadence — typically annually plus on any material model change — is.
A nine-step methodology that holds up
The following is the methodology we use for our own internal audits and recommend HR teams ask vendors to follow.
Step 1 — Define the decision the tool actually makes
A common audit failure is auditing the wrong thing. If the AI produces a score that a human recruiter then uses as one input into a decision, the unit of analysis is the score, not the final hire. If the AI produces a binary pass/fail recommendation, the unit of analysis is the pass/fail. Both need to be audited, but the methods differ.
Step 2 — Assemble the dataset
The audit needs a sample of candidates who have completed the tool, with self-reported demographic data and the resulting scores. The sample should be large enough to produce statistically meaningful subgroup estimates — typically several hundred candidates per protected group as a minimum, more for intersectional analysis.
If real applicant data is insufficient, LL144 permits the use of synthetic test data, with that fact disclosed.
Step 3 — Compute selection rates by group
For each protected group, compute the proportion who would be "selected" under the tool's recommendation. This is the four-fifths rule applied directly.
Step 4 — Compute impact ratios
For each group, divide its selection rate by the selection rate of the highest-selected group. Ratios below 0.80 are flagged for further investigation, per the EEOC convention.
Step 5 — Look at the underlying score distributions
Selection rates can mask important differences. Two groups can have the same pass rate but very different score distributions — for example, one group bunched at the threshold and another spread wider. Plot the distributions and report differences in means, medians, and tails.
Step 6 — Intersectional analysis
The most important contribution of LL144's methodology is its insistence on intersectional analysis — sex by race/ethnicity. Aggregate-level fairness can hide subgroup-level disparate impact. The audit should compute selection rates and impact ratios for each intersectional category.
Step 7 — Investigate any flagged disparities
A ratio below 0.80 is not a verdict. It is a flag. The next question is: is the disparity driven by a real difference in job-relevant skill in the sample, or by something the system is picking up that it should not be?
Investigation techniques include:
- Item-level analysis. Which interview questions or which scorecard competencies are driving the disparity?
- Feature analysis. What signals from the candidate's response is the model weighting most heavily?
- Counterfactual testing. Holding everything else constant, does changing the candidate's audio characteristics (accent, pitch) change the score?
Step 8 — Document mitigations
If a disparity is real and driven by something the system should not be relying on, document the mitigation. Common mitigations include reweighting features, removing items, retraining on a more representative dataset, or adjusting the decision threshold.
Do not "fix" bias by quietly applying group-specific score adjustments. That creates a different category of legal exposure under disparate-treatment doctrine in many jurisdictions.
Step 9 — Publish a summary, retain the full report
LL144 requires a public summary. The full audit, with methodology, dataset description, results, and mitigations, should be retained internally and provided to enterprise customers under NDA on request.
What a real audit summary should contain
A defensible audit summary, at a minimum, includes:
- The name of the tool and version audited.
- The audit date and the date range of the underlying data.
- The auditor's name and statement of independence.
- The total sample size and the breakdown by protected category.
- Whether real applicant data or test data was used.
- The selection rates and impact ratios for each protected category and each intersectional category.
- A description of any disparities flagged and the mitigations applied.
- The cadence of the next planned audit.
Anything shorter is marketing.
What HR teams should actually ask a vendor
When evaluating an AI interview platform, four questions cut through the noise:
- "Can I see the full text of your most recent bias audit, not just the summary?"
- "What was the sample size for each protected category in the audit?"
- "What disparities did the audit surface, and how did you respond?"
- "When is the next audit scheduled, and what triggers an off-cycle audit?"
Vendors who answer all four with specifics are at the level you want. Vendors who deflect on any of them are not.
How structured AI interviews help in the first place
It is worth saying explicitly: structured AI-assisted interviews tend to audit better than unstructured human interviews, not worse. Decades of selection research, summarised in Schmidt and Hunter's meta-analyses, show that structured interviews have higher predictive validity and produce less subgroup variance than unstructured ones.
The bias risk in AI hiring is real, but the counterfactual is not a perfectly fair human process. The counterfactual is a process whose biases are harder to measure because there is no audit trail. A structured AI interview with a published bias audit is, on most reasonable readings of the evidence, a fairer process than a recruiter screen with no records.
Where to go next
If you want to see what a structured, evidence-backed AI interview looks like before evaluating its audit profile, the Voxxhire demo walks through a complete candidate flow in under three minutes.
For a complementary view on candidate data handling under regional law, see our pieces on Saudi PDPL candidate data handling and UAE labour law and AI interviews in 2026.
For an applied example of structured AI interviews in a higher-education context, see the University of Birmingham Dubai pilot case study.
This article is general guidance on bias audit methodology for AI hiring tools. It is not legal advice. The specific obligations applicable to any AI hiring system depend on jurisdiction, role, and use case. Always consult qualified employment law and data protection counsel.