AI Algorithmic Bias Audit — Hiring Tool Playbook
A Fortune 100 company uses an AI-powered resume screening tool that filters 40,000 applicants annually to a pool of 2,000 for human review. An internal data scientist flagged a 22% pass-through rate disparity between male and female applicants for technical roles. The CHRO has 30 days before the next board meeting.
When to use this playbook
- Use this playbook when the decision looks like the situation above: A Fortune 100 company uses an AI-powered resume screening tool that filters 40,000 applicants annually to a pool of 2,000 for human review.
- It is a fit when you have source files in hand and need a structured, reviewable analysis — not a generic chat answer about "Algorithmic Bias Audit — Hiring Tool".
- Do not use it as a substitute for licensed, legal, clinical, or authorized official judgment in the domain.
What you'll need
- Resume screening tool output (applicant ID, features scored, pass/fail, demographic data where available)
- Ground truth hiring outcomes for the past 2 years (who was hired, from which pool)
- Tool vendor documentation including features used and training data description
- EEOC Uniform Guidelines on Employee Selection Procedures
- Adverse impact ratio calculations from the data scientist's flagged analysis
Attachments: Documents (Documents)
The Prompt
You are an AI governance specialist conducting an algorithmic bias audit for a Fortune 100 hiring tool. I am attaching: Work only from the attached source files. If a conclusion is not supported, say so. Produce: 1. Confirm or dispute the 22% pass-through disparity using the EEOC 4/5ths rule and statistical significance testing—is this legally actionable adverse impact? 2. Identify which features in the scoring model are most correlated with the disparity (proxy variable analysis). 3. Assess whether the disparity persists after controlling for job-relevant qualifications—is it differential prediction or differential treatment? 4. Calculate the potential litigation exposure under Title VII and what the remediation cost estimate is. 5. Tell me what the CHRO should present to the board and what immediate operational changes should be made pending a full audit. Call out where independent models are likely to disagree, and list follow-up documents a reviewer should request.
What to expect
- Adverse impact statistical analysis with legal threshold comparison
- Proxy variable identification
- Differential prediction vs. differential treatment assessment
- Title VII litigation exposure estimate
- Board presentation language and operational remediation steps
Review before you act
- Validate this output against source files before relying on it: Confirm or dispute the 22% pass-through disparity using the EEOC 4/5ths rule and statistical significance testing—is this legally actionable adverse impact?.
- Validate this output against source files before relying on it: Identify which features in the scoring model are most correlated with the disparity (proxy variable analysis).
- Validate this output against source files before relying on it: Assess whether the disparity persists after controlling for job-relevant qualifications—is it differential prediction or differential treatment?.
- Validate this output against source files before relying on it: Calculate the potential litigation exposure under Title VII and what the remediation cost estimate is.
- Confirm every cited figure, date, counterparty, or requirement against the attached originals — models compress and can drop a qualifier.
- Treat disagreement between models as a review item, especially on classification, materiality, and recommended next action.
- Do not authorize an operational, clinical, legal, credit, or enforcement action solely because the models agree.
Why compare models on this
For Algorithmic Bias Audit — Hiring Tool, running the same attachments across independent models is useful because the hard part is classification and completeness, not fluency. The workflow is already designed to surface adverse impact statistical analysis with legal threshold comparison; proxy variable identification; differential prediction vs. differential treatment assessment; title vii litigation exposure estimate. Those are comparison artifacts — they only exist if more than one model runs. Risk-tier assignments and 'high-risk system' calls vary with how a model reads a use-case description. Comparison exposes those classification fights before they reach an exam.
See governed multi-model AI on your own prompt
Compare GPT-5, Claude, and Gemini side by side, with human review and a decision record built in.

