← All articles
Recruitment AI · 10 min read

AI Resume Screening Bias: What Recruiters Must Know

One peer-reviewed study found AI resume screening favored white-associated names in most tests. Here is the real evidence, and how to screen fairly.


The fear behind most recruiters' hesitation about AI screening is not abstract. It is a specific, well-founded worry that a model trained on messy, human-generated data will quietly repeat that data's worst habits, at scale, without anyone noticing until a candidate or a regulator does. That worry is backed by real, peer-reviewed research, not just anecdote. It is also a solvable problem, not a reason to avoid AI screening altogether. This piece walks through the actual evidence and what to do about it.

It is worth separating two questions that usually get merged into one. The first is whether AI resume screening can be biased. The research below answers that clearly: yes, and the bias is measurable, specific, and worse for some groups than others. The second, harder question is whether that makes AI screening worse than the alternative, which for most agencies is unstructured human screening with no audit trail at all. That second question has a more complicated answer, and it is the one this piece actually tries to settle.

Is AI Resume Screening Really Biased? What the Research Shows

Yes, in the tools and studies that have been tested so far, though the picture is more specific than the headlines suggest. The clearest evidence comes from a 2024 peer-reviewed study out of the University of Washington, and it is worth reading the actual finding rather than the rounded-off version that circulates online.

85.1%

of statistical tests for racial bias in AI resume screening significantly favored white-associated names over Black-associated names, across three embedding models and nine occupations.

Source: Wilson & Caliskan, AIES 2024 (University of Washington)

That number needs one clarification most secondhand summaries leave out. The researchers ran 27 statistical significance tests in total, three language models tested across nine occupations. White-associated names were significantly preferred in 85.1 percent of those 27 tests. It is a finding about how consistently the bias showed up across models and job categories, not a claim that 85 percent of individual resumes or applicants were affected. The underlying study drew on more than three million resume-to-job comparisons to build that picture, which is a large and serious dataset, but the 85.1 percent figure specifically describes the tests, not the comparisons. Getting that distinction right matters if you are going to cite this to a client or use it to justify a compliance decision.

The same study found gender bias in the opposite direction, and smaller: male-associated names were preferred in 51.9 percent of gender tests, against female-associated names being preferred in 11.1 percent. The most severe effect showed up at the intersection: Black male candidates were disadvantaged in up to 100 percent of the relevant tests, worse than either race or gender bias alone. Bias does not distribute evenly. It concentrates hardest at the overlap of more than one protected characteristic.

How AI Bias Actually Happens

AI screening tools do not decide to be biased. They inherit it from three sources, usually stacked on top of each other.

  • Training data. A model trained on historical hiring decisions learns whatever pattern is in that history, including who got hired for the wrong reasons in the past.
  • Proxy variables. Even when race, sex, or age are never directly used, a model can learn to treat zip code, school name, or graduation year as a stand-in for a protected characteristic, reproducing the same discrimination indirectly.
  • Name signals. Names carry strong demographic associations in the data language models are trained on, which is exactly what the University of Washington study measured directly.

None of these require anyone to have intended discrimination. That is precisely what makes it dangerous. A biased result can come out of a tool nobody built with bad intent, which is also why "the AI decided, not me" has never held up as a legal or ethical defense.

A concrete version of the proxy variable problem shows up constantly in recruiting specifically. A model asked to favor "culture fit" or "strong communication skills," phrases that sound neutral, can quietly learn to associate those labels with graduates of a small list of universities, or with resumes that use a certain writing style more common among native English speakers. Nobody told it to discriminate by school or by first language. It found the pattern in the data anyway, because the data itself was shaped by decades of human decisions that were not always fair to begin with.

Beyond That One Study: What Other Research Confirms

The University of Washington findings are not an outlier. A 2025 study published in PNAS Nexus tested five commercial language models against 361,000 real resumes and found a related but distinct pattern: models favored female candidates overall, while still specifically penalizing Black male applicants, echoing the intersectional finding above from a different angle. Bloomberg ran its own investigation into GPT-based resume ranking in 2024 and found bias that varied by job type rather than applying uniformly, which is a useful reminder that "is this tool biased" does not have one universal answer across every role you screen for.

Two more findings are worth knowing specifically because they complicate the simple version of this story. A 2025 University of Washington follow-up study found that humans who were shown biased AI recommendations followed them roughly 90 percent of the time, even when the recommendation was demonstrably skewed. That is the strongest evidence available that a human reviewing AI output is not automatically a safeguard, unless that human is actually trained and expected to catch it. And a 2026 study found that bias direction and severity varies by model generation, with newer models tested showing smaller or reversed patterns compared to older ones. Bias is not a fixed, universal property of "AI." It depends on which model, trained when, tested on what.

For context, biased human screening predates AI by decades. A well-known 2004 study sent matched resumes under stereotypically white names like Emily and Greg, and stereotypically Black names like Lakisha and Jamal, to real job postings, and found the white-sounding names received about 50 percent more callbacks. That was pure human bias, no software involved. AI screening did not invent this problem. The real question is whether a given tool makes it better, worse, or just faster.

The Legal and Compliance Stakes for Recruiting Agencies

Bias in an AI screening tool is not just a research finding. It is the exact scenario several current laws exist to catch. New York City's Local Law 144 specifically requires a bias audit measuring outcomes by race and sex before you can rely on an automated tool for hiring decisions involving NYC-based roles. Illinois's HB 3773, in force since January 2026, makes it a civil rights violation if AI use has a discriminatory effect, whether or not anyone intended it. If a tool produces the kind of skew documented above and you have not checked for it, you are not just facing a research problem. You are facing exactly the exposure those laws were written to address.

Small agencies sometimes assume compliance rules like these are an enterprise problem, written for companies with in-house counsel and a compliance department. The rules do not actually distinguish by agency size. A five-person shop placing candidates into NYC-based roles carries the same bias audit obligation as a five-hundred-person staffing firm doing the same thing. Nobody enforces it more gently because your agency is small. If anything, that makes getting this right on a small desk more urgent, not less, since there is no compliance team downstream to catch a mistake before it reaches a candidate.

A biased result coming out of a tool nobody built with bad intent is still a biased result. That is what makes it dangerous, and why checking for it cannot be optional.

What Not to Feed an AI Screening Tool

A few concrete habits meaningfully reduce risk before you even get to auditing a tool formally.

  • Do not let a tool see a candidate's name, photo, or age during initial screening if the platform allows blind review.
  • Watch for zip code, school name, and graduation year being weighted heavily. All three commonly stand in for protected characteristics.
  • Never let a tool auto-reject without a visible reason attached. If you cannot see why a candidate was screened out, you cannot check whether the reason was fair.
  • Test the tool with a matched set of resumes that are identical except for a name signal, the same method the University of Washington study used, before trusting it on real candidates.

How to Screen Resumes with AI Responsibly

The goal is not to avoid AI screening because bias is possible. Human screening carries the same risk, demonstrated for over two decades, without any of the paper trail an AI tool can actually produce. The goal is a tool that shows its reasoning and keeps a person in the loop who is actually positioned to catch a problem, not just rubber-stamp the output.

Creo Access

Skill: screen

"Screen these 40 CVs against the senior backend brief."
Strong match6 candidates, ranked with reasons shown
Worth a look11 candidates, flagged with why, not silently dropped
Not a fit23 candidates, reason logged, nothing auto-deleted

Every candidate gets a visible reason. Nothing is rejected silently.

Before you trust a screening tool with real candidates

  • Require a visible reason for every recommendation, not just a ranked list.
  • Test with matched resumes that differ only by name before rolling a tool out.
  • Keep a human reviewing outcomes on a schedule, not just when something looks wrong.
  • Check whether your jurisdiction requires a formal bias audit, and if so, get one done before you rely on the tool.
  • Re-test periodically. Bias varies by model version, and the tool you audited a year ago may not be the tool you are using today.

Screening fairly with AI is possible. It just requires treating the tool the way you would treat a new junior recruiter: useful, worth training, and never left completely unsupervised on a real decision.

Creo Access

Creo Access is in private beta.

Request access and you'll receive product updates, testing opportunities and launch information.

Request beta access

Keep reading