The hidden risk in AI recruiting: training tomorrow on yesterday’s bias

Blog

When firms talk about bias in AI recruiting, they usually start with the information a system can see. Remove age, gender, and candidate photos, and the assumption is that the process becomes more objective.

Those restrictions help. They also leave a harder question unanswered: what is the system learning to treat as a good decision?

There are two primary channels by which bias may enter an AI system: via the data it’s fed, and the ways in which that data’s used.

A purpose-built recruiting platform can make input restrictions easier to apply consistently. For firms using generic tools, those protections depend on the process around the tool: what recruiters may submit, which criteria they use, and how recommendations are reviewed.

Recruiters decide who gets contacted, submitted, interviewed, and placed. When those decisions become training data or feedback for future recommendations, the system can absorb the preferences behind them, including preferences nobody has examined. Yesterday’s recruiting habits can become tomorrow’s definition of a strong candidate.

Why restricting fields is only a starting point

Removing sensitive fields can reduce the risk of bias, but it does not remove every clue about a candidate. A graduation year may suggest age, while a name or location may lead to assumptions about background. Firms still need to examine how the system uses the information that remains and whether it is relevant to the role.

The same principle applies to generic AI tools. Firms can redact documents, standardize inputs, and limit approved uses. Without an established process, however, those decisions fall to individual recruiters each time they use the tool. The risk grows when the firm assumes someone else has already handled it.

An image that reads, "AI can reproduce existing recruiting biases and introduce new patterns of its own"

Imagine a team that consistently favors candidates from a handful of familiar firms. Those candidates receive more outreach, more interviews, and more opportunities to be placed. If a system later trains on those placements as examples of success, it may reinforce the team’s original preference. Candidates outside that pattern receive fewer opportunities to produce the outcomes the system rewards.

Even without retraining, biased instructions and examples can still shape a tool’s recommendations.

Restricting sensitive inputs does not, by itself, correct biased selection criteria or feedback. A placement record tells you who was placed. It cannot, by itself, tell you whether the strongest candidate was considered.

When the score missed the judgment

In law, accounting, and financial advisory, credentials and experience matter. So does the ability to exercise judgment when a situation becomes complicated. If an evaluation rewards titles, tenure, and credentials without defining how to assess judgment, the ranking can look precise while overlooking something essential.

Define the situations the person must handle, gather evidence through consistent questions or work samples, and require the evaluation to connect its conclusions to that evidence. Leaving judgment undefined creates room for both automated assumptions and human favoritism.

An image that says, "A system can get better at predicting your decisions without getting better at identifying talent."

In an experiment using roughly 361,000 fictitious resumes, researchers found that five language models assigned different scores in response to names signaling race and gender, after controlling for qualifications. These were experimental screening results, rather than observed hiring outcomes, but they demonstrate how seemingly ordinary resume content can affect an evaluation.

Applied at the scale of the US labor force, the same researchers estimated that comparable bias in a single model could affect the hiring outcomes of hundreds of thousands of qualified applicants, even in systems used only for entry-level screening.

The bias built into the incentive

Not every distorted recommendation begins with a demographic preference. Some begin with how recruiters get paid.

Consider a compensation structure that pays a recruiter more for placing a candidate with their own client than with a colleague’s. That creates an incentive to route strong candidates toward the recruiter’s accounts, even when another opportunity may be a better fit. If a system uses those placements to learn which candidates belong with which clients, it can absorb the effects of the compensation structure.

The record shows a completed placement. It may say nothing about the better opportunity that was passed over. Before using outcomes to improve recommendations, firms need to examine the incentives that produced those outcomes.

Governance has to be demonstrated

A purpose-built platform can provide consistent inputs, defined workflows, and records of decisions. Firms should verify that those controls exist, understand their limits, and confirm that people use them.

For generic tools, the questions are equally concrete. What information may recruiters submit? What criteria must they use? Which recommendations require review? What record connects the input, the output, and the eventual decision?

The label on the software cannot answer those questions.

An image that says, "Restricting sensitive inputs does not, by itself, correct biased selection criteria or feedback."

Within a firm, oversight weakens when recruiters stop questioning the output. A review step provides little protection if the reviewer simply accepts the ranking. Without a reliable record connecting the information submitted, the recommendation, and the eventual decision, managers cannot readily investigate how a pattern developed.

A study by researchers at Stanford, Chapman, and Northeastern examined more than four million applications across 156 employers using algorithms from one vendor. Among applications from candidates who identified as Black, about 26% went to positions that met the study’s adverse-impact criteria. These findings concern screening recommendations, not confirmed final hiring decisions.

The researchers also found repeated rejection recommendations across positions more often than expected under their comparison model. For employers, that raises a question about shared infrastructure: how much independent consideration does a candidate receive when different companies rely on similar screening systems?

What managing the risk requires

Firms need controls that address both the information entering the process and the decisions coming out of it.

There are four clear actions for doing so: control the inputs, define success carefully, build tool literacy, and audit the process over time.

  • Control the inputs. Limit unnecessary sensitive information across fields, documents, and prompts. Test whether other information still produces unjustified differences.
  • Define success carefully. Examine what submissions, interviews, and placements actually measure before using them as learning signals. Include evidence of performance and retention where appropriate, while checking those measures for bias, too.
  • Build tool literacy. Train recruiters to distinguish evidence from inference, challenge unsupported conclusions, and document meaningful reasons for overriding recommendations.
  • Audit the process over time. Review outcomes by role and stage, including samples of candidates screened out. Investigate patterns and correct the data, criteria, workflow, or model responsible. Check again after changes.

These responsibilities continue after implementation. A change in the model, the workflow, or the definition of success can change who gets considered.

Bias must be managed continuously

AI can reproduce existing recruiting biases and introduce new patterns of its own. Managing that risk requires examining the decisions a system relies on, the outcomes it rewards, and the candidates its recommendations leave behind.

The question extends beyond whether the technology works as designed. Firms also need to ask whether they have taught it to value the right things.

A system can get better at predicting your decisions without getting better at identifying talent. Firms need to know which improvement they are paying for.

Get Started with Bear Claw ATS Platform Today.

We’re here to help. Try us free today or schedule a 1:1 customized demo with our team to see precisely how Bear Claw ATS platform can help you and your teams reach your goals.

Schedule My Demo