Are AI Hiring Assessments Biased? How PRODICTA Approaches Fairer Candidate Assessment
The honest answer is that some are, in specific and well-documented ways, and that no vendor can promise you a hiring process with no bias in it. What an employer can reasonably ask is narrower and more useful: what does this assessment actually measure, what information reaches the thing that produces the score, what is done to check the outcomes across groups, and who makes the decision at the end. This article answers those four questions for PRODICTA. It starts with the concern, because the concern is legitimate.
The concern is real, and it predates AI
Most of the assessment market still runs on self-report. The candidate is shown statements about themselves and asked how strongly they agree, and a personality profile is inferred from the pattern. The problems with this in a selection context were documented long before anyone put a model on top of it.
The first problem is that the format itself can disadvantage people for reasons that have nothing to do with the job. In a 2023 commentary in Industrial and Organizational Psychology, Wegmeyer and Speer examine how conventional personality testing behaves when the person sitting it is neurodivergent. Their argument, in outline, is that these instruments are built around typical patterns of self-description, so an autistic candidate or another neurodivergent applicant may answer items in a way that produces a poor profile without that profile saying anything about how they would perform in the role. The disadvantage comes from the measure, not from the work, and because the scoring is opaque the candidate has no way to see or contest it. The authors argue for assessment that is tied to the actual demands of the job rather than to a generic picture of the desirable employee (Wegmeyer, L., and Speer, A., 2023, Examining personality testing in selection for neurodiverse individuals, Industrial and Organizational Psychology, 16(1), 61 to 65, DOI 10.1017/iop.2022.102).
The second problem is reliability rather than fairness, and it is separate. People applying for a job do not always answer a personality questionnaire the way they would answer it privately. In a study published in Personnel Review, Griffith, Chmielowski and Yoshita compared the personality scores of real applicants with scores from the same people re-tested later under instructions to answer honestly. Depending on the confidence interval applied, between roughly a third and a half of applicants had elevated their scores when applying, and the elevation was large enough to change the rank order that a hiring decision would have been based on (Griffith, R. L., Chmielowski, T., and Yoshita, Y., 2007, Do applicants fake? An examination of the frequency of applicant faking behavior, Personnel Review, 36(3), 341 to 355). Later work indexed on PubMed Central builds on this and finds that the ability to fake well is itself linked to cognitive ability and job knowledge, which means a self-report score partly measures how well the candidate understood what the employer wanted to hear.
Put those two findings together and the shape of the problem is clear. A self-report instrument can penalise a candidate whose self-description is atypical and reward a candidate whose self-description is strategic, and in neither case has it looked at the work.
AI adds a third layer on top of whatever it is given. A model trained on who was hired or promoted before will learn who was hired or promoted before, including every pattern in that history that the employer would not choose to repeat. A model given a CV can read age from dates, sex from a name and background from a postcode without being told to. None of that is exotic. It is what happens when the inputs carry the signal and nobody removes it.
Self-description versus behavioural evidence
The design choice that matters most is what the assessment asks the candidate to produce. PRODICTA does not ask candidates to rate themselves against statements. It builds a situation from your vacancy and puts the candidate into decisions that have no single right response: conflicting priorities, incomplete information, competing consequences, someone waiting on a call that costs something whichever way it goes. The candidate has to decide, explain and prioritise, and the report scores what they did.
This addresses both research findings directly, not by being cleverer about self-report but by not using it. There is no personality profile to infer, so an atypical way of describing oneself has nothing to score against. There is no obviously desirable answer to a situation with real trade-offs, so there is much less to be gained from working out what the employer wants to hear. What the candidate submits is a record of choices in a context built from the job, and the scoring criteria are the behaviours the role needs: what they prioritised, what they traded off, whether they took ownership of a specific outcome, how their judgement held up when the situation moved.
The scoring instructions are explicit about what is out of bounds. The model is told never to score personality type, communication style as a preference, confidence, enthusiasm, cultural fit as a personality judgement, or any trait that could correlate with a protected characteristic under the Equality Act 2010, and never to penalise spelling, grammar or writing style. It is told to score decisions, actions and reasoning. Every strength and watch-out in the report quotes what the candidate actually wrote, so an employer can check the evidence rather than take the score on trust. The longer comparison is in scenario-based assessment vs psychometric testing.
One standard, a different situation for each candidate
When a vacancy is created, PRODICTA builds a profile of the role and a set of scoring criteria from it. Each candidate who opens their assessment link is then generated their own scenarios from that same profile, so two people applying for the same role face different situations built to the same role-specific standard and scored against the same underlying criteria.
The fairness point here is the one that gets missed. A single shared test rewards whoever saw it first. Answers travel between candidates, coaching services build model responses, and the second week of a campaign is assessed on preparation rather than judgement. Generating each candidate their own situation removes most of what could usefully be passed on. What has to stay fixed is the standard, and it does: the criteria, the role profile and the seniority calibration are shared, and only the situation varies. Equivalent rather than identical is the design goal, because identical is the version that is easy to share and different-without-a-standard is the version where one candidate got an easier draw.
Where a candidate uses AI to write their answers, the report says so. Response integrity analysis looks for AI-assisted text, pasted content, rushed submissions and inconsistent quality across scenarios, and flags what it finds for the employer to probe at interview rather than acting on it.
What the scoring model is never given
The call that produces a candidate's scores receives the role, the seniority the role was set at, the scoring criteria, the candidate's responses and the timing of those responses. It is not given the candidate's age, sex or ethnicity. It is not given their name. It is not given a CV, because PRODICTA does not take one. Where an employer chooses a spoken response format, the scoring works from the transcript. The candidate is not asked for a date of birth, a nationality or an education history as part of the assessment.
This is a narrower claim than "the AI is blind" and it is meant to be. The scoring model cannot use a characteristic it was not given, and the characteristics that carry the most risk are simply absent from its input. What it can see is the candidate's own writing, and writing carries some signal about a person. The scoring rules address that by directing the model at the decision and away from the style, and by quoting the evidence for every judgement so that a reviewer can see what was weighed.
Adverse impact monitoring: what PRODICTA does and does not do
Removing characteristics from the input is necessary but it is not sufficient. A process can be blind to sex and still produce different outcomes by sex, which is the whole point of indirect discrimination under the Equality Act 2010. So the outcomes have to be checked. This is where it is important to describe exactly what the product does, because compliance claims in this market are routinely overstated.
Before an assessment starts, and on a page that says in terms that the assessment has not begun and that skipping changes nothing, candidates are invited to share an age band, a gender and an ethnicity. It is optional. Those answers are not shown to the employer against the candidate's name and they play no part in the score. Their only use is monitoring: they let PRODICTA compare how groups fared on the same assessment. They are asked before rather than after the assessment so that someone who starts and does not finish is still represented, because a monitoring sample made up only of people who completed can never see a group that dropped out.
For an employer, the monitoring view then applies the four-fifths rule: a group's pass rate on an assessment should be at least 80% of the highest group's rate. It runs per assessment, only once at least ten candidates on that assessment have both been scored and self-reported, and it only compares groups of five or more. Smaller groups are reported as insufficient data rather than compared, because a pass rate over three people is noise. Alongside the group comparison, three checks run on the shape of the score distribution itself. Where the sample is large enough, the result can be exported as a dated document that records which checks ran, what gate they were computed on and whether each passed or needs review.
Three limits are worth stating plainly. First, the pass rate is currently computed on a score threshold, which is a proxy for the hiring decision rather than the decision itself, and the document says so rather than implying it monitored who was hired. Second, when the rule cannot be applied because there is not enough demographic data, the result is recorded as unknown, not as a pass. A blank sample is not a finding of no adverse impact and the product does not treat it as one. Third, the four-fifths rule is a screening statistic, not a legal test, and it says nothing about cause. A flag means look closer. The fuller guide to what the rule can and cannot tell you, and what to do when it flags, is in Equality Act 2010 in hiring: adverse impact and how to check for it.
What PRODICTA does not do is certify your hiring process as fair. It monitors the assessment stage it runs, on the data candidates chose to give it, and gives you a documented result you can act on. The rest of your process, from the advert to the interview panel, is yours to monitor.
The employer decides
PRODICTA does not make hiring decisions. It does not progress, hold or reject anyone on its own, and it does not communicate a decision to a candidate on your behalf. The report gives you scores, the evidence behind them, strengths and watch-outs to probe, and questions to ask at interview. A person on your team reads it and decides what happens next.
That is a design position rather than a limitation. An assessment can produce better evidence than an interview or a CV, and the research above is part of why. It cannot know what you know about your team, the role's history or the candidate in the room, and a system that acted on its own output would turn every limit of the assessment into an automatic outcome for a real person. The value is in giving the decision-maker something better to decide with, and in being able to show afterwards what the decision was based on.
What this article does not claim
It does not claim that PRODICTA is free of bias, because no process that involves people and language can honestly claim that. It does not claim that scoring behaviour rather than self-description removes every way a candidate could be disadvantaged; it removes specific, documented ones. It does not claim that the adverse impact monitoring has audited your hiring, because it monitors one stage on voluntary data and tells you what it found. What it claims is that the assessment scores what a candidate did in a situation built from the job, that the characteristics the law protects are neither requested as part of the assessment nor available to the scoring model, that the outcomes are checked across groups where there is enough data to check them, and that a human on your side makes the decision.
Where PRODICTA fits
PRODICTA helps employers decide who should progress and who to hire, using evidence from the work candidates will actually have to do. The fairness design above is part of how that evidence is produced, not a compliance feature bolted on afterwards: evidence from decisions rather than self-description is what makes the report worth deciding on and what keeps the characteristics that should not matter out of it. How it works covers the steps from job description to report, and PRODICTA for employers sets out which hiring decisions it is built to support. The demo walks a real role through so you can see what the scoring saw and what it did not.