How to Write a Screening Rubric That Survives a Bias Audit
A rubric fails an audit for boring reasons: criteria nobody defined, ratings nobody can reconstruct, and traits that were never part of the job. How to write anchored, job-related criteria, which proxies to cut, and how to audit your own scoring before someone else does.
By the InterviewAgent.ai team
July 2026 · 9 min read
First-round interview
Candidate consented · AI-conducted00:00 · AI Interviewer
Run the sample interview to watch the AI ask, follow up and score against your rubric.
Scored report
RubricThe report assembles after the interview: overall score, rubric, highlights and a recommendation. You make the final call.
Highlights
Recommendation only · a recruiter makes the final decision
Ranked shortlist
Live, interactive · consent-first · no signup needed
Structured & consistent · bias-audited (EEOC / NYC Local Law 144) · you make the final call
A screening rubric survives a bias audit when it scores job-related behavior, uses anchored levels instead of a vague 1 to 5 scale, is written before anyone is interviewed, and keeps the evidence behind each score. An auditor is checking whether your selection rates differ by sex and by race or ethnicity, and whether you can explain the scores that produced them. Rubrics fail audits for boring reasons: criteria nobody defined, ratings nobody can reconstruct, and traits that were never part of the job.
Most hiring teams write a rubric the week before a role opens, use it inconsistently for a month, and never look at it again until something goes wrong. That is fine right up until an auditor, a plaintiff's lawyer or your own general counsel asks why the two candidates with identical experience scored a 4 and a 2. At that point the rubric is either your best evidence or your biggest problem, and which one it is was decided months earlier.
This is a practical guide to writing the first kind. It assumes you are screening at some volume, that scores are helping decide who advances, and that you would rather find the weaknesses yourself than have someone else find them.
What does a bias audit actually check?
Under NYC Local Law 144, a bias audit is an independent evaluation of whether an automated employment decision tool produces different selection rates by sex and by race or ethnicity, including the intersections of those categories. It has to be conducted within the previous 12 months by someone independent of both the employer and the vendor, and a summary has to be published where candidates can find it. The duty sits with the employer using the tool.
The audit measures outcomes, not intentions. Nobody reads your rubric and pronounces it fair. They compare who advanced against who applied, broken down by category, and a gap prompts the follow-up question: what were these scores based on? That is the moment a rubric earns its keep. A criterion like "communication" with no definition cannot answer it. A criterion that reads "explained a technical decision so a non-specialist could follow it, with a concrete example" can.
Two things follow from this. Documentation is not paperwork you do after the fact, it is the substance of the defense. And a rubric that produces a disparity is not automatically illegal: it is a signal that the criteria need to be justified as job related and consistent with business necessity, which is much easier when they were derived from the job in the first place.
Write the rubric before you see any candidates
This is the single highest-value rule and the one most often broken. A rubric written after you have met a few applicants is a rubric shaped by them, usually by the one everybody liked. Criteria drift toward that person's background, and the standard stops being "what the job needs" and becomes "how similar is this person to the one we already liked".
Start from the work instead. List the five or six things somebody in this role does in a normal month, in plain language. For each, name what a strong answer sounds like and what a weak one sounds like. If a criterion cannot be traced back to something on that list, it does not belong in the rubric, however much it feels like a signal.
Cap it at five or six criteria. Long rubrics look rigorous and score worse, because interviewers stop reading them and fall back on impressions while filling in numbers afterward to match. A short rubric people actually use beats a thorough one they ignore.
Use anchored levels, not a 1 to 5 feeling
An unanchored five point scale is where consistency dies. Two interviewers using the same scale on the same answer routinely differ by two points, because a 3 means "fine, I suppose" to one person and "did not impress me" to another. Neither is measuring the candidate.
Anchoring means writing, for each criterion, what earns each level. Three levels are usually enough. For a criterion like "handled a customer escalation":
- Level 1: Describes the situation in general terms, no specific example, unclear what they personally did.
- Level 2: Gives a specific example, describes the actions they took, outcome stated but not evidenced.
- Level 3: Specific example with their own actions distinguished from the team's, describes what they would do differently, outcome supported by a detail or number.
Notice that none of these mention confidence, polish, enthusiasm or culture fit. Every level describes the content of an answer. That is what makes the score reconstructable by somebody who was not in the room, which is exactly what an audit requires. Our interview scorecard template has a fuller worked structure to start from.
Which criteria fail an audit?
Some are obviously off limits, and no serious team includes them. The ones that cause real trouble are the proxies: criteria that look neutral and correlate closely with a protected characteristic.
- Culture fit. Almost never defined, and in practice measures similarity to the existing team. If you mean specific working behaviors, name those instead and score them.
- Communication skills, unqualified. Scores fluency and accent rather than clarity. Fix by scoring whether the point was made clearly, not how it sounded.
- Years of experience as a scored criterion rather than an eligibility gate. Correlates with age and with career breaks, which disproportionately affect women.
- Enthusiasm, energy, executive presence. Scores personality and self-presentation norms that vary by culture and by neurotype.
- Anything inferred from face or voice. Vocal tone and facial expression analysis is the weakest science in this field and the fastest route into Illinois and Maryland trouble. Illinois HB 3773 also specifically names zip code as an impermissible proxy, which is a useful reminder of how indirect a proxy can be.
Salary expectations deserve a separate mention, because it is the criterion teams forget is a criterion. Screening candidates against what they currently earn imports every historical pay gap directly into your funnel, and it is exactly the practice pay transparency laws were written to interrupt. Decide the range from the role rather than the candidate: if you have not set defensible pay bands and posted salary ranges yet, do that before it becomes a screening question at all.
Score the answer, then the candidate, and keep the evidence
Score each answer as it happens, against its criterion, before forming an overall view. Halo effects are strong and fast: one impressive answer lifts every subsequent score if you let a general impression form first. Per-answer scoring, then summing, is a small procedural change that measurably reduces this.
Attach the evidence to the score. For a human interviewer that means a quote or a close paraphrase of what the candidate actually said, written at the time. For an AI screening interview it means the transcript sentence that produced the rating. Six months later, "3, good example of ownership" is worthless and "3, said he rewrote the reconciliation process after the Q2 miss and cut close time from nine days to four" is a defense. This is where structured tooling earns its cost, because the evidence is captured automatically rather than depending on whether anyone had time to type notes between calls. An AI interview agent attaches the transcript to every rubric score as the interview happens. Our page on candidate interview scoring covers how the scoring and evidence link works in practice.
Ask every candidate the same questions
Rubric consistency is undone by question inconsistency. If one candidate was asked about a failed project and another about their strengths, their scores on "self-awareness" are not comparable, and no rubric can repair that after the fact. Same questions, same order, for every candidate for a role. Follow-ups can and should adapt to what someone actually says, but the core set does not move.
This is also the practical argument for structured AI screening: consistency is trivial for software and genuinely hard for humans, who are tired by the fourth call and warm to candidates they like. If you are considering that route, the compliance duties that come with it are set out on our AI hiring compliance page, and the question of what makes a tool a regulated automated employment decision tool is covered in our AEDT explainer.
Audit your own rubric before someone else does
Run the check yourself, quarterly, on data you already have. For each role, compare the pass rate at the screening stage by sex and by race or ethnicity where you collect it. You are not looking for perfection, you are looking for a gap large enough that you would want an explanation ready.
When you find one, look at the criterion level rather than the total. Usually one or two criteria carry the disparity, and inspecting the anchors for them is more productive than rewriting everything. Common culprits: a communication criterion scoring fluency, an experience criterion penalizing career breaks, a scenario question assuming a background that only some candidates have.
Then check the boring mechanical things, because they cause more failures than bias does. Are scores actually being entered, or is half the pipeline blank? Do two reviewers scoring the same candidate land within a point of each other? Is the rubric in use the version you think it is? A rubric that nobody follows is worse than no rubric, because it creates a paper record contradicted by the decisions.
Frequently asked questions
Does a good rubric guarantee we pass a bias audit?
No. An audit measures outcomes, so it is possible to have a well-written rubric and still show a disparity, usually from an upstream sourcing pattern rather than the scoring itself. What a good rubric guarantees is that you can explain the outcome, identify which criterion produced it, and show the reasoning was job related. That is the difference between a finding you can address and one you cannot answer.
How many criteria should a screening rubric have?
Five or six for a first-round screen. Fewer than four tends to be too coarse to separate candidates, and more than about seven degrades in practice because interviewers stop reading the anchors and score on impression. Weight them explicitly if some matter more, rather than adding extra criteria to inflate the importance of one area.
Can we use the same rubric for every role?
Only for genuinely shared criteria, and even then be careful. A shared frame is useful for something like structured problem solving, but the anchors have to be rewritten per role, because a strong answer for a warehouse supervisor and a strong answer for a controller do not resemble each other. Reusing anchors across unrelated roles is how criteria stop being job related, which is precisely the thing an audit tests.
Who should write the rubric?
The hiring manager, with a recruiter making sure it is scoreable and someone checking the criteria are job related. The hiring manager knows what the work requires, the recruiter knows whether a criterion can be assessed in 20 minutes, and the third pass catches the proxies that slip in when you are close to the role.
Do we need a bias audit if we do not use AI?
Local Law 144 attaches to automated employment decision tools, and the definition is broader than most people expect: a computational process that produces a simplified output like a score or a ranking and substantially assists or replaces discretionary decision making. A weighted scoring spreadsheet can qualify. Whether yours does is a question worth asking a lawyer rather than assuming, and our guide to US state AI interview laws sets out the tests each jurisdiction applies.
See InterviewAgent.ai screen candidates
The agent interviews every applicant with role-tailored questions, scores against your rubric, and ranks a shortlist for your recruiters. The agent advances candidates, your team decides.