InterviewAgent.ai

AI screening · Candidate interview scoring

Interview scorecard software: automate interview scorecards and candidate interview scoring

The short answer

Candidate interview scoring means grading each answer against a written rubric rather than a gut impression, so two candidates who gave the same quality answer get the same number. InterviewAgent.ai scores every first-round interview against anchored 1 to 5 criteria you define, links each score to the transcript line that earned it, and ranks the shortlist. A recruiter reviews the evidence and decides.

Last updated August 2026

Candidate interview scoring goes wrong the moment two interviewers grade the same answer differently. Without a shared rubric, scores reflect the reviewer mood and biases as much as the candidate, and the results are impossible to compare or defend.

InterviewAgent.ai brings rubric-based candidate interview scoring to the first round. The agent conducts a structured screen, then scores each answer against the same rubric for every candidate, producing comparable scorecards with the transcript evidence behind each rating. Recruiters review the ranked results and make the final decision, while the scoring stays transparent, consented, AI-disclosed, and bias-audited for EEOC and NYC Local Law 144.

Interview · score · rank · recruiter reviews · from $149/mo

Interview Studio

First-round interview

Candidate consented · AI-conducted

00:00 · AI Interviewer

Run the sample interview to watch the AI ask, follow up and score against your rubric.

Scored report

Rubric

The report assembles after the interview: overall score, rubric, highlights and a recommendation. You make the final call.

score

Highlights

Recommendation only · a recruiter makes the final decision

Ranked shortlist

Live, interactive · consent-first · no signup needed

Score /100 Ranked #1 of Transcript + highlights ready

Structured & consistent · bias-audited (EEOC / NYC Local Law 144) · you make the final call

ROLE-TAILORED RUBRIC-SCORED ATS-READY

Human-in-the-loop you decide

EEOC and LL144 bias-audited

Why it works

What your team gets with candidate interview scoring

One rubric for all

Every answer is scored against the same rubric, so candidates are compared on identical criteria rather than on who interviewed them.

Evidence per score

Each rating links to the transcript moment that earned it, so a scorecard is explainable rather than a number you have to trust blindly.

Defensible results

Consistent, transparent scoring with bias auditing gives you results you can stand behind under EEOC and NYC Local Law 144.

What it handles

Interviewed, scored and shortlisted on autopilot

The agent invites each applicant, runs a role-tailored screening interview by voice or video, asks smart follow-ups, scores every answer against your rubric, and advances the strongest candidates into a ranked shortlist for your recruiters to review.

  • Scores every answer against a consistent rubric
  • Produces comparable scorecards across all candidates
  • Links each score to transcript evidence
  • Ranks a shortlist for recruiter review
  • Keeps scoring transparent and bias-audited
SHORTLIST Interviewing
#1 Maya R. 92
#2 Devon L. 88
#3 Priya S. 81 Review
Rubric scored 48 screened this week

Why InterviewAgent.ai

One agent that runs the whole first round

Not a one-way video tool, not a six-figure assessment suite, and not a staffing agency. Interview, score, rank and hand off in one place, shaped to the roles and rubric you already hire on.

Interviews every applicant

A role-tailored screening interview runs by voice or video, with smart follow-ups, on the candidate's schedule. Applicants consent and are told they are speaking with AI, so no qualified person waits days for a first call.

Scores to your rubric

Every answer is scored against the same structured rubric, with transcripts and highlights, so candidates are compared consistently and the scoring stays bias-audited against EEOC guidance and NYC Local Law 144.

Ranks the shortlist

The strongest candidates are advanced into a ranked shortlist your recruiters review. The agent never auto-hires or rejects, it only surfaces who to talk to next, and your team makes every decision.

At a glance

A 1 to 5 interview scoring scale, with anchors that mean something

Score What it means What the answer looked like
1 Well below the bar No relevant example. Talks about the topic in general terms, never about anything they did
2 Below the bar A real example, but thin: no context, no specifics, no outcome, and it falls apart under a follow-up
3 Meets the bar A concrete example with situation, action and result. Answers the follow-up without prompting
4 Above the bar Specific, quantified, and shows reasoning about the trade-offs they chose between
5 Well above the bar All of the above, plus what they would do differently and what the experience changed about how they work

Anchors are written before screening starts. A scale without anchors is just five different opinions.

How do you score a candidate in an interview?

You score against defined criteria, not against the other candidates and not against your impression of the person. Pick the three to five competencies the job genuinely requires, write what a weak, adequate and strong answer looks like for each one, and grade every candidate on that. The scale itself matters less than the anchors: a 1 to 5 with written descriptions beats a 1 to 10 where nobody knows what a 7 means.

Score answer by answer while the evidence is in front of you, and record the quote that drove the rating. That last part is what turns a score into something you can defend, whether the person asking is a hiring manager who disagrees or, less comfortably, a lawyer. We publish a copyable interview scorecard template with the anchors already written.

Interview answer scoring is the part teams most often skip, and it is the part that decides whether the rest of the process means anything. Grading the whole interview with one overall impression at the end reintroduces exactly the bias the rubric was meant to remove, because the last strong answer colors everything before it. Score interview answers individually, against the anchors, before you form a view of the person.

What is an interview scorecard?

An interview scorecard is a single sheet, per candidate, that lists the competencies being assessed, the score given for each, the evidence behind that score, and an overall recommendation. It exists so that the hiring conversation is about what candidates actually said rather than who is remembered most vividly by whoever spoke last.

A good scorecard has weights, because not every competency matters equally. If technical judgment is twice as important as written communication for the role, say so on the card and weight the score. What must never appear on a scorecard is anything about a protected characteristic, or a proxy for one: age, accent, where someone went to school if it is not a genuine requirement, or a note about culture fit that really means the person felt unfamiliar.

Can AI score an interview fairly?

It can be more consistent than a panel of humans, which is a lower bar than it sounds. Human interview scores drift with fatigue, with who was interviewed immediately before, and with how similar the candidate feels to the interviewer. Applying one rubric identically to 400 people is precisely the kind of task software does not get bored of.

Consistency is not the same as fairness, though, and this is where the diligence goes. A model trained on past hiring can inherit past bias, which is why the scoring here is audited for adverse impact under EEOC guidance and NYC Local Law 144, why every score shows its evidence, and why the agent never rejects anyone. It ranks and explains; a human reads the transcript and decides.

What should you never score a candidate on?

Anything not required to do the job. That covers the obvious protected characteristics under US law, and it also covers the proxies people slip into scorecards without noticing: how polished someone sounds, whether they seemed confident, whether they would be fun to grab a beer with. Confidence is not competence, and comfort is not a competency.

Scoring appearance, tone of voice or facial expression is a specific trap in video interviewing, and one of the reasons some vendors quietly retired their facial analysis features. We score the substance of the answer against your rubric, from the transcript. If a criterion cannot be written down and defended as job-relevant, it does not belong on the card.

How do you automate interview scorecards?

You automate a scorecard by moving the scoring to the moment the answer is given rather than the moment someone finds time to write it up. The card is generated as the interview happens, each criterion is scored against anchors you wrote before applications opened, and the transcript line that produced each rating is attached to it. Nobody types a card from memory two days later, which is where most of the inconsistency in interview scoring actually comes from.

What you cannot automate is the rubric itself. Someone has to decide which four to six competencies the role genuinely requires and write what a 1, a 3 and a 5 look like for each, in job-related terms. That is an hour of work per role family and it is the hour that determines whether the automation produces something defensible or something that merely looks tidy. A candidate assessment scorecard filled automatically against vague criteria is faster, not better.

The practical difference shows up at the debrief. When every card was produced the same way and cites evidence, the conversation is about what candidates said. When cards were written from recollection, the conversation is about who remembers the interview most vividly, which is usually whoever spoke last.

  • Write four to six job-related competencies with anchored 1 to 5 levels before applications open
  • Score each answer as it is given, not from memory afterwards
  • Attach the transcript line that produced every rating
  • Weight the competencies, because they are rarely equally important
  • Keep protected characteristics and their proxies off the card entirely

What is the best interview scorecard software?

The best interview scorecard software is whichever kind fills the card in for the failure your team actually has, and there are three kinds. ATS-native scorecards from Greenhouse, Lever, Ashby and Workable are already in the tool you own and depend on interviewers completing a form. Interview intelligence tools such as Metaview, BrightHire and Fabric score from a recording of an interview a person ran. AI interview agents conduct the interview themselves and score each answer as it happens.

That framing matters because most teams searching for scorecard software already own scorecards and do not know it. If your interviewers reliably fill in the ATS form, buying a second scoring product adds a step rather than removing one. The purchase only makes sense when the human step is the thing that keeps failing: cards left blank on a busy week, ratings written from memory three days later, or first-round interviews that never happen because nobody had the hours.

Price separates the groups less than capability does. Metaview publishes $60 per user a month for AI Notes and Fabric publishes $199 to $499 a month, while several vendors in the category publish nothing at all. We compare all three groups tool by tool, with each price dated, in interview scorecard software compared.

What is a candidate assessment scorecard?

A candidate assessment scorecard is the single document that lists the competencies a role is hired on, the rating scale for each one, and the evidence behind every rating. It is filled in the same way for every applicant to the same req, which is what makes two candidates comparable at the debrief instead of merely memorable.

A usable card has four parts. The competencies, four to six of them, chosen because the job genuinely requires them. An anchored scale describing in concrete behavior what a 1, a 3 and a 5 look like for each competency. The interview questions mapped to the competency they produce evidence for. And a field holding the actual answer, quote or transcript line that justified the score.

The fourth part is the one teams skip and the one that does all the work later. A card showing "Communication: 4" is an opinion. A card showing "Communication: 4" beside the candidate explaining a database migration to a non-technical stakeholder without jargon is a finding, and it is what lets a hiring manager disagree with the rating on the merits rather than on instinct.

How do you build an interview scoring system?

Build it backwards from the job rather than forwards from a template. Start with what someone actually does in the role in their first six months, reduce that to four to six competencies you would genuinely reject a candidate for lacking, then weight them, because they are almost never equally important. A support role that weights empathy and de-escalation the same as spreadsheet skill is not describing the job.

Then write the anchors before you write the questions. For each competency, describe the observable behavior at a 1, a 3 and a 5 in language a new interviewer could apply without asking what you meant. This is the step that determines whether your scores mean anything, and it is the step that gets skipped under time pressure. Unanchored scales drift toward 3 out of 5 for everyone and produce a ranking with no separation in it.

Only then attach questions, one or two per competency, asked of every candidate in the same order. If you are starting from a job description rather than a finished competency list, an interview question generator will draft that set for you to cut down. Pilot the system on five candidates and check for two failure signs: competencies where everybody scores the same, which means the question is not discriminating, and competencies where two interviewers disagree by two points or more, which means the anchors are still ambiguous. Fix those before the system is used on a live req rather than after.

What is a candidate interview scorecard, and what belongs on one?

A candidate interview scorecard is the single record that turns an interview into evidence: the competencies you agreed to assess, a rating on each one, and the specific thing the candidate said that produced each rating. Anything less than that is a note. A scorecard for interview rounds only does its job when a second reader who was not in the room can look at it and understand why the number is the number.

Four fields carry almost all the weight. The competency and its anchored definition, so everyone is rating the same thing. The rating itself on a fixed scale. The evidence, quoted or timestamped, because a rating without a source cannot be challenged or defended. And the overall recommendation held separately from the ratings, so a strong gut feeling cannot quietly rewrite the numbers that preceded it. Free text belongs at the end, not at the top, for the same reason.

Most teams already have somewhere to put this. ATS interview scorecards in Greenhouse, Lever, Ashby and Workday all support competency-based cards, and if your applicant tracking system has them, use them rather than a parallel spreadsheet. What the ATS does not solve is the part that actually fails: getting every interviewer to fill the card in the same way, on the same day, with evidence attached. That is a behavior problem, and it is the reason automated scoring exists.

  • Competency with a written anchor for each level, not a bare label
  • A rating on a fixed scale, the same scale for every candidate on the req
  • The transcript line or timestamp that produced the rating
  • The recommendation recorded separately from the ratings
  • The date and the interviewer, so the card is auditable a year later

How to score an interview so two reviewers reach the same number

Score against written anchors, in the moment, one competency at a time. Two reviewers disagree when they are rating an impression rather than a behavior, so the fix is to describe the behavior in advance: what a 1 looks like, what a 3 looks like, what a 5 looks like, in words specific enough that they could only apply to this job. Then rate each competency on its own before forming any overall view, because a single strong answer otherwise lifts every other rating with it.

The second half of the fix is timing. Candidate scoring done from memory at the end of the day is not scoring, it is recall, and recall favors whoever interviewed last and whoever was easiest to talk to. Automated interview scoring closes that gap by producing the rating while the answer is still on screen, with the evidence attached to it. That is the whole of interview scoring automation: not a machine deciding, but the rating and its source captured at the only moment they are both reliable.

This is where interview scoring software earns its cost, and it is a narrow claim. It will not tell you which competencies matter, it will not fix a rubric written in adjectives, and it should never advance or reject anyone by itself. What it does is remove the variance that comes from twelve people writing up twelve interviews at twelve different times, which in most hiring teams is a larger source of unfairness than any individual interviewer.

  • Rate each competency separately, then form the overall view last
  • Write anchors in observable behavior, never in adjectives like "strong" or "good fit"
  • Capture the rating during the answer, not at the end of the week
  • Keep the evidence with the rating so a disagreement is about facts
  • Calibrate on two real candidates before the req opens, not after the debrief goes wrong

Good questions

Questions about candidate interview scoring

Score each answer against a written rubric with anchored levels, not against your impression of the person or against the last candidate you saw. Pick three to five job-relevant competencies, define what a weak, adequate and strong answer contains, and record the quote behind every score you give.
It grades every candidate on the same defined criteria instead of subjective impressions, which removes much of the variance between reviewers. The scoring is also bias-audited for EEOC and NYC Local Law 144 and reviewed by a human before anyone advances.
Yes. You define the criteria and weights that matter for each role, and the agent scores every candidate against that rubric so the results map directly to what you care about. Weighting matters, because competencies are rarely equally important.
Every rating links back to the transcript moment that earned it, so you can read the exact answer behind a 2 or a 5. A score you cannot explain is a score you cannot defend to a hiring manager, or to a candidate who asks. The full mechanism is broken down in how AI interview scoring works.
Answer by answer, against criteria written before the first candidate applied. For each question you define what a weak, an adequate and a strong response contains, then rate the response you actually got against those anchors. The ratings roll up into a per-competency profile rather than one overall verdict, which is what lets you see whether a candidate was strong on the thing the job turns on. Scoring an overall impression at the end of a conversation is the version that does not survive scrutiny.
Well-run teams score each answer immediately after it is given, using a written rubric and a narrow scale, before moving to the next question. Scoring at the end of the interview lets the last answer color the earlier ones, and scoring after several interviews lets candidates be graded against each other instead of against the standard. The other habit that matters is recording the evidence: a note on what the candidate said that justified the rating, so the number can be reconstructed later.
Yes, when the rubric is defined before the interview runs. InterviewAgent.ai auto-scores each interview response against the anchored criteria you set for that role, and every score links to the transcript line that produced it. Auto-scoring ranks candidates for recruiter review. It never rejects anyone automatically, which is what keeps it decision support under EEOC guidance and NYC Local Law 144.
A single sheet that lists the competencies a role requires, the questions that test each one, an anchored rating scale, and space for the evidence behind each rating. It exists so every interviewer grades on the same criteria and so a hiring decision can be explained months later. A scorecard without written anchors is a form, not an assessment, because two interviewers will use the same numbers to mean different things.
A 1 to 5 scale with a written anchor for each level, and no half points. Every level needs a sentence describing what an answer at that level contains, phrased in terms of what the candidate said rather than how they came across. Wider scales create false precision that interviewers cannot apply consistently, and even-numbered scales force a side when the honest rating is that an answer was adequate.
A 1 to 5 scale with written anchors, in almost every case. Wider scales create false precision: nobody can consistently distinguish a 6 from a 7, so reviewers cluster in the middle. Five levels, each with a description of what that answer looks like, produce far more consistent scores.

Explore more

More ways hiring teams screen with InterviewAgent.ai

Stop running first-round calls by hand. Put screening on autopilot.

Set your role and rubric and the agent interviews every applicant, scores each answer, and ranks a shortlist for your team. The agent advances candidates, your recruiters make every hiring decision.

See pricing

Role-tailored questions · bias-audited to EEOC and LL144 · human-in-the-loop