Behaviorally Anchored Rating Scales: What They Are and Why They Work
BARS replace subjective impressions with observable behavior. Here is how they are built and why they improve scoring consistency across interviewers.
Most interview scoring is impressionistic. An interviewer finishes a conversation, forms a general feeling about the candidate, and then assigns a number. The number reflects the feeling, not the evidence.
Behaviorally anchored rating scales (BARS) are designed to fix this. They replace impressions with observable behavioral descriptions, one at each rating level. The interviewer is not asked to judge the candidate. They are asked to match what the candidate said to a described behavior.
What a BARS looks like in practice
A BARS for the competency “Communicates technical information to non-technical audiences” might look like this:
5 - Exceeds expectations: Candidate described translating complex technical constraints into business terms without prompting, checking for understanding, and adjusting the explanation when the audience indicated confusion. Proactively identified what the audience needed to know vs. what was technically interesting.
3 - Meets expectations: Candidate described simplifying technical content for a non-technical audience, using analogies or examples. Response did not include evidence of checking comprehension or anticipating gaps in audience knowledge.
1 - Below expectations: Candidate described explaining technical information but used technical jargon throughout, or did not tailor the explanation to the audience. No evidence of adapting based on feedback.
The interviewer does not decide whether the candidate was “good” at communication. They find the anchor that most closely matches what the candidate actually said.
Why BARS improve interrater reliability
When two interviewers score the same candidate differently on the same competency, that variability is measurement error. It reduces validity and introduces the kind of inconsistency that creates legal exposure.
BARS reduce that variability because the scoring criterion is behavioral, not evaluative. Two interviewers may disagree about whether a candidate is “an excellent communicator,” but they are more likely to agree on whether the candidate’s response matches the level-5 anchor description or the level-3 description.
Research on structured interviews consistently shows that adding behavioral anchors to rating scales improves interrater reliability. Not to perfect agreement, but enough to meaningfully reduce the noise in the data.
How BARS are developed
Building a BARS from scratch requires involvement from subject matter experts who know the job.
The standard process:
- Identify the competency from job analysis. Know what you are measuring before you build the scale.
- Collect critical incidents from supervisors and incumbents: examples of effective and ineffective performance for this competency in this role.
- Draft behavioral anchors for each rating level based on those incidents.
- Have SMEs retranslate the anchors back to performance levels independently. If SMEs consistently assign a given anchor to the same level, it is a good anchor. If they disagree, revise it.
- Pilot the scale with a small group of interviewers before deployment.
The SME involvement is not optional. Anchors that are developed without people who do the job tend to describe idealized or generic behavior rather than what effective performance in this specific role actually looks like.
The limits of BARS
BARS are a scoring tool, not an interview design tool. A well-built BARS attached to a poor question produces limited improvement. The question has to elicit the behavior you are trying to score.
BARS also do not eliminate subjectivity in evidence gathering. They standardize the scoring of evidence once it has been gathered, but what the interviewer records from a conversation still involves judgment. This is why standardized note-taking guidance, alongside the BARS, further improves consistency.
Talent Systems AI
See structured, scored interviewing running on your roles.
Validated competency rubrics. Adverse impact monitoring. Full audit trail from day one.
Book a Demo