Designing AI Interview Questions Tied to Specific Competencies
Competency-tied interviews predict job performance twice as well as gut-feel hiring.

A behavioral interview built around real competencies predicts job performance at.51 validity. An unstructured interview, the kind most hiring teams still run on gut feel and rapport, is.28 (Schmidt, Oh, & Shaffer). That's not a small gap; it's the difference between a hiring process that works and one that mostly doesn't. AI interview tools don't close that gap by themselves. They just apply whatever definition of "good" you feed them, at scale and without fatigue, which means a bad question design gets worse, not better, once AI is running it.
What a competency is and why most job descriptions don't contain one
A competency is a specific, observable behavior pattern that predicts performance in a role. Most job descriptions don't contain a definition like that.
"Strong communicator" "Strong communicator" is a trait, and traits are self-reported, not observed. It's a trait, and traits are self-reported, not observed. "Five years of experience" isn't a competency either, it's a credential, and credentials tell you what someone was exposed to, not what they did with it. "Manages the pipeline" is a duty. It describes a slot in an org chart, not a behavior that predicts the person filling that slot will actually succeed.
Job descriptions list responsibilities and preferred qualifications because those are easy to write and easy to agree on in a meeting. Neither one describes the behavioral evidence that separates someone who thrives in the role from someone who quietly struggles through it for eighteen months before leaving.
Credential shortcuts appear here. "Ex-FAANG" or "MBA preferred" on a job posting isn't a competency substitute, it's a pattern-recognition shortcut standing in for one. It tells the hiring team what kind of resume to feel good about, not what kind of behavior to expect on the job. And if an AI tool gets trained on job descriptions built that way, it inherits the shortcut. It will keep optimizing for the pattern it was shown rather than the performance the team actually needs.
A real competency needs three components, and skipping any one of them is how teams end up with an interview that feels rigorous but isn't:
A behavioral anchor is a plain description of what this looks like when someone does it well. A level calibration, meaning what looks like at this specific stage, team size, and context, because at a 12-person startup and at a much larger company are not the same behavior. And a failure mode is a description of what "present but weak" looks like, so an evaluator, human or AI, can actually tell the difference between someone who has the competency and someone who's just talking around it convincingly.
Calibrating what "exceptional" looks like before writing a single question
Calibration is a deliverable. It's a deliverable. The team needs a written, shared definition of what a strong answer looks like before the first candidate walks in, not after three candidates have already been scored on three different internal rubrics that never got written down.
Skip that step and a few predictable failures occur. Contrast bias occurs when a mediocre candidate looks great only because they interviewed right after someone worse. Panelists carry different private definitions of "great," so one interviewer's 4 is another's 2, and nobody notices until debrief turns into an argument about vibes. And seniority in the room starts overriding evidence, where the most senior voice in the debrief effectively decides the outcome regardless of what the actual transcript shows.
One practical approach is replacing subjective interview notes with structured scoring sheets. Interviewers score each candidate on a 1 to 5 scale per competency, and surfacing meaningful score gaps between panelists before the debrief rather than during it keeps the conversation grounded in evidence. That's a small mechanical change, but it forces disagreement into the open instead of letting it resolve itself by whoever talks first.
What calibration produces that a job description simply cannot: a written definition of the competency at each score level, agreement on what kind of story counts as evidence (which decision context, which outcome type), and a failure-mode description specific enough to stop a polished-but-thin answer from sailing through on charisma alone.
Translating a competency into a question that produces evaluable behavioral signal
A common answer framework, situation, a specific element, action, result, is a floor, not a ceiling. It gives an answer structure, but structure alone doesn't guarantee the behavioral signal a specific competency actually requires. The real work is in constraining the situation type the question asks about, not just asking for a story shaped like STAR.
Four design levers do that work. Situation constraint means specifying the exact context that tests the competency. If the competency is "initiates under ambiguity," ask about a moment when the path forward was genuinely unclear, not a moment of difficulty in general terms. Evidence focus means asking what the person did, not what they thought or felt or would hypothetically do. Feelings and hypotheticals don't produce behavioral signal, actions do. Outcome specificity means building in a follow-up that asks what actually changed as a result, which separates candidates who took real action from candidates who narrate busy-sounding activity with no consequence attached. And a failure probe, a question about what didn't work or what they'd do differently, is where polished responders tend to collapse while genuine performers lean in.
Take "initiates under ambiguity." A weak version of this question sounds like: "Tell me about a time you showed initiative." That's open enough that any story fits it, which means it filters nothing. A competency-tied version instead asks: "Tell me about a situation where you had to move a project forward but the goal or the process wasn't clearly defined. What specifically did you do to get started, and how did you decide that was the right move?" Follow it with: "What happened, and what would you do differently now?" Same competency, but now the question constrains the situation and forces action narration instead of general reflection.
AI fluency is worth a second worked example because it's becoming a baseline expectation fast. A St. Paul survey of 500 hiring managers found AI fluency ranked as the single top hiring priority for 2026, with more than one in five employers treating AI proficiency as a baseline expectation, not a bonus. That creates a real problem for non-technical roles where there's no portfolio to check.
"Are you comfortable with AI tools?" answers nothing, because everyone says yes. A competency-tied version instead asks: "Walk me through a specific workflow where you used an AI tool. What did you use it for, how did you verify its output, and where did you decide a human judgment call was necessary?" That structure forces specificity, and specificity separates a rehearsed general answer from a real one, since specificity separates a rehearsed general answer from a real one in ways evaluators can act on. Add the failure probe: "Tell me about a time the AI output was wrong or misleading. What did you do?" Anyone who's actually used these tools day to day has an answer to that. Anyone who hasn't, doesn't.
Building the scoring rubric that tells the AI what signal to weight
A rubric is not a checklist. A checklist rewards completeness: it gives credit for touching every part of STAR whether or not any of it demonstrates the competency. A rubric weights the specific behavioral evidence that actually predicts performance, which is a different and harder thing to build.
A working rubric for a competency-tied interview looks something like this. Absent scores a 1: the candidate describes the situation but narrates no specific personal action, or the action described was entirely prompted by someone else. Score 2, weak: an action is described but it's generic, with no evidence of the specific behavioral pattern the competency requires, and the outcome is vague or unattributed. Score 3, present: a specific action with a clear decision rationale, a described outcome, and at least one dimension of the competency clearly shown. Score 4, strong: a full behavioral story with context, decision logic, consequence, and reflection, with multiple dimensions of the competency visible. Score 5, exceptional: everything in a 4, plus evidence the candidate can name what they learned and apply it to a different context entirely, which is what separates a one-time story from a transferable standard.
What does an AI system actually do with a rubric like this? It evaluates the response against those defined levels and surfaces the specific moment in the transcript that matches each score, an evidence-backed match rather than a black-box number. Platforms like Willo describe their evaluation approach in similar terms, surfacing how a candidate's response maps against a defined rubric rather than just asserting a rating.
That only works if the rubric itself is right, though, and the rubric has to be human-authored and human-approved. The AI applies the standard. It doesn't get to set it. Whatever score levels get written into that rubric are the team's own standards for what "good" means in this specific role, not a vendor's generic default template.
Where AI-assisted evaluation helps and where human judgment is irreplaceable
AI does a specific set of things well in a structured, competency-tied interview process. It applies the same rubric to every candidate without fatigue, without the fifth interview of the day getting graded more harshly than the first. It flags score discrepancies between panelists that would otherwise just get smoothed over by whoever speaks first in debrief. It produces evidence-backed summaries that hand the hiring manager a specific reason for a score instead of a bare number with no context. And it can run the mechanical parts of the loop, question delivery, response capture, rubric application, so a human's attention goes toward candidates who've already cleared a meaningful bar.
What it can't do is decide whether the competency definition itself is the right one for this team, at this moment, in this market. That call requires judgment and context that lives with the hiring manager, not the platform. It also can't reliably catch what a candidate doesn't say: the evasive answer, examples that all come from low-stakes situations, or a response that technically scores well against the rubric but rings hollow against everything else known about the role. And the final call on an offer has to stay with a person, both as a matter of principle and as a practical check against any blind spot baked into the rubric itself.
Recruiters overwhelmingly plan to expand AI use in 2026, roughly 93% by one count, and 59% already say AI has surfaced candidates they wouldn't have found otherwise. Recruiters overwhelmingly plan to expand AI use in 2026, roughly 93% by one count, and 59% already say AI has surfaced candidates they wouldn't have found otherwise. That's a real result. But finding a candidate and evaluating them correctly are two separate problems, and the second one is entirely dependent on the design of the interview and rubric behind it.
Which points to the actual risk in this cycle. AI sourcing tools are good at widening the pool, at surfacing the non-obvious candidate who doesn't have the familiar résumé shape, by identifying candidates whose profiles diverge from that shape. But if the interview rubric behind that expanded pool is still built around credential signals rather than behavioral evidence, that same non-obvious candidate gets screened right back out at the evaluation stage. Wider sourcing paired with a shallow rubric just moves the bias one step downstream and makes it harder to see. It just moves the bias one step downstream and makes it harder to see.
A practical design sequence for building competency-tied AI interview questions from scratch
Start with the competencies, not the questions. Identify three to five that actually predict success in the role, not more. The honest way to find them is to ask the hiring manager directly: what does the person who thrives here do differently from the person who struggles? That answer rarely matches the job description, and that gap is the whole point. Cross-reference it against the top performers already on the team. What behavioral patterns do they actually share that never appear in their titles or their résumés? Keep the list to five competencies or fewer. Push past that and every interview ends up covering each one shallowly instead of testing any of them properly.
Once the competencies are set, write the behavioral anchor for each one before drafting a single interview question. The anchor is a plain-language description of what this looks like in action, specific enough that two different interviewers, reading the same candidate story, would agree on whether it counts as evidence. If two reasonable people can't agree the anchor demonstrates the competency, the anchor isn't specific enough yet, and no amount of clever question-writing downstream will fix that. The question design comes after this groundwork is done, never before it.


