AI screening tools are only as good as the criteria they're scoring against. A system that evaluates candidates against vague requirements doesn't produce better shortlists than a vague human review. It produces vague shortlists faster. Here's what makes a requirement evaluable by an AI system.
The evaluability test
Before you finalize any requirement, apply the evaluability test: could a careful reader of a resume find concrete evidence for or against this criterion? If the criterion requires inference that depends on the reader's personal experience or industry context, it's not evaluable. If it requires specific evidence that a resume either contains or doesn't, it is.
Not evaluable: "Strong communication skills." Every candidate who didn't fail out of school has "strong communication skills" somewhere on their resume. An AI system will find the phrase or a synonym and score it positively, which is meaningless.
Evaluable: "Has produced written technical documentation intended for a non-technical audience, such as product specs, executive summaries, or customer-facing technical guides." This is checkable from project descriptions and work samples. It's specific enough that absence of evidence is meaningful.
The evaluability test also catches requirements that are evaluable in theory but too vague to produce consistent scores in practice. "Has led a team" is evaluable -- it's checkable from job titles and role descriptions. But it's broad enough that a consistent score is hard to define: does "led a team" mean formal management of direct reports, cross-functional coordination without direct authority, or something in between? Adding specificity -- "has directly managed 3 or more engineers for at least one full product cycle" -- produces a criterion that an AI system can apply consistently across a large candidate pool.
The output anchor
The most reliable scoring criterion pattern is the output anchor: describe what the candidate should have made, built, or shipped, rather than describing abstract skills. "Has built and deployed a production ML model that handles real-time inference" is evaluable. "Experience with machine learning" is not, because experience can mean anything from a classroom project to leading a production ML infrastructure.
Output anchors work because resumes are largely collections of outputs. People list projects, products they shipped, features they owned. An AI system can match claimed outputs against the output you specified.
The output anchor approach also surfaces an important distinction between outputs and tools. "Experience with Kubernetes" is a tool requirement. "Has designed and operated a containerized production system handling real traffic" is an output requirement that encompasses Kubernetes among several valid approaches. The output-anchored criterion is both more accurate (it captures candidates using equivalent tools) and more relevant (it tests whether the candidate can do the thing, not whether they've used one specific implementation of it).
Context clues that improve scoring accuracy
Specify the context in which you'd expect to find the evidence. "Has managed a recruiting process from sourcing to offer, independently, without a coordinator" is more specific than "recruiting experience." The context clue (independently, without a coordinator) significantly narrows the population that meets the criterion and makes it harder to score a false positive.
Similarly: "At a company under 100 employees" or "in a regulated industry" as a contextual modifier changes the population substantially and makes scoring more accurate for your actual needs.
Context clues are particularly valuable for roles where domain specificity matters. "Has built customer-facing software" is a broad criterion. "Has built customer-facing software for financial services firms, working within compliance review requirements" is specific enough that a candidate who has done this will have distinctive language in their resume that clearly differentiates them from a candidate who has built generic consumer software. The specificity does the discrimination work that vague criteria can't.
Structuring the full rubric
A complete rubric for AI scoring typically has 5 to 8 criteria: 2 to 3 that are genuinely disqualifying (must-haves), 3 to 4 that are important but not disqualifying (strong signals), and optionally 1 to 2 that are differentiating factors among otherwise equivalent candidates. Each criterion needs a weight that reflects its relative importance, not a weight that makes the rubric look balanced.
The must-have criteria should be specific and observable. The strong-signal criteria can be slightly broader but still output-anchored. The differentiating criteria are often the most interesting to specify because they reflect what would make a genuinely exceptional hire distinct from a good one.
What to do with soft criteria
Some genuinely important criteria aren't evaluable from a resume. "Collaborative working style" and "intellectual curiosity" matter for some roles but can't be scored from resume evidence. The answer isn't to remove them; it's to evaluate them at a different stage (phone screen, structured interview) and not ask the resume screening to do work it can't do. A screening rubric that mixes evaluable and unevaluable criteria produces incoherent scores.
The practical result of keeping soft criteria in the screening rubric is that AI systems either ignore them (finding no evidence, they score zero or average) or score them by proxy (finding phrases like "team player" or "collaborative environment" that candidates include specifically to satisfy expected criteria). Neither is useful. Soft criteria belong in your interview rubric, with structured interview questions designed to surface evidence for them, not in your resume screening rubric where they can't be evaluated reliably.


