Judgment, disciplined: the research behind the MyCandidate scoring system
Most hiring mistakes are made on the second look, not the first. The research on noise, bias, and structured judgment, and the four design choices behind the MyCandidate scoring system.
Most hiring mistakes are not made at the top of the funnel. They are made on the second look.
By the time a search reaches a shortlist, the sourcing problem is largely solved. What remains is an act of judgment: does this candidate meet the standard the role requires, and can that conclusion withstand a client's scrutiny? It is the least examined step in most search processes, and it fails in two analytically distinct ways. The first is bias, a systematic tilt for or against particular candidates. The second, and the more neglected, is noise: ordinary inconsistency, the same evidence read two ways on two days by two readers. Bias has a direction. Noise has none. Both degrade the quality of the hire, and a scoring system worth the name must be designed against both.

Fig 1. Noise: the same input, two outputs.
How much judgment drifts
The scale of the problem, even among trained professionals, is difficult to overstate. Early in Noise: A Flaw in Human Judgment, Daniel Kahneman, Olivier Sibony, and Cass Sunstein describe a noise audit at an insurance company in which underwriters independently priced the same policies. Asked what they expected, the firm's executives put the difference between any two underwriters at roughly 10 percent. The measured median difference was 55 percent: five times what leadership assumed, and wholly invisible until someone thought to measure it (Kahneman, Sibony, & Sunstein, 2021).
There is no reason to expect hiring to be quieter than underwriting. Two experienced partners read the same candidate and reach different verdicts, each internally coherent, neither aware of the gap between them. This is not a failure of competence, and it cannot be trained away. It is structural, and the structural remedy is to change what the judgment is anchored to. That, in a sentence, is what a scoring system is for: not to displace human judgment, but to give it a fixed reference so that it stops drifting.
The instinct to trust the personal read is durable, and the evidence against it is uncomfortable. In one well-known demonstration, interviewers formed impressions just as confidently when the people they interviewed answered questions at random as when they answered honestly, and never noticed anything was amiss; even observers who were told that half the interviews they had watched were random still judged most of the random ones to be genuine (Dana, Dawes, & Peterson, 2013). The interview feels informative. The feeling is not evidence that it is.
Four design choices, and the research behind them
A defined, versioned standard. Every search in the system begins with what we call a Role Model: a written, versioned account of what the role actually requires, drawn from the brief and the firm's own criteria. Every candidate is then assessed against that same standard, rather than against the shifting benchmark of whichever candidate was read last.
In personnel psychology this is the distinction between structured and unstructured assessment, and it is among the most replicated findings in the discipline. Schmidt and Hunter's synthesis of eighty-five years of research placed structured methods well above the unstructured interview for predicting job performance (Schmidt & Hunter, 1998), drawing its interview estimates from a meta-analysis of more than 86,000 individuals (McDaniel, Whetzel, Schmidt, & Maurer, 1994). The specific validity figures from that era have since been revised, and the revision deserves to be stated rather than discovered. Sackett and colleagues (2022) demonstrated that decades of meta-analyses had systematically over-corrected for range restriction, inflating the field's most-cited numbers. What is telling is what the revision left standing: once re-estimated, the structured interview emerged as the strongest single predictor of job performance in their analysis. The magnitude was in dispute. The direction never was.
Structure also narrows the aperture. By holding every candidate to the same job-relevant criteria, it reduces the weight of the things that should carry none. In the meta-analytic evidence, racial subgroup differences in interview ratings run smaller under high structure than under low, d = .23 against .32, a gap roughly a quarter narrower (Huffcutt & Roth, 1998). It is worth being exact here, because the temptation to overstate is real: structure reduces bias, it does not eliminate it. Anyone who promises otherwise is selling something. What structure does is shrink the gap and make whatever remains visible enough to inspect.
Anchored scales. Within the Role Model, each requirement is rated on a defined scale, from absent to exceptional, in which every level carries a concrete, written meaning. This is a deliberately old idea. Behaviorally anchored rating scales, introduced by Smith and Kendall (1963), attach a specific behavioral description to each point so that "strong" means the same thing to two different assessors. A sixty-year-old instrument, built for precisely the purpose it serves here: converting an impression into a rating on which two assessors can agree.

Fig 2. The anchored scale: each requirement rated against a written definition.
Consistent combination. Once the evidence is gathered, it must be combined into a single judgment. The intuitive method, allowing an experienced reader to weigh everything holistically, is, on nearly a century of evidence, the wrong one. This is the clinical-versus-actuarial question, first framed by Paul Meehl (1954) and resolved the same way ever since. Grove and colleagues' meta-analysis of 136 studies found mechanical prediction about ten percent more accurate on average, and, more revealing, found human judgment clearly superior in only six to sixteen percent of the studies examined (Grove, Zald, Lebow, Snitz, & Nelson, 2000). Dawes, Faust, and Meehl (1989), writing in Science, had reached the same verdict across domains a decade earlier.

Fig 3. The formula matches or beats the expert in the large majority of studies.
The obvious objection is that hiring is different from the medical and forensic settings much of that work covers. The evidence does not grant the exemption. Kuncel and colleagues (2013) studied selection and admissions decisions specifically and found that combining the same candidate data mechanically rather than holistically improved the prediction of job performance by more than fifty percent, an advantage that held even when the judges were experts familiar with the jobs and organizations in question.
None of this is an argument for automation, and it is worth saying so plainly. The finding is narrow: given the same evidence, a consistent rule combines it more accurately than a shifting intuition. The rule does not decide who to hire. It produces a defensible, repeatable ranking that a person then interrogates. In the system, must-have requirements act as gates and known derailers are surfaced early, but the recruiter still makes the call. The point is only that the call should start from the same place every time.
Visible reasoning. A score is worth only as much as it can be audited. Every number the system produces traces back to the specific language in a resume or interview that supports it. Part of the reason is defensibility: when a client asks "why her," the answer should be evidence, not conviction. The subtler reason is psychological. People abandon an algorithm faster than they abandon a person, reverting to their own judgment after seeing it err, even when the tool remains measurably better, a pattern the literature calls algorithm aversion (Dietvorst, Simmons, & Massey, 2015). The antidote is not a more polished black box. It is no black box at all. A score you can open, question, and overrule is one you will actually come to trust. A score you cannot interrogate is just an opinion with a decimal point.

Fig 4. In the product: candidates scored and ranked against the Role Model, each conclusion traceable to its evidence.
What a scoring system is, and what it is not
Understood this way, a scoring system is not a machine that replaces the recruiter's judgment. It is an instrument that disciplines it: the same standard, the same scale, the same method of combination, applied to every candidate, with the evidence attached and the reasoning exposed. It does not source candidates, conduct interviews, or make the hire. It makes the one judgment a person still owns more consistent, less biased, and easier to defend, on the first search and on the hundredth.
That is the discipline the system exists to keep. Run every search to your standard.
References
Dana, J., Dawes, R., & Peterson, N. (2013). Belief in the unstructured interview: The persistence of an illusion. Judgment and Decision Making, 8(5), 512–520.
Dawes, R. M., Faust, D., & Meehl, P. E. (1989). Clinical versus actuarial judgment. Science, 243(4899), 1668–1674.
Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114–126.
Grove, W. M., Zald, D. H., Lebow, B. S., Snitz, B. E., & Nelson, C. (2000). Clinical versus mechanical prediction: A meta-analysis. Psychological Assessment, 12(1), 19–30.
Huffcutt, A. I., & Roth, P. L. (1998). Racial group differences in employment interview evaluations. Journal of Applied Psychology, 83(2), 179–189.
Kahneman, D., Sibony, O., & Sunstein, C. R. (2021). Noise: A flaw in human judgment. Little, Brown Spark.
Kuncel, N. R., Klieger, D. M., Connelly, B. S., & Ones, D. S. (2013). Mechanical versus clinical data combination in selection and admissions decisions: A meta-analysis. Journal of Applied Psychology, 98(6), 1060–1072.
McDaniel, M. A., Whetzel, D. L., Schmidt, F. L., & Maurer, S. D. (1994). The validity of employment interviews: A comprehensive review and meta-analysis. Journal of Applied Psychology, 79(4), 599–616.
Meehl, P. E. (1954). Clinical versus statistical prediction: A theoretical analysis and a review of the evidence. University of Minnesota Press.
Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274.
Smith, P. C., & Kendall, L. M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology, 47(2), 149–155.