We leave the meeting with a fairly firm conclusion, although we may not be entirely sure what it rests on. The answers matter, of course, but so do fluency, presence, perceived similarity, the chemistry of the conversation, and our memories of other people who once made the same impression.
This is how confident statements begin to take shape: “They have potential.” “I can’t see them leading people.” “They seemed superficial.” “This is exactly who we need.” Sometimes we are right. On other occasions, it is only after the person has been hired or promoted that we discover how much of the missing information we filled in ourselves.
In 2022, Paul Sackett, Charlene Zhang, Christopher Berry, and Filip Lievens published an extensive reassessment of personnel selection research in the Journal of Applied Psychology. Their study challenged several conclusions that had long been treated as settled knowledge and recalculated the predictive validity of methods commonly used to estimate future job performance. Structured interviews reached an average validity of approximately .42, while unstructured interviews stood at .19. This does not mean “42% accuracy”; the coefficient describes the statistical relationship between assessment results and performance observed later. The difference is still difficult to dismiss. A conversation becomes far more useful when the evaluator knows beforehand what they are looking for and assesses answers against comparable criteria. In fact, Sackett and colleagues’ study placed structured interviews at the top of the methods examined in terms of mean validity.
This does not require a cold conversation read from a script. Structure begins before the candidate enters the room: What are we trying to find out? How is it connected to the role? What kind of evidence would give us reasonable grounds to believe that this person can meet that particular demand?
Consider two candidates competing for their first managerial position. The first speaks exceptionally well about leadership. They respond quickly, use the organization’s language with ease, and offer a convincing account of how they would combine empathy with accountability when dealing with an underperforming colleague. The second candidate takes longer to think. They do not have the same presence or such a well-rehearsed vocabulary. In a free-flowing interview, the first candidate will probably leave the stronger impression.
Now give both candidates the same situation: a 15-minute conversation with a team member who repeatedly misses deadlines and becomes defensive when challenged. The comparison suddenly moves into more relevant territory. We are no longer listening only to how they talk about management. We can observe what they do when they have to practice it.
The highly persuasive candidate may show considerable understanding and protect the relationship, yet end the conversation without making clear what needs to change. The other may listen less impressively, but distinguish between genuine obstacles and recurring explanations, state the expectation without becoming aggressive, and close with a concrete agreement. We have not uncovered either candidate’s “true personality.” We have gathered a limited piece of evidence from a constructed situation. Still, it is evidence about behaviors that matter in the role, which makes it more valuable than the impression that one of them simply “looks like a manager.”
This is where competency-based decision making begins. Competencies become useful only when they leave the presentation slide and enter the reality of the job. “Strategic thinking,” “collaboration,” “customer orientation,” and “accountability” all sound appropriate, but they can mean almost anything. Two assessors may agree that they are looking for strategic thinking while one rewards the ability to speak convincingly about the future and the other looks for evidence that the candidate understands how a decision will affect other functions.
In one role, strategic thinking may involve recognizing that a solution which benefits one department merely transfers costs and problems to two others. Somewhere else, it may mean noticing that the incident everyone is trying to fix is only the visible expression of an older pattern. Sometimes it appears in the decision to decline an attractive short-term opportunity because it would consume resources needed for a more important priority. A widely cited paper on competency modeling by Michael Campion and his colleagues emphasizes precisely this connection with the work itself, the organization’s strategy, and the practical uses of the model. Without that connection, we are left with an elegant collection of words. Campion et al., 2011
When behaviors are defined concretely enough, the conversation between assessors changes too. “They seemed strategic” can become: “They identified the effect of the decision on two other departments without being prompted.” And “They don’t think long term” can be expressed more honestly as: “They proposed several options but did not compare their consequences beyond the next three months.” Judgment does not disappear. It simply becomes easier to examine and, where necessary, challenge.
The evaluator’s experience still matters. It would be absurd to pretend that an instrument can replace years spent watching people succeed, fail, learn, or return to the same mistakes. Structure does, however, ask experience to slow down for a moment and show what it is relying on.
“They interrupted the other person four times” is an observation. “They don’t know how to listen” is already an interpretation. “They lack emotional intelligence” takes us even further away from what actually happened in the room. The mind crosses the distance between these statements so quickly that we rarely notice the leap. A behavioral methodology makes it visible: What did we observe? What do we think it means? What other evidence supports that conclusion? The classic research by Campion, Palmer, and Campion identified 15 components through which interviews can be structured, ranging from job-analysis-based questions asked consistently across candidates to response scoring and behaviorally anchored rating scales. Campion, Palmer & Campion, 1997
The same problem appears after someone has joined the organization. “I’ve known him for seven years” contains a great deal of information, but it does not automatically answer whether he can lead a more complex team. Neither does “She is our best specialist.” An exceptional expert may become an excellent manager, or may continue taking over every difficult problem personally because that is exactly the behavior that earned recognition in the first place. A manager who is deeply appreciated by the team may create trust and psychological safety. That same person may also delay the conversations most likely to disturb those relationships. Both can be true at once.
There is an opposite temptation as well: responding to uncertainty with enormous assessment processes. Tests, questionnaires, interviews, simulations, scores, and reports accumulate until the sheer quantity of information creates the impression of rigor. Yet a crowded process is not necessarily a well-designed one. It can generate a great deal of data about secondary questions and very little about the behaviors on which the decision actually depends.
Even the authors of the 2022 reassessment warn against turning rankings of average validity into a universal recipe. Results vary across contexts, and the choice of method also needs to account for the role, cost, scale, candidate reactions, and potential implications for diversity. The applied analysis published by Sackett and colleagues in 2023 preserves precisely this nuance.
Before choosing the instruments, then, it is worth making the decision question as concrete as possible: What would we need to see this person doing to feel more confident that they can succeed in the next role? If the answer remains “demonstrate leadership” or “show strategic thinking,” the criterion is not yet helping us very much. If we can describe the situation, the behavior we expect, and the conditions under which it would be appropriate, we already have the foundation of an assessment that does not depend solely on who impressed us most during a conversation.
Sources
Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068. https://doi.org/10.1037/apl0000994
Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2023). Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors. Industrial and Organizational Psychology, 16(3), 283–300. https://doi.org/10.1017/iop.2023.24
Campion, M. A., Palmer, D. K., & Campion, J. E. (1997). A review of structure in the selection interview. Personnel Psychology, 50(3), 655–702. https://doi.org/10.1111/j.1744-6570.1997.tb00709.x
Campion, M. A., Fink, A. A., Ruggeberg, B. J., Carr, L., Phillips, G. M., & Odman, R. B. (2011). Doing competencies well: Best practices in competency modeling. Personnel Psychology, 64(1), 225–262. https://doi.org/10.1111/j.1744-6570.2010.01207.x
