Psychometric Testing in Hiring: What the Science Actually Says, What Companies Get Wrong, and What to Do Instead
A deep look at one of talent acquisition’s most used and least understood tools, with cited research, documented program examples from the UK and US, and a practical framework for getting it right.
There is a phrase I hear in hiring conversations that should make anyone who takes this work seriously stop and ask a follow-up question.
It goes: “We use psychometric testing, so our process is objective.”
I have been in recruiting for over 12 years, working across early-stage companies, Y-Combinator backed start-ups, and institutional employers on both sides of the Atlantic. And I can tell you with complete confidence that “we use psychometric testing” tells you almost nothing about whether a hiring process is fair, predictive, or defensible. What matters is which tools, built on what construct, validated against what evidence, deployed in what configuration, monitored for what outcomes.
The gap between how psychometric tools work in peer-reviewed research and how they get deployed inside real hiring pipelines is one of the most consequential mismatches in modern talent acquisition. This article is my attempt to close that gap. Not with vendor-friendly summaries, and not with a blanket dismissal of psychometrics as pseudoscience either, because that is equally wrong. With the actual evidence, the actual documented programs, and the actual practitioner checklist that separates defensible systems from expensive theatre.
If you are a company using psychometric testing, this is what you need to audit. If you are a candidate trying to understand why you keep getting filtered out before a human has reviewed a single line of your application, this is the machinery.
First, get the definitions right
Psychometric testing is not one thing. It is a family of standardized assessment methods used to infer job-relevant attributes, and the family is large enough that conflating its members is where most of the confusion starts.
The AERA/APA/NCME Standards for Educational and Psychological Testing define validity around the intended interpretation and use of scores. That means before you can evaluate whether psychometric testing “works,” you have to specify which kind, for which construct, used in which way. The Standards are explicit that a valid test is not a globally valid test. Validity is specific to an intended use.
With that in mind, the major categories are:
Cognitive ability tests measure reasoning, memory, perceptual speed and accuracy, arithmetic, and reading comprehension. The EEOC defines them this way in its guidance on employment tests and selection procedures. These are what most people picture when they hear “psychometric assessment,” and they are also where the most heated validity debates have played out.
Personality inventories measure typical behavioural tendencies, usually organized around the Big Five framework: openness, conscientiousness, extraversion, agreeableness, and emotional stability. Barrick and Mount’s meta-analysis found that conscientiousness has a relatively consistent relationship with job performance across occupations. The key word is typical. Personality tests measure what someone usually does, not what they can do at their best.
Situational judgment tests present job-relevant scenarios and ask candidates to choose or rate responses. McDaniel and colleagues’ meta-analytic work found meaningful criterion-related validity, and subsequent research established that what an SJT actually measures depends heavily on whether it asks what a candidate would do or should do. Knowledge-based SJTs correlate more strongly with cognitive ability. Behavioural tendency SJTs correlate more with personality dimensions like agreeableness and conscientiousness. That distinction matters enormously for how you use them and what you claim they are measuring.
Integrity tests are more powerful than their reputation suggests. Ones, Viswesvaran, and Schmidt’s meta-analysis found an estimated mean operational validity of .41 for predicting supervisory ratings of job performance. They are not just theft screens. They predict a broad range of counterproductive workplace behaviours and, used appropriately in the right roles, have a stronger evidence base than many more fashionable tools. The EEOC treats them as personality-like measures aimed at assessing dependability and the likelihood of counterproductive conduct.
Work samples require candidates to perform or simulate job tasks directly: a writing exercise, a coding challenge, a customer call simulation. They offer the strongest face validity because the connection to job performance is transparent. Their main limitation is scalability and feasibility, and I will come back to a finding about their real-world validity that most practitioners have not heard.
Assessment centres combine multiple exercises with multiple trained assessors. OPM describes them as multi-method systems particularly effective for higher-level managerial and leadership competencies. They are resource-intensive and logistically demanding, but for the roles where they are used properly, they are more informative than any single psychometric instrument.
Each of these operates on different logic, has different evidence, carries different legal risks, and is appropriate in different contexts. Anyone claiming psychometric testing “works” without specifying which tool is making a category error.
What the evidence actually says
Here is where we need to spend serious time, because the validity literature in personnel selection is both more robust and more nuanced than most practitioners appreciate.
The canonical starting point is Schmidt and Hunter’s 1998 meta-analysis, which synthesized decades of validity research and found general cognitive ability to be among the strongest predictors of job performance and training success across occupations. That study became the foundation for much of the modern case for cognitive testing in selection. Structured interviews, work samples, integrity tests, and assessment centres all featured alongside cognitive ability in the rankings. OPM’s Assessment and Selection site still cites the Schmidt and Hunter work as part of its foundational framing.
For much of the 2000s and into the 2010s, this evidence base was treated as settled. Vendors built their sales materials around it. HR functions cited it in procurement conversations. The mean validity estimates from Schmidt and Hunter became reference points that were repeated without much scrutiny.
Then Sackett, Zhang, Berry, and Lievens published their re-analysis, and the picture shifted.
Their core finding: many of the historical validity estimates had been substantially overestimated, because earlier meta-analyses had applied range restriction corrections in ways that inflated effect sizes by roughly .10 to .20 across many methods. When they applied more appropriate corrections, most mean validity estimates dropped, and the rankings changed meaningfully. Structured interviews moved to the top of the ranking. The relative advantage of cognitive ability tests over other well-designed methods shrank.
Let me be precise about what this does and does not mean. It does not mean psychometric testing does not work. It means the effect sizes were oversold, the gap between the top-ranked methods and alternatives was narrower than commonly claimed, and the case for cognitive-heavy screening systems used with hard cut-offs is weaker than a decade of vendor presentations suggested. As Sackett and colleagues put it, the research produced validity-diversity trade-off analyses rather than a declaration that cognitive tests outperform everything else by a margin that makes the adverse impact worth absorbing.
For specific test families, the revised picture is instructive.
On work samples: Roth, Bobko, and McFarland found validity closer to .33 corrected for attenuation than the historically cited .54. Still meaningful. But the “work samples solve everything” argument, which was already too simple, is harder to sustain when the effect size drops by a third and when the same researchers noted larger Black-White differences in applicant work-sample scores than most practitioners expected. The assumption that work samples cleanly solve adverse impact concerns does not hold.
On integrity tests: Ones, Viswesvaran, and Schmidt’s .41 estimate for supervisory ratings has held up better in subsequent work and is worth taking seriously. These tools are underused in development contexts and over scrutinized as if they were lie detectors.
On personality: Barrick and Mount found conscientiousness to have the most consistent relationship with performance across roles, but the overall validity of broad personality measures for most job performance criteria is moderate at best. Hogan, Barrett, and Hogan’s research with real applicants argued that response distortion may be less damaging in practice than laboratory faking studies imply, but that does not resolve the deeper question of whether personality self-report is a sufficiently strong predictor to carry significant weight in selection decisions.
The most important practical conclusion from the revised evidence base: the best-supported approach combines multiple methods, weights structured interviews more heavily than convention often suggests, and does not assume that any single psychometric instrument carries enough validity to act as a standalone gate.
The methodology that separates defensible systems from expensive theatre
The authoritative standards for psychometric testing in employment settings form a coherent framework across jurisdictions. The AERA/APA/NCME Standards, the SIOP Principles for the Validation and Use of Personnel Selection Procedures, EEOC’s Uniform Guidelines on Employee Selection Procedures, OPM’s assessment guidance, the BPS testing standards in the UK, and Ofqual’s reliability research all converge on the same logical sequence.
Start with job analysis.
OPM is direct on this point: job analysis identifies the duties and competencies needed for effective performance and makes the link between job requirements and assessment tools more transparent and fair. Before a single vendor demo, before a single contract signature, job analysis should be complete and current. It specifies what you are trying to predict. Without it, construct choice is arbitrary.
From job analysis, define the construct. What are you measuring? Inductive reasoning? Conscientiousness? Judgment under pressure? The Standards require that the intended interpretation of a score be clearly specified, and that validity evidence be collected in relation to that specific interpretation. “We used a validated psychometric” is not a construct definition. It is a category. The construct is what the test is actually measuring and why that construct matters for the job.
Then establish reliability. Ofqual’s research defines reliability as the consistency of outcomes that would be observed if assessment were repeated under equivalent conditions, and the Standards emphasize that evidence of reliability must be appropriate for each specific intended use. For pass/fail screening decisions, this means precision at the specific score region where classification is happening. A test with high overall internal consistency can still have large standard errors of measurement around a cut score, which means small score differences near that threshold translate to genuinely unreliable classification decisions. Most employers who use cut scores have never looked at the standard error of measurement around the threshold they are using. That is a significant oversight.
Then validate, locally. The Standards emphasize multiple sources of evidence: content, response processes, internal structure, relations to other variables, and consequences. Transportability of validity evidence from one context to another is never guaranteed. A tool validated on one role family in one country with one applicant pool is not automatically valid for your role, your country, your pool. The SIOP Principles make the same point. Local validation studies and ongoing monitoring are not optional rigour. They are the difference between a defensible system and a liability.
Standardization matters throughout. The Standards require administrators to follow standardized procedures specified by the developer, and warn that departures significant enough to affect score interpretation should be treated as evidence quality problems, not administrative inconveniences. If candidates are sitting your assessment in different conditions with different technical setups and different amounts of time, your scores mean different things.
Norms and cut scores are policy decisions. The Standards are explicit that cut scores are never purely technical. They embed value judgments about the consequences of false positives and false negatives, and they have direct fairness implications because they determine who advances. Documentation of how cut scores were set, what norm group was used, and what false-positive and false-negative rates result is not bureaucratic box-ticking. It is the evidentiary foundation for any disparate impact defence.
And then monitor, continuously. Differential item functioning analysis, or DIF, examines whether specific items behave differently across demographic groups who are similar on the target construct. The Standards describe DIF as part of fair item development. Adverse impact monitoring looks at outcome rates by protected characteristic across the full selection funnel. Neither is a one-time exercise. Roles change, labour markets shift, platforms update, applicant pools evolve. A validation study from five years ago on a different platform for a job that has since changed is not current evidence.
The public sector has this more figured out than most private employers
The most rigorously documented psychometric programs in both the UK and US are in the public sector, and that is not a coincidence. Legal obligations, procurement frameworks, transparency requirements, and scrutiny from oversight bodies create the governance conditions that good assessment systems need.
In the UK, the Civil Service Fast Stream uses online work-based scenarios, a situational judgment questionnaire, numerical tests, and work-based scenario assessments in early rounds, followed by a virtual half-day assessment center for candidates who progress. That assessment centre consists of a Written Advice Exercise, a Stakeholder Communication Exercise, and a Personal Development Conversation, with each exercise marked independently by different assessors. The Civil Service online tests guidance published on GOV.UK explains the test types, scoring methodology, and reasonable adjustment processes in transparent, candidate-facing detail that most private employers never produce.
The UK government’s Responsible AI in Recruitment guidance, published by the Department for Science, Innovation and Technology in March 2024, goes further. It requires equality impact assessments, algorithmic impact assessments, data protection impact assessments, and clear signposting to candidates when automation is involved in selection decisions. This is sector-leading governance documentation that sets a standard most private employers have not reached.
In the US, OPM’s USA Hire platform evaluates nearly 1 million applicants each year and is used by more than 80 federal agencies. The platform is described as providing valid, reliable, and fair skills-based assessments developed by OPM’s industrial-organizational psychologists. It delivers computer-adaptive tests, job simulations, and online interviews, available proctored or unproctored, on multiple devices. The OPM Assessment and Selection site provides a full public-facing explanation of the philosophy, methods, and evidence behind the approach, including a direct citation of Schmidt and Hunter’s foundational meta-analysis.
What these public sector examples share is not just rigor. They share documentation. Procurement evidence, algorithmic transparency records, technical basis for tool selection, and candidate-facing information about what is being assessed and why. These are the governance habits that private employers need to build.
In the private sector, the evidence base is thinner and more vendor-mediated. Talogy case studies document deployments at organizations like Boots with reported reductions in time-to-hire and improved candidate completion rates. HireVue’s case studies report scale screening deployments across retail with high candidate satisfaction scores. Walmart’s deployment of a Virtual Job Try-out covers scale hiring with vendor claims of more than 400 validation studies for the job try-out class, though not all are Walmart-specific and none have been independently evaluated. These are useful operational data points. They are not psychometric validation evidence.
The distinction matters enormously. When a vendor tells you their tool has been “validated,” the question is: validated where, on which population, against which criterion, using which correction methodology, and how recently? A vendor brochure that cites aggregate meta-analytic findings is not an answer to any of those questions.
The 4 risks most organizations have never properly assessed
Risk 1: Adverse impact is the employer’s problem, not the vendor’s.
The EEOC’s guidance on employment tests is explicit: when a selection procedure produces disparate impact on a protected group, the employer must justify it with job-relatedness and business necessity, and must consider whether there is a less discriminatory alternative that would work as well. The EEOC makes equally clear that algorithmic tools used by vendors are still selection procedures for which the employer carries legal responsibility. A vendor contract does not transfer liability.
The Uniform Guidelines on Employee Selection Procedures, adopted in 1978 and still governing employment test use under Title VII, provide the 4/5ths rule as a working threshold for detecting meaningful disparate impact: if the selection rate for any protected group is less than 4/5ths of the rate for the highest-selecting group, that is evidence of adverse impact requiring justification.
Cognitive-heavy screening used early in a funnel with hard cut-offs consistently produces the largest subgroup mean differences of any test family. The validity-diversity trade-off Sackett and colleagues analysed is not hypothetical. It is the actual data from real hiring systems showing that maximum predictive accuracy and minimum adverse impact rarely point to the same design choice. Organizations need to know which direction they have made that trade-off and whether the business case justifies it.
And the work sample assumption needs correcting here. Roth, Bobko, and McFarland found larger Black-White differences in applicant work-sample scores than most practitioners had assumed. The conventional wisdom that work samples solve the adverse impact problem by virtue of their job realism is not supported by the applicant-level data.
Risk 2: Accessibility is a legal requirement that cannot be retrofitted.
The UK Equality Act requires reasonable adjustments in recruitment and employment. The Americans with Disabilities Act requires accommodations for qualified applicants. The EEOC’s ADA guidance covers testing contexts specifically. The ICO’s guidance on automated decision-making in recruitment adds data protection and transparency obligations that interact with accommodation requirements.
Timing accommodations, alternative format provision, assistive technology compatibility, and accessible platform design are not features you add after launch. They are legal obligations that need to be built into the assessment architecture before the first candidate sits the test. The most common failure I see is organizations that procure an assessment platform, run it for a year, and then discover through a complaint or a legal review that the accommodation process was inaccessible, undocumented, or inconsistently applied. The Civil Service GOV.UK guidance sets an example here: reasonable adjustment processes are explained clearly to candidates before they sit any test, and the page is updated regularly.
Risk 3: Automation opacity is both a legal exposure and a candidate trust problem.
In the UK, the picture here has sharpened considerably. The ICO published a major report on automated decision-making in recruitment in March 2026, drawing on voluntary engagement with more than 30 employers. Its central finding: many employers characterize their use of automated recruitment tools as decision support, but in practice, their tools were making solely automated decisions with no meaningful human involvement. The ICO is unambiguous that a human reviewer who simply approves whatever the algorithm outputs, without the authority, time, or information to reach a different conclusion, is not providing meaningful human involvement.
The UK government’s Responsible AI in Recruitment guidance from DSIT (March 2024) requires equality impact assessments, algorithmic impact assessments, and data protection impact assessments for hiring systems that use automation. Candidates must be informed when automation is involved. Contestability, meaning the right to challenge an automated decision, must be built in.
In the US, the EEOC’s guidance on employment tests reinforces that “the algorithm decided” is not a defence against a Title VII or ADA claim. An employer who cannot explain how their automated screening tool scores applicants, what construct it claims to measure, and what adverse impact monitoring it has received is in a precarious legal position.
Beyond the legal dimension, candidate experience research is clear on the trust question. Hausknecht, Day, and Thomas found that applicants who hold positive perceptions about selection are more likely to view the organization favourably, more likely to accept offers, and more likely to recommend the employer to others. Opaque automated screening consistently drives negative perceptions. In competitive labour markets where employers are selling as much as they are selecting, that is a strategic problem as much as a legal one.
Risk 4: Vendor evidence is not local validation.
The most common procurement error in psychometric testing is treating a vendor’s published validity evidence as proof that the tool will work for your role, in your context, with your applicant pool. The AERA/APA/NCME Standards are unambiguous on this: validity evidence must support the specific intended use. The SIOP Principles make the same requirement explicit.
A meta-analysis with a mean validity of .35 based on studies across diverse industries and roles tells you almost nothing about whether a specific tool will predict performance in your specific role at the level needed to justify its use. Local validation requires your own data, your own criterion measures, and your own adverse impact analysis. Most employers never do it. Most vendor contracts do not require it. That gap is where the legal and predictive failures accumulate.
What to use instead, or alongside
The most important reframing from the revised validity literature: the alternative to psychometric testing is not intuition. It is structure. And the most defensible structure starts with something most organizations are not investing in nearly enough.
Structured interviews are at the top of Sackett and colleagues’ revised validity ranking. McDaniel and colleagues found that interview validity depends heavily on content and structure, with structured interviews substantially outperforming unstructured ones. For most organizations, the single highest-return investment in selection quality is not a new assessment platform. It is converting ad hoc interview conversations into structured, job-anchored processes with consistent questions, behavioural anchoring, scoring rubrics, and interviewer calibration. This does not require vendor contracts or new technology. It requires discipline and training.
Work samples and job trials remain valuable for roles where the core tasks can be demonstrated in a realistic time window. They offer strong face validity and candidate acceptance, and they reduce the inferential distance between predictor and criterion. Their limitations are scalability, feasibility for abstract managerial roles, and the adverse impact reality that their reputation has not caught up with. Use them in later-stage selection where the pool is smaller and the role is concrete enough to be sampled.
Assessment centres are both an alternative and a complement to psychometric testing for senior, specialist, or leadership roles. OPM and the UK’s Civil Service Fast Stream both use them as gateways for roles where decision quality justifies the cost. The logic is sound: deploy more diagnostic, more expensive methods where the hire’s impact is largest. Use simpler screens for volume roles. That is rational segmentation, not a compromise.
Structured reference checks are treated by most organizations as a formality and by the evidence as an underused supplementary method. OPM notes that adding structure to reference checking increases its validity and usefulness. A standardized telephone reference with specific behavioural questions aligned to the competency model is better evidence than an open-ended “would you rehire?” call. It is also a risk management tool for negligent hiring claims.
Competency frameworks, built properly with SME input and validated against actual role performance, are not selection tools by themselves. But they are the foundation that makes every other selection method more defensible. An organization with a well-built, current competency framework can design structured interviews, work samples, assessment centre exercises, and psychometric selection more coherently than one without. Cambridge Assessment’s work on validating competence frameworks treats the framework itself as an object requiring evidence and review, not an unexamined premise that downstream tools are built on.
AI-based screening tools are the most prominent contemporary addition to this landscape and require the most governance. The ICO’s 2026 recruitment report, the DSIT Responsible AI in Recruitment guidance, and EEOC guidance all make the same point: automation does not waive the obligations that applied to traditional selection. It layers new transparency, impact assessment, and accountability requirements on top of them. AI-assisted simulation scoring may be genuinely useful in the right governance context. Unguided algorithmic ranking of applicants based on resume language patterns or video sentiment analysis is in a significantly riskier position, both legally and predictively.
A practitioner checklist for organizations using these tools
Run through these questions honestly. Not the vendor’s answers. Your own.
Check 1: Do you have a current job analysis? Not from three years ago. Not from when the role was created. One that reflects current duties and performance requirements. If not, your construct validity argument starts from zero. OPM’s assessment guidance is direct that job analysis precedes tool selection, not the other way around.
Check 2: Have you read the technical manual? Not the brochure. The technical manual. It shows reliability coefficients, validity study conditions, norm group composition, standard errors of measurement, and fairness evidence. If a vendor will not provide a full technical manual, that is a material red flag.
Check 3: What is the precision at your cut score? A test with good overall internal consistency can still have high measurement error at the specific score region where you are making classification decisions. The AERA/APA/NCME Standards require that evidence of reliability be appropriate for each intended score use. If you do not know the standard error of measurement around your threshold, you do not know how reliable your pass/fail decisions actually are.
Check 4: Have you run adverse impact analysis on your own data? Not the vendor’s claim. Not a published meta-analysis. Your applicant pool, your selection funnel, your outcome data, broken down by protected characteristic. The EEOC’s Uniform Guidelines require this analysis when a selection procedure is used, and the 4/5ths rule provides a working threshold for detecting meaningful disparate impact. If you have not run this, you do not know whether your system is discriminating.
Check 5: Are your accommodation processes designed in, not bolted on? Reasonable adjustments under the Equality Act and ADA-required accommodations need to be part of the assessment design, not an exception pathway created after complaints. Timing accommodations, alternative formats, assistive technology compatibility, and clear candidate-facing instructions about how to request support are legal requirements. The EEOC’s guidance is clear on employer obligation under the ADA for pre-employment testing contexts.
Check 6: What have you told candidates about what is being assessed and why? Hausknecht, Day, and Thomas show that perceived relevance and transparency correlate with favourable organizational perceptions and offer acceptance intentions. The UK government’s Responsible AI in Recruitment guidance requires clear signposting when automation is involved. The ICO requires transparency about automated decision-making and the right to contest. These are not just good practices. They are legal and ethical obligations with documented candidate-experience consequences.
Check 7: When did you last revalidate? The AERA/APA/NCME Standards and SIOP Principles both treat ongoing monitoring as a requirement of responsible assessment practice. Selection systems decay when roles change, labor markets shift, platforms update, or applicant pools evolve. If your validation evidence predates significant changes to the role, the platform, or the applicant population, it is not current evidence.
Check 8: Have you audited your automation for ADM compliance and bias? In the UK, the ICO’s recruitment report found that many employers relying on solely automated decision-making were doing so without adequate safeguards. The DSIT guidance requires data protection impact assessments, equality impact assessments, and algorithmic impact assessments. In the US, the EEOC’s guidance covers algorithmic selection tools under existing anti-discrimination law. If you have procured an AI-enabled assessment platform without completing these assessments, you have a compliance gap.
The bottom line
Psychometric testing is not pseudoscience, and it is not a silver bullet. It is a family of methods whose value depends almost entirely on disciplined design, disciplined deployment, and continuous monitoring. The authoritative guidance from AERA/APA/NCME, SIOP, EEOC, OPM, the ICO, and the UK government’s DSIT all converge on the same conclusions, even though they come from different regulatory traditions and different continents.
Start from job analysis. Define the construct with precision. Demand reliability evidence for the specific use, not just general alpha. Validate locally or at minimum scrutinize the transportability of vendor evidence to your context. Build accommodation and accessibility in from day one. Monitor adverse impact continuously in your own data. Document cut score rationale as the policy decision it is. Subject automated components to impact assessments and maintain human review and contestability.
The organizations doing this well are running better hiring processes, making more defensible decisions, and producing better candidate experiences. The organizations that bought a platform because a competitor uses it, set a cognitive cut-off because the vendor said so, and called the whole thing “objective” without running a single adverse impact analysis: they are sitting on a liability, and they are almost certainly losing candidates they should have hired.
The most consequential insight from a decade of revised psychometric research is not that any single tool has been discredited. It is that the best-performing system in nearly every context is a combination of structured interviews, job-relevant work evidence, and carefully selected psychometric inputs, governed by ongoing monitoring, and anchored to a clear competency model. That is not glamorous. It does not come with a vendor demo. But it is what the evidence actually supports.
Organized sources and further reading
Regulatory and standards frameworks
AERA/APA/NCME: Standards for Educational and Psychological Testing The definitive technical standards for test development, validity, reliability, fairness, and appropriate use in employment contexts. Now available open access. The authoritative reference for any validity argument in personnel selection.
SIOP: Principles for the Validation and Use of Personnel Selection Procedures, 5th ed. (2018) Professional guidelines for personnel selection, covering validation methodology, job analysis, and fairness. Approved by the APA Council of Representatives as an official policy statement.
EEOC: Employment Tests and Selection Procedures Guidance on lawful use of employment tests under Title VII, ADA, and ADEA. Covers cognitive tests, personality tests, integrity tests, work samples, and algorithmic selection tools. Includes best practices and enforcement context.
EEOC: Uniform Guidelines on Employee Selection Procedures — Questions and Answers The 1978 UGESP Q&A document explaining adverse impact, the 4/5ths rule, validation methods, and employer obligations. Joint guidance from EEOC, DOJ, OPM, and DOL. Still the governing framework for disparate impact in US employment testing.
OPM: Assessment and Selection The US federal government’s comprehensive public resource on assessment design, job analysis, competency modelling, and the evidence behind skills-based hiring.
OPM: USA Hire Platform page for the federal government’s online assessment system. Covers adaptive tests, job simulations, online interviews, delivery options, and the agencies using the platform.
UK DSIT: Responsible AI in Recruitment guidance (March 2024) UK government guidance on procuring and deploying AI in HR and recruitment. Covers equality impact assessments, algorithmic impact assessments, DPIAs, transparency obligations, and contestability requirements.
ICO: Automated decision-making in recruitment — “Recruitment Rewired” (March 2026) ICO report on employer use of automated decision-making in hiring, based on engagement with 30+ employers. Sets out regulatory expectations under UK GDPR and the Data (Use and Access) Act, including requirements for meaningful human involvement, transparency, bias monitoring, and candidate rights.
Documented public-sector programs
Civil Service Fast Stream: online tests (UK) SJT-style scenarios, numerical tests, work style questionnaires, and work-based scenario assessments used at scale in UK Civil Service hiring, with published candidate guidance and reasonable adjustment processes.
Civil Service Fast Stream: Assessment Centre (UK) Virtual half-day assessment centre using Written Advice, Stakeholder Communication, and Personal Development exercises. Each exercise is marked independently by different assessors.
GOV.UK: Civil Service online tests guidance Public-facing explanation of test types, scoring methodology, norm percentiles, and accommodation processes for Civil Service recruitment. The transparency standard most private employers do not match.
Key academic and meta-analytic research
Schmidt & Hunter (1998) General cognitive ability found to be among the strongest predictors of job performance across occupations. The foundational meta-analysis that shaped two decades of practice. Sets the baseline the Sackett re-analysis later revised downward.
Sackett, Zhang, Berry & Lievens (re-analysis) Found that many historical validity estimates had been overestimated by .10 to .20, due to overcorrection for range restriction in earlier meta-analyses. Structured interviews moved to the top of the revised validity rankings. The most important recent revision of the evidence base.
Ones, Viswesvaran & Schmidt Integrity tests show a mean operational validity of .41 for supervisory ratings of job performance, and predict a broader range of counterproductive behaviors beyond theft. The strongest published evidence for integrity test use in selection.
Barrick & Mount Among the Big Five personality dimensions, conscientiousness has the most consistent relationship with job performance across occupations. The foundation of personality-based selection research.
Roth, Bobko & McFarland Work-sample validity closer to .33 corrected than the widely repeated .54 estimate. Also found larger Black-White applicant score differences than most practitioners assumed. Corrects two common myths: work samples are not as valid as often claimed, and do not automatically reduce adverse impact.
McDaniel et al. Interview validity depends heavily on structure and content. Structured interviews substantially outperform unstructured ones. Core evidence for the structured interview as a primary selection tool, consistent with Sackett’s revised rankings.
Hausknecht, Day & Thomas Interviews and work samples are perceived more favourably by applicants than cognitive tests, which are in turn perceived more favourably than personality inventories and honesty tests. Shows that a technically valid process can still damage employer brand and offer acceptance rates if candidates find it opaque or irrelevant.
Hogan, Barrett & Hogan Response distortion in real applicant personality testing may be less damaging in practice than laboratory faking studies imply. Nuances the faking debate without eliminating validity concerns for high-stakes selection use.
If you are going through a hiring process and the assessment stage keeps stopping your application, the answer is usually not the psychometric test itself. It is your positioning before you get there. The resume and approach that get you past initial screening to the point where your actual ability can be assessed: that is what we fix together. Work with me at Career Launch Campus and let’s sort the foundation before your next application.




My view is that psychometric testing should rarely be used as a standalone gate.
A stronger process usually combines structured interviews, job-relevant work evidence, and carefully selected assessments.
Do you agree, or do you think testing deserves more weight in hiring decisions?
For anyone working in recruitment or HR:
When was the last time your company actually validated its assessment process against real job performance, candidate outcomes, and adverse impact?
Not the vendor’s evidence. Your own data.