Psychometrics Tomorrow
Psychometrics is the science of measuring psychological characteristics: abilities, aptitudes, personality traits, attitudes, interests and other aspects of human behaviour. Its methods have shaped education, employment, clinical psychology and scientific research for more than a century.
The foundations of the discipline remain familiar. A psychological test must produce sufficiently reliable scores, measure what it claims to measure, distinguish meaningfully between individuals and support decisions that are appropriate for the purpose for which the test is being used.
These requirements will not disappear. But the technologies through which psychological information is collected, analysed and interpreted are changing rapidly. The psychometrics of tomorrow will therefore be both continuous with the discipline we know and markedly different in its methods.
From Paper Tests to Digital Assessment
The first major transformation has already occurred. Tests that were once administered with paper, pencil, stopwatches and printed manuals are now routinely delivered online. This change may initially appear superficial. A questionnaire presented on a screen can seem little different from the same questionnaire printed on paper. But digital assessment makes possible forms of measurement that paper tests could never provide. A computer can record not only whether an answer is correct, but how long the respondent took, whether an answer was changed, which alternatives were considered, how consistently the person responded and how performance developed throughout the assessment. The resulting data can reveal aspects of the response process that were previously invisible. Psychometrics will
Adaptive Testing
Traditional tests usually present the same questions to everyone. This is convenient, but often inefficient. Some questions are much too easy for one respondent and much too difficult for another. omputer-adaptive testing offers an alternative. After each response, the system estimates the respondent’s current level and selects the next item accordingly. A person who answers correctly may receive a more difficult question, while someone who struggles may receive an easier or more informative one. Item Response Theory provides the statistical foundation for many such systems. Rather than treating a test simply as a fixed collection of questions, IRT models the relationship between individual respondents and individual items. Characteristics such as item difficulty and discrimination can be estimated, while a respondent’s score is treated as an estimate of their position on the ability or personality characteristic being measured.
This makes it possible for different people to answer different questions while still receiving scores on a common scale. It also allows the precision of each score to be estimated and the test to concentrate its questions where they will provide the most information. Adaptive tests can therefore be shorter, more precise and less frustrating than conventional fixed tests. They may also assess people across a wider range of ability without requiring everyone to complete a very long examination. A fuller introduction to the underlying models and their application can be found in IRT and Adaptive Testing, which traces the development of Item Response Theory from the Rasch model to contemporary adaptive assessment, multidimensional measurement, fairness monitoring and machine-learning approaches to item selection. The challenge is to preserve transparency, comparability and fairness. Test users must still be able to understand what a score means, how accurately it has been estimated and whether the assessment functions equivalently for different groups.
More Natural Forms of Assessment
Many psychological characteristics are difficult to measure through multiple-choice questions alone. Communication, creativity, judgement, collaboration and practical problem-solving are often better demonstrated through behaviour than through self-report. Future assessments are therefore likely to include more realistic tasks. Candidates may be asked to respond to workplace scenarios, analyse complex information, participate in simulations, write explanations or solve problems in interactive environments. These tasks can provide richer evidence, but they also create new psychometric difficulties. The more realistic an assessment becomes, the more complicated its scoring may be. Performance can be influenced by prior experience, language, confidence, familiarity with technology and many other factors unrelated to the characteristic being measured. The task of the psychometrician will be to determine which features of performance provide valid evidence and which introduce irrelevant variation. Realism by itself does not guarantee validity. A sophisticated simulation can still be a poor test. The quality of an assessment depends not on how impressive it looks, but on the evidence supporting the interpretation and use of its scores.
Artificial Intelligence in Test Development
Artificial intelligence is likely to influence almost every stage of test construction. It can already assist with generating possible test items, producing alternative versions, checking wording, identifying duplicated content and estimating the likely difficulty of questions. It may help test developers create large item banks more quickly and adapt assessments for different languages, populations or contexts. AI may also support the scoring of written answers, spoken responses and complex performances that previously required human raters. This could make certain forms of assessment faster and more widely available. However, automatically generated items and scores cannot simply be assumed to be valid. An AI system may produce questions that appear plausible to experts but contain hidden ambiguities, excessive reading demands, unintended clues or assumptions that make them unnecessarily difficult for the people who will actually take the test.
There is a related danger in the growing use of online participant panels to pilot and calibrate new items. Such samples may be demographically diverse while remaining unusually experienced with online studies, complex instructions and abstract test formats. Items that appear straightforward when answered by practised online participants may prove confusing or substantially more difficult when administered to cognitively less experienced populations. These problems are examined more fully in Current Limits to AI-Assisted Test Development. The chapter argues that modern test development must retain cognitive realism: items should be evaluated not only through clean statistical data, but through evidence about how they are read, interpreted and attempted by the full range of people for whom the test is intended.
Human expertise will therefore remain essential. AI can propose, classify and analyse, but psychometric validation must establish whether the resulting items measure the intended characteristic, whether their difficulty has been estimated realistically and whether their scores support the interpretations placed upon them. The arrival of AI does not remove the need for psychometrics. It increases it.
Assessing Performance in an AI-Assisted World
Psychometrics will also need to respond to a more fundamental change. People increasingly use digital assistants and generative AI when studying, writing, analysing information and solving problems. This raises a difficult question: should an assessment measure what a person can do unaided, or what they can achieve using the tools that will be available in real life? There will be no single answer. Some assessments will continue to require unaided performance because they are intended to measure underlying knowledge or skill. Others will need to examine how effectively a person can use AI, evaluate its suggestions, detect its errors and combine automated assistance with independent judgement. The distinction resembles that between testing mental arithmetic and testing the ability to solve a real financial problem with access to a calculator. Both can be legitimate, but they measure different things. Future assessments will need to state much more clearly what forms of assistance are permitted and what competence the resulting score represents.
Validity Will Become More Important
As assessment technology becomes more powerful, the central psychometric question will remain unchanged:
What conclusions are we entitled to draw from this score?
Large datasets and complex algorithms can produce highly accurate predictions without explaining what has been measured. A system may predict examination performance, employee turnover or purchasing behaviour, but prediction alone does not establish psychological meaning. A correlation can be useful without identifying a stable human characteristic. It may reflect temporary circumstances, social inequalities, access to technology or features of the data-collection process. Psychometrics must therefore resist the temptation to treat every successful prediction as a psychological measurement. Construct validity, criterion validity, reliability, measurement invariance and careful examination of alternative explanations will remain indispensable. Indeed, they will become more important as models grow more complicated and their outputs become harder to inspect.
Fairness and Bias
Tests can create opportunities, but they can also restrict them. Their consequences are particularly serious in education, employment, healthcare and legal decision-making. Future psychometrics must therefore pay close attention to fairness. This includes examining whether items function differently across groups, whether scores have comparable meanings, whether prediction errors are distributed unequally and whether apparently neutral procedures reproduce existing disadvantages. Statistical fairness is necessary, but it is not sufficient. A test may satisfy a technical definition of fairness while still being used for an inappropriate purpose. Conversely, differences between groups do not automatically demonstrate that a test is biased. Fairness must be investigated through evidence, interpretation and the consequences of use. It cannot be established by a single statistic. Test developers must also recognise that populations, language and social conditions change. Validation is not an event completed when a test is first published. It is a continuing process.
Privacy and Consent
Digital assessment can collect much more information than a conventional test. Response times, keystrokes, speech, facial movements, patterns of revision and other behavioural signals can all potentially be recorded. Some of these data may improve measurement. Others may be intrusive, unreliable or unnecessary. The fact that information can be collected does not mean that it should be. Test takers should know what is being recorded, why it is relevant, how long it will be retained and who will be able to use it. Psychometric value must be weighed against privacy. Measures based on obscure behavioural traces may appear objective while being difficult for the individual to understand or challenge. Responsible assessment will require restraint as well as innovation.
The Changing Role of the Psychometrician
The psychometrician of tomorrow will still need expertise in psychological theory, research design and statistical modelling. But the profession will increasingly overlap with computer science, data science, artificial intelligence, interface design, ethics and law. Psychometricians will need to understand how algorithms are trained, how digital platforms influence behaviour and how automated decisions are made. At the same time, technical specialists will need to understand why measurement is not merely a matter of collecting data and fitting a predictive model. The distinctive contribution of psychometrics is its concern with the meaning of scores
A number produced by a computer is not automatically a measurement. It becomes meaningful only when there is evidence connecting the observed performance to a clearly defined interpretation.
What Should Remain Constant
The technology of assessment will continue to change. Tests may become adaptive, interactive, automated and embedded within everyday digital environments. But the fundamental responsibilities of psychometrics should remain constant:
- to define clearly what is being measured;
- to collect evidence that the interpretation is justified;
- to estimate the precision and limitations of the resulting scores;
- to investigate fairness across individuals and groups;
- and to ensure that assessments are used only for purposes they can properly support.
Psychometrics tomorrow will possess tools of extraordinary power. Whether those tools improve human decision-making will depend upon the scientific standards and professional judgement with which they are used. The future of psychometrics is therefore not simply a future of more data, larger models or faster testing. It is a future in which measurement must become more adaptive and informative while remaining accountable to the people whose lives may be affected by it.