AI-Supported Statistical Inference
Model-informed priors, calibration and human anchoring
This page develops one possible response to the methodological problems examined in Current Limits to AI-Assisted Test Development. That chapter argues that modern test-development samples may be demographically diverse while remaining cognitively selective. Online participants are often unusually experienced in reading instructions, navigating tests and completing abstract tasks. Consequently, the item difficulties and structures estimated from them may not transfer accurately to less practised operational populations. Artificial intelligence does not automatically solve this problem. Synthetic respondents inherit the assumptions, biases and omissions of the models that generate them. Nevertheless, AI may provide a new source of provisional information during the early stages of test development.
The status of this proposal
The proposal examined here is therefore not that simulated people should replace real people. It is that model-generated evidence may provide useful structural hypotheses and prior estimates that can subsequently be tested, corrected and anchored using carefully selected human data. The central question is:
Can model-informed priors reduce the amount of human data required during early psychometric development without compromising validity, fairness or psychological meaning?
This is presented as a research programme rather than an established method.
From synthetic samples to provisional inference
Psychometric test development has traditionally treated human responses as the principal source of information about:
- item difficulty and discrimination;
- dimensional and factor structure;
- population distributions;
- subgroup differences;
- reliability and validity;
- norms and decision thresholds.
This remains indispensable wherever a test is intended to make claims about human beings. Generative models introduce a second possible source of information. A large language model has encountered extensive descriptions of human behaviour, attitudes, abilities, social roles and cultural differences. When prompted to simulate responses, it may reproduce some regularities found in human datasets. Research on synthetic or “silicon” samples suggests that language models can sometimes approximate aggregate response patterns and relationships between variables. However, these simulations also exhibit systematic distortions. They may compress variance, underrepresent unusual responses, reproduce cultural stereotypes, or generate patterns that are internally coherent but unrepresentative of any actual population. The appropriate use of model-generated data is therefore not as a substitute population. It is as a fallible source of prior information.
A formal statement of the problem
Let:
- (P^*) denote the true joint distribution of item responses in a specified target population;
- (D_H) denote a finite dataset collected from human respondents in that population;
- (M) denote a trained generative model;
- (P_M) denote the distribution of responses generated by that model under a specified set of prompts, personas and sampling conditions;
- (\theta) denote the psychometric parameters to be estimated.
Classical psychometric inference estimates (\theta) primarily from (D_H).
AI-supported inference uses information derived from (P_M) to construct provisional expectations about (\theta), while retaining (D_H) as the source of empirical anchoring.
In Bayesian form: [ p(\theta \mid D_H, M) \propto p(D_H \mid \theta),p(\theta \mid M) ]
Here, (p(\theta \mid M)) is a model-informed prior. It may express provisional expectations about item difficulty, discrimination, dimensionality, covariance, local dependence or subgroup functioning.
The human data then determine how far those expectations survive contact with the target population. This formulation makes an essential distinction. A model can generate an arbitrarily large number of simulated responses, so random simulation error can be made small. But repeated sampling does not create unlimited independent information. Bias arising from the model, its training data, the prompts or the simulated personas remains, regardless of how many responses are generated. A million synthetic respondents generated by the same model are not equivalent to a million independently sampled people.
Three possible forms of AI-supported inference
1. Structural prototyping
The most immediate use of model-generated evidence may be exploratory rather than normative. AI simulations can be used to investigate:
- whether candidate items appear to form the intended dimensions;
- whether two items are semantically or functionally redundant;
- whether some items depend on unintended background knowledge;
- whether particular phrasings introduce local dependence;
- whether response options are interpreted consistently;
- whether provisional factor structures are plausible;
- whether proposed constructs have distinguishable subcomponents;
- whether some items are likely to generate ceiling or floor effects.
At this stage, the model is not being asked to establish a population norm. It is being used as a rapid and inexpensive environment in which weak hypotheses can be discarded before scarce human participants are recruited. This is analogous to simulation elsewhere in science and engineering. A simulation does not certify that a design will work in the physical world, but it may identify obvious failures and improve the design of the empirical study. Structural prototyping may be particularly useful when thousands of AI-generated candidate items must be reduced to a manageable number for human piloting.
2. Model-informed prior formation
A more ambitious use is to derive provisional item or structural parameters from model-generated responses. For example, the model might suggest that:
- one item is probably easier than another;
- an item is likely to discriminate most strongly near a particular trait level;
- two attitude items may share residual variance;
- a proposed scale may contain two dimensions rather than one;
- an item may function differently across simulated cultural or occupational contexts.
These estimates should not initially be treated as measurements of the human population. They are prior hypotheses about what human data may show. Their usefulness is empirical. A model-informed prior is valuable only to the extent that it improves prediction when tested against independent human observations. Different models, prompts and persona-construction methods may produce different priors. Their performance can therefore be compared. Priors that repeatedly fail to predict human data should be weakened or rejected; those that transfer successfully may be retained and progressively refined. The model is not a psychometric oracle. It is a provisional inference system whose claims remain open to correction.
3. Criterion-anchored provisional calibration
A third form of provisional inference is particularly relevant to educational and ability testing. New items are not always created in a complete informational vacuum. National curricula, examination specifications, developmental expectations, teacher knowledge and established tests contain evidence about what children at different ages or stages are ordinarily expected to understand. This evidence can be used to construct provisional priors for new items before full standardisation. For example, an item may require:
- vocabulary normally introduced at a particular curriculum stage;
- a mathematical operation expected by a specified school year;
- a level of reading comprehension associated with an age band;
- the coordination of concepts taught at different points in the curriculum.
AI can analyse these requirements and compare them with curriculum expectations. It can then help estimate a provisional difficulty range or age location for the item. Such estimates would not constitute norms. Nor would curriculum expectations establish the actual population mean or standard deviation. They would provide an initial, criterion-anchored expectation that could guide item generation and early piloting. This approach differs from simply asking an AI model to simulate children. The prior is anchored partly in an external educational framework rather than solely in patterns reconstructed from the model’s training history. Teacher judgements, curriculum evidence and performance on existing anchor items could all be incorporated. Formal human standardisation would still be required before operational use, but early item development could proceed from an informed starting point rather than from complete uncertainty.
Where might model-informed priors work best?
The value of AI-supported inference is likely to vary greatly across domains. It may be most useful for constructs that:
- are extensively represented in ordinary language;
- concern attitudes, preferences, values or social judgements;
- have well-described behavioural manifestations;
- have already been measured in many published studies;
- can be expressed through natural-language scenarios;
- do not depend heavily on sensory, motor or time-sensitive performance.
Examples might include attitudes towards technology, political cynicism, occupational preferences, interpersonal styles or broad personality descriptions. Greater caution is required where performance depends on:
- perceptual discrimination;
- spatial rotation;
- motor coordination;
- processing speed;
- working-memory capacity;
- unfamiliar non-verbal stimuli;
- developmental or clinical processes poorly represented in text;
- rare characteristics or extreme score ranges.
A language model may know how people describe spatial ability without reproducing the processes by which people actually solve spatial problems. The distinction is between representing discourse about a psychological characteristic and representing the characteristic itself.
Human data remain the anchor
The purpose of model-informed priors is not to demote human evidence. It is to use human evidence more deliberately. A carefully designed human study may be more informative than a much larger convenience sample if it represents the cognitive and operational conditions under which the test will actually be used. Human anchoring may be required to:
- determine whether the proposed structure occurs in the target population;
- estimate discrepancies between model predictions and actual responses;
- establish the score metric;
- test measurement invariance;
- examine differential item functioning;
- recover variance compressed by synthetic samples;
- include low-frequency and extreme response patterns;
- assess comprehension and cognitive accessibility;
- observe hesitation, misunderstanding and abandonment;
- connect scores with external criteria and outcomes;
- establish operational norms and cut-scores.
There can be no universally appropriate human sample size for this stage. The amount of evidence required will depend on the number of items and dimensions, the complexity of the model, the heterogeneity of the population, the intended score interpretation, subgroup analyses and the consequences of error. A low-stakes research scale and a high-stakes clinical, educational or occupational test require very different levels of validation.
A possible development sequence
1. Construct specification
Human theorists define the proposed construct, its conceptual boundaries, its facets and its expected relationships with other variables. This stage cannot be delegated entirely to statistical modelling. A stable factor solution does not by itself establish that a psychologically coherent construct has been measured.
2. AI-assisted item generation
A generative model produces candidate items under explicit constraints concerning:
- construct coverage;
- reading level;
- cultural accessibility;
- working-memory demand;
- response format;
- undesirable overlap between items;
- potential sources of bias;
- intended difficulty range.
Human experts review the items for conceptual and ethical adequacy.
3. Structural prototyping
One or more models simulate responses under varied prompting conditions. These simulations are used to identify possible redundancy, dimensionality, local dependence and gross difficulty differences. Disagreement between models should be retained as information rather than averaged away automatically. Strong disagreement may indicate that an item is ambiguous or highly dependent on contextual assumptions.
4. Provisional parameter formation
The synthetic results, curriculum evidence, expert judgement and findings from related instruments are combined to form provisional priors. The provenance of each prior should be recorded. A prior based on curriculum expectations is epistemically different from one based on a simulated demographic persona or a related published scale.
5. Targeted human anchoring
Human participants are recruited to test the assumptions most likely to fail. Sampling should include the people whose cognitive behaviour is least well represented in online development panels, including where appropriate:
- inexperienced test-takers;
- people with limited literacy;
- second-language users;
- individuals at the lower and upper ends of the intended ability range;
- participants using the devices and interfaces expected in practice;
- relevant cultural, educational and occupational groups.
The objective is not merely to obtain clean responses. It is to discover how and why items fail.
6. Prior-data conflict analysis
The model-informed predictions are compared with human observations. This stage asks:
- Which item parameters transferred successfully?
- Where did the model underestimate or overestimate difficulty?
- Which forms of variance were absent from the synthetic sample?
- Which subgroups were misrepresented?
- Did the model reproduce covariance without reproducing distributions?
- Did prompting choices affect the estimates?
- Are discrepancies random, or do they reveal systematic model bias?
Large prior-data conflicts should not be hidden by regularisation. They are central findings about the limits of the model.
7. Operational validation
Before consequential use, the instrument must be validated under conditions that reproduce its intended administration ecology. This includes the stakes, timing, devices, instructions, motivation, supervision and emotional context of actual use. The resulting evidence, not the synthetic data, must support the final interpretation of scores.
8. Continuing Bayesian revision
As genuine operational data accumulate, item and structural parameters can be updated. Model-informed priors may be progressively weakened as direct evidence grows. Where population change occurs, the system can also test whether older priors remain applicable. AI-supported inference is therefore not a one-time substitution for standardisation. It is a staged process in which provisional knowledge is gradually exposed to human reality.
What can go wrong?
Variance compression
Synthetic samples often produce answers that are more moderate, orderly and internally consistent than human responses. Rare combinations, extreme positions, misunderstandings and apparently irrational responses may be underrepresented. This is especially serious where a test must measure the tails of a distribution.
Invariance illusions
A model may produce an elegant and stable factor structure across simulated demographic groups because those groups are all being generated by the same underlying system. Apparent invariance within a model is not evidence of invariance among human populations.
Prompt dependence
Synthetic results may vary with subtle changes in instructions, persona descriptions, response order, temperature settings or model versions. The prompt and sampling procedure therefore become part of the measurement method and must be documented.
Training-data opacity
A model-informed prior implicitly depends on the model’s training history, data selection, fine-tuning and safety procedures. These may be only partly known. Changes made by a model provider may alter the generated response distribution without warning.
Stereotype substitution
When asked to simulate a demographic or cultural group, a model may reproduce textual stereotypes about that group rather than the response processes of actual members. Adding more elaborate persona descriptions does not necessarily solve this problem. It may simply produce more elaborate stereotypes.
Model convergence
Using several models does not guarantee independent evidence if they were trained on overlapping internet corpora, similar preference data or related synthetic outputs. The apparent agreement of models may reflect shared ancestry rather than shared accuracy.
Feedback and model collapse
If future models are increasingly trained on earlier model-generated material, synthetic psychometric results may become self-reinforcing. A model may eventually reproduce patterns because previous models generated them, not because they correspond to human behaviour.
Premature authority
The most serious risk is that provisional estimates acquire the appearance of established psychometric parameters merely because they are precise. Precision within a model is not the same as validity outside it.
Validity and governance
AI-supported inference expands the validity argument required for a test. In addition to describing the human sample, developers must document:
- the model and model version used;
- the date of generation;
- prompt construction;
- sampling settings;
- simulated persona definitions;
- the source of curriculum or criterion information;
- the relationship between model and human estimates;
- prior-data conflicts;
- subgroup failures;
- changes introduced during updating;
- which conclusions depend primarily on synthetic rather than human evidence.
Responsibility also becomes distributed. Bias may enter through the construct definition, item generator, model training data, prompts, simulated personas, human anchor sample, statistical model or operational context. This does not make responsibility impossible to assign. It makes provenance and auditability indispensable.
The philosophical change
Classical psychometrics locates much of its empirical reference in a defined sample collected at a particular time. Model-informed psychometrics introduces a more distributed form of reference. The prior may contain traces of many populations, documents, theories and historical descriptions incorporated during model training. This creates three philosophical questions.
What is being measured?
If a model reproduces a plausible structure of attitudes, does that structure belong to contemporary human beings, to the historical textual record, or to the model as a socio-technical artefact?
Where does the evidence reside?
The relevant epistemic unit is no longer only the visible response dataset. It includes the model’s training history, prompting method and alignment procedures.
What is the status of the synthetic respondent?
A synthetic persona is not a participant drawn from a population. It is a conditional production of a model. Treating it as an ordinary respondent risks confusing simulation with observation. These questions are not peripheral philosophical complications. They determine what conclusions the resulting statistics can legitimately support.
Testable predictions
The proposal generates empirical predictions.
- Model-informed priors should improve early parameter estimates when human samples are small, but their advantage should decline as representative human evidence increases.
- Transfer should be better for linguistically represented attitudes and social constructs than for perceptual, motor, speeded or non-verbal abilities.
- Synthetic samples should reproduce some covariance and factor structures more successfully than score distributions, variances and extreme response patterns.
- Priors derived from curriculum and established criterion frameworks should perform differently from priors derived solely from synthetic personas.
- Model errors should not be evenly distributed. They should cluster around culturally specific meanings, cognitively unusual response strategies and groups poorly represented in the training data.
- Prompt and model variation should produce measurable variation in provisional item parameters.
- Human samples selected for cognitive and operational representativeness should reveal larger model discrepancies than highly practised online-panel samples.
These predictions make the proposal open to refutation.
A research agenda
The immediate research task is not to ask whether AI can replace psychometric standardisation. It is to determine which forms of provisional model-based information survive independent human testing. Research should compare:
- model-informed and weakly informative priors;
- synthetic-persona and curriculum-anchored priors;
- single-model and multi-model estimates;
- online-panel and cognitively representative human samples;
- attitudinal, personality and ability domains;
- group-level and individual-level prediction;
- model predictions before and after human anchoring;
- low-stakes and high-stakes applications.
The most important outcome is not whether the synthetic data look realistic in isolation. It is whether they improve prediction of independently collected human evidence.
Conclusion
AI-supported statistical inference may offer psychometrics a useful intermediate stage between theoretical item construction and full human standardisation. Its most defensible role is not to manufacture an artificial population and declare it normative. It is to generate provisional structures, hypotheses and priors that make subsequent human research more targeted and informative. Human data remain indispensable because tests make claims about people, not about models. However, human evidence need not always enter the development process as an undifferentiated mass sample used for every purpose. AI may assist with exploration. Curricula and established knowledge may provide provisional anchors. Carefully selected human participants can then expose error, restore missing variance, establish meaning and connect the measurement system to lived performance. The central scientific question is therefore not whether synthetic respondents can replace human beings. It is:
Which psychometric assumptions can be provisionally informed by artificial intelligence, how much carefully chosen human evidence is required to correct them, and where must the authority of the model end?
References
-
Argyle et al. (2023) show that GPT-3–based “silicon samples” can approximate opinion distributions of real subpopulations when conditioned on demographic backstories. They introduce the idea of algorithmic fidelity: does a model reproduce not just means but correlation structures similar to real data? Cambridge University Press & Assessment
-
Multiple follow-ups (e.g. public-opinion simulations, policy preference surveys) show that LLM-generated samples often match aggregate survey patterns, but with systematic distortions: reduced variance, more positivity, and item-level heterogeneity. ACM Digital Library
-
Pellert et al. (2024) treat LLMs as objects of psychometrics—administering inventories to models and analysing factor structure, reliability, and validity—essentially building a bridge between classical psychometrics and LLM behaviour. PubMed SAGE Journals
-
Google DeepMind researchers have found ways to generate large corpora while preserving informational content and avoiding model collapse. These corpora contain massive collections of digital text and data used to train, test and analyse large language models, including books, websites, articles, and code, and serve as the raw material that enables LLMs to learn patterns, grammar, and information about the world. Business Insider arXiv
-
Studies detecting AI bots passing as humans in online surveys; panel contamination is no longer hypothetical. The Times
-
Variance collapse – synthetic samples often under-represent extreme or rare patterns compared to real data. ACM Digital Library
- Feedback / model collapse – if you train tomorrow’s models on data generated by yesterday’s models, you risk a form of informational inbreeding. DeepMind’s GDR work is explicitly a response to this threat. Business Insider