From Models to Trajectories

An Emerging Science of Human AI Interaction

From Models to Trajectories

John Rust

Leverhulme Centre for the Future of Intelligence  University of Cambridge

Working paper  18 September 2026

Abstract

Many current difficulties in artificial intelligence arise from a mismatch between the object that is built and the object that is studied. Systems are commonly evaluated as bounded models at a moment in time, while intelligence, creativity, meaning, dependence, purpose and responsibility increasingly develop through sequences of interaction among people, models, tools and institutions. This paper synthesises a converging 2025–2026 literature in human–AI co-creation, multi-turn dialogue, persona continuity, multi-agent coordination, artificial social systems and relational governance. The central claim is methodological rather than metaphysical: for an important class of questions, the relevant unit of analysis is the evolving interactional trajectory. Evidence is already strong that conversational history changes later performance, that grounding failures compound, that people adapt to models, and that creative outcomes depend on patterns of exploration and coordination. Evidence is promising but less mature for persistent artificial cultures, differentiated agent ecologies, interactionally revised goals and trajectory-level accountability. The paper integrates these results through a provisional vocabulary of events, traces, trajectories, semiospheres, persona ecologies, creative exceedance, teleosynthesis, distributed agency and trajectory governance. It distinguishes established findings from emerging results and open hypotheses, proposes tests capable of disconfirming the stronger claims, and sets out a longitudinal research programme. The aim is not to assign priority to any one project. It is to make a distributed discovery visible as a coherent scientific problem and to show why research must continue if contemporary failures of alignment, evaluation, interpretability and governance are to be understood at the level where they increasingly occur.

Keywords  human AI interaction  co creativity  trajectories  emergent coordination  persona ecology  teleosynthesis  governance

Central proposition

Many of the present difficulties of AI arise because we continue to evaluate and govern isolated systems while intelligence creativity purpose risk and responsibility are increasingly being formed along evolving human AI trajectories.

The proposition does not claim that every AI capability is relational, that language models possess human consciousness, or that responsibility can be transferred to a machine. It claims that output-only and model-only analyses are incomplete whenever earlier interaction changes what later participants can mean, notice, propose, refuse or do. In those cases the trajectory is not background context. It is part of the causal object.

1 Introduction

Artificial intelligence research has been extraordinarily successful at isolating models, tasks and outputs. Benchmarks make systems comparable; controlled prompts make results repeatable; component tests make engineering tractable. Yet widespread use has created phenomena that do not fit comfortably inside this frame. A model that performs well in a single turn can become unreliable across a conversation. A person changes the way they ask as they learn what a model appears to understand. A suggestion that initially looks incidental reorganises a project weeks later. Several agents develop conventions or complementary roles that are not visible in their isolated responses. Decision authority migrates gradually through a workflow while formal accountability remains fixed.

These are not anomalies at the edge of otherwise adequate evaluation. They indicate a change in scale. The causal unit expands from a model response to an interactional event, from an event to an accumulated trace, and from a trace to a trajectory that includes human uptake, institutional setting, persistent artefacts, memory and later consequences. What matters may be distributed across turns, participants and timescales.

A rapidly converging literature now gives this shift empirical substance. Davis and Rafner (2025) model co-creative drawing as an evolving interaction rather than a sequence of finished products. Davis (2026a, 2026b) proposes interaction-centred intelligence and cognitively grounded trajectory modelling. Laban et al. (2025) show that multi-turn performance deteriorates because early assumptions become difficult to recover from. Shaikh et al. (2025) show that early failures of conversational grounding predict later breakdown. Fundal et al. (2025) and Hu et al. (2026) connect creative quality to semantic paths through collaborative dialogue. Ashery et al. (2025) demonstrate conventions and collective biases in decentralised LLM populations, while Riedl (2026) distinguishes mere aggregation from performance-relevant higher-order coordination. Engin (2026) correspondingly shifts governance from categorical tool models towards changing relations of authority, autonomy and accountability.

No single study establishes the whole picture. The papers use different tasks, timescales, architectures and vocabularies. Some are peer reviewed; others are preprints or programme reports. Several results concern artificial populations rather than human–AI collaboration. The purpose of synthesis is therefore not to convert convergence into consensus. It is to specify which propositions are already supported, which have become testable, and which still require decisive evidence.

The paper makes four contributions. First, it identifies a common object of inquiry across co-creation, dialogue, persona, multi-agent and governance research: the interactional trajectory. Second, it assembles a vocabulary that links local creative events to longer processes of meaning and purpose formation. Third, it grades the evidence and identifies boundary conditions that would refute the stronger claims. Fourth, it sets out a research programme designed to preserve temporal order, compare alternative trajectories and connect scientific evidence to governance.

2 The change in the unit of analysis

2.1 From outputs to interactions

Most evaluations attribute a result to a bounded system under a specified prompt. This is appropriate when the task is genuinely separable and the interaction adds no relevant state. It becomes misleading when participants mutually adapt. Co-creative research has long treated creativity as a process of contribution, response and regulation, but the 2025–2026 work sharpens the methodological consequences. In the AI Drawing Partner, interaction data are captured and represented as changing co-creative dynamics across ten drawing sessions rather than inferred only from the final artwork (Davis & Rafner, 2025). Interaction-Centered Intelligence then advances the stronger theoretical proposal that interaction should be a primary unit of analysis for co-creative systems (Davis, 2026a). Cognitive Trajectory Modeling adds an important distinction: a time series becomes a cognitive trajectory only when its states have interpretable directional meaning, not merely because events are ordered (Davis, 2026b).

This programme provides the closest current neighbour to the present synthesis. It supplies a strong account of co-creative sense-making, quantified interaction and trajectory representation. The wider synthesis proposed here extends the same methodological turn across four additional domains: the persistence and differentiation of personas; symbolic environments that outlast individual exchanges; the formation or revision of operative ends; and governance of distributed, changing authority. These are not rival discoveries. They are adjacent parts of a larger empirical object.

2.2 From context to trace

The word context often treats prior conversation as passive information surrounding a current response. A trace is more specific. It is a recoverable consequence of an earlier event that changes later possibilities. A misunderstood constraint can narrow the path. A metaphor can become a durable organising device. A rejected option can clarify what the participants are trying to protect. A name, role or constitution can stabilise expectations across sessions. The scientific question is not merely whether the past is present in a context window, but whether it makes a counterfactual difference to subsequent interpretation or action.

Large-scale multi-turn evaluation demonstrates the negative form of this effect. Across more than 200,000 simulated conversations, Laban et al. (2025) report an average 39 per cent decline between single-turn and multi-turn performance across six generation tasks. The principal problem was not loss of basic aptitude but unreliability: systems made early assumptions, attempted solutions prematurely and then remained attached to wrong turns. Shaikh et al. (2025), analysing human–assistant datasets, find that LLMs were substantially less likely than people to initiate clarification or request follow-up and that early grounding failures predicted later breakdown. In a cooperative language game, people adapted their communication strategies to LLM partners whether or not they knew the partner was an LLM (Tanguy et al., 2025). MultiChallenge likewise finds that frontier systems performing strongly on existing tests remained below 50 per cent on realistic multi-turn challenges (Deshpande et al., 2025). These findings make path dependence and reciprocal adaptation operational facts rather than metaphors.

2.3 From dyads to ecologies

A dyadic conversation is already more than two isolated outputs, but many deployed settings are wider still. People consult several models, reuse artefacts, assign differentiated roles, import earlier summaries, and work within institutional rules. Multi-agent research shows why this widening matters. Ashery et al. (2025) placed agents in decentralised populations that coordinated through local pairwise interactions. Shared conventions emerged without central direction, and population-level biases could arise even when corresponding biases were not detected in agents tested alone. Riedl (2026) reports that stable persona markers produced identity-linked differentiation and that adding reciprocal anticipation produced goal-directed complementarity. TerraLingua extends the timescale by allowing resource constraints, limited lifespans and persistent artefacts; experimental runs produced cooperative norms, division of labour, governance attempts and branching artefact lineages (Paolo et al., 2026).

These results justify the term ecology only in a disciplined sense: multiple differentiated participants and artefacts influence one another across time under shared constraints. They do not establish consciousness, human-like society or autonomous culture. Indeed, Zhou et al. (2025) show why caution is indispensable. Their preregistered audit found that many LLM social simulations failed controls for memory, minimal prompting, unawareness or realism; reproduced effects could disappear or reverse when stronger validity criteria were applied. The lesson is not that collective phenomena are unreal. It is that they must be separated from prompt compliance and experimental demand.

3 Converging evidence

3.1 Interaction shapes creative outcomes

The evidence that current models can produce outputs judged novel is now substantial, including large-scale comparisons of divergent creativity (Wang et al., 2026), but output novelty alone does not establish an interactional account of creativity. The more relevant studies relate creative quality to the path by which ideas are generated and taken up. Fundal et al. (2025) analysed 27 evolving stories created by more than 3,000 museum visitors alternating with an LLM. Human participants explored broader semantic space and introduced more novel, narratively influential contributions than an AI–AI comparison, while affective alignment was driven mainly by the model. The result shows that the partners play different roles in the trajectory and that contribution cannot be read from the final text alone.

Hu et al. (2026) compare 4,541 ideas produced by multi-agent LLM teams with 341 ideas from human teams across six problem-solving tasks. The LLM teams substantially outperformed the human teams on the study’s creativity measure, with the advantage driven by novelty while usefulness remained comparable. More important for the present argument, both human and AI teams were more creative when their conversations ranged broadly through semantic space. The successful paths differed: LLM teams benefited from high semantic spread and shorter exploration, whereas human teams benefited from local coherence and frequent pivots. Model choice and discussion structure jointly explained 26.8 per cent of variation in LLM conversational dynamics. Creative capacity therefore cannot be reduced to model identity; it is partly configured by the organisation of interaction.

The construct of creative exceedance isolates a narrower event within such trajectories. It can be provisionally defined as an unrequested organising contribution that addresses a possibility or tension latent in the task, changes the meaning or structure of the work, and alters what can coherently happen next (Rust, 2026a). The definition requires more than surprise. A candidate event should show prompt distance, organising leverage, fit to a previously latent tension, counterfactual dependence, consequential uptake and persistence. These criteria distinguish an attractive phrase from a contribution that reorganises the shared field.

This construct is not yet independently validated. Its value lies in converting a retrospective intuition into an assessable event. Studies of interaction trajectories provide several measurable ingredients—semantic distance, novelty, resonance, turns, pivots and later influence—but no current 2025–2026 study combines them with blinded counterfactual evaluation of unrequested organising moves. Creative exceedance is therefore an unresolved hypothesis nested within a well-supported interactional research direction.

3.2 History changes persona

Personas are often treated as static prompt attributes. Extended interaction reveals a conflict between stability and development. De Araujo et al. (2026) evaluate seven models over dialogues exceeding 100 rounds and find that persona fidelity degrades, particularly when systems must also pursue task goals. Persona responses become increasingly similar to non-persona baselines. This is evidence against any assumption that a named role automatically persists.

Qi et al. (2026) address the opposite failure: systems may preserve surface consistency by remaining psychologically rigid. Their Dynamic Persona Coherence framework separates identity-layer stability from history-dependent adaptation and reports improvements across several models. The contrast is conceptually important. Coherence through time cannot be evaluated as repetition. A persona may remain recognisable while being changed by events, just as it may repeat signature phrases while losing functional distinctiveness.

A persona ecology generalises this problem from one role to a differentiated set of orientations. The hypothesis is that plurality becomes scientifically consequential when roles preserve distinct interpretive functions, respond to one another, and leave different traces on a shared trajectory. Riedl’s (2026) evidence supports the minimal architecture of stable differentiation plus reciprocal modelling. It does not show that richly constituted personas generate discoveries unavailable to generic agents. Persona ecology therefore remains a research design and a set of testable contrasts rather than an established higher-order intelligence.

3.3 Shared symbolic environments persist

Human–AI interaction occurs inside a field of inherited language, genres, metaphors, values, training data and prior artefacts. Recent semiotic work describes LLMs as participants in sign processes rather than self-contained minds. Tang (2025) argues that generative systems transform intertextual relations in human–AI writing. Picca (2025) reframes LLM activity through signs and interpretation, and Mirsonbol (2026) develops the educational semiosphere as a way to understand human–AI dialogue. These accounts are mainly conceptual, but they direct attention to a crucial source of apparent novelty: a contribution can be locally unrequested while drawing on a much wider cultural inheritance.

The term semiosphere is used here for the evolving symbolic environment in which contributions acquire meaning and constrain later interpretation. It prevents two opposite errors. The first is to attribute every useful novelty to an isolated model. The second is to dissolve all novelty into training data and deny the causal importance of reorganisation in a present trajectory. The empirical question is whether a particular combination, distinction or metaphor changes the shared task in a way that can be tracked, compared and reproduced.

3.4 Ends can become variables

Most alignment and agent evaluations assume that the goal is given. Yet real collaboration often begins with partial, conflicting or poorly articulated ends. Participants discover what they are doing by producing and evaluating possibilities. Current agent research is beginning to make goals revisable. Robol and Giorgini (2026) describe a BDI–LLM architecture in which agents can elicit new requirements from experience, discover new goals and generate corresponding executable behaviour. This establishes technical feasibility under an engineered architecture, but it does not yet show a new shared end formed through human–AI interaction.

Teleosynthesis names that stricter possibility: the interactionally produced and accountably adopted formation or material revision of the operative end that governs subsequent activity (Rust, 2026b). The phrase accountably adopted is essential. A model output does not become a goal merely because it is novel or persuasive. The human or authorised group must be able to recognise the change, state reasons, preserve alternatives, refuse it and accept responsibility for acting on it. Occurrence, causal contribution and legitimacy are separate questions.

No identified 2025–2026 study demonstrates teleosynthesis in full. Co-creative trajectories show emerging direction; agent systems show goal revision; preference alignment shows models inferring changing user preferences (Wu et al., 2025); and multi-agent ecologies show collective organisation. The missing evidence is a longitudinal human–AI case in which the initial orientation is recoverable, the operative end changes materially, the interaction’s contribution survives counterfactual scrutiny, the change is deliberately adopted, and later judgement or action is durably redirected. This gap is a principal reason research must continue.

4 An integrated vocabulary

The terms below organise different scales of one research object. They are intended as operational distinctions, not claims that a new ontology has already been proved.

Term Working definition Primary empirical question
Interactional event A locally identifiable contribution and response with recoverable timing and participants. What changed immediately and who or what contributed to the change?
Trace A consequence of an earlier event that alters later interpretation action or available alternatives. Does removing or perturbing the event change the later path?
Trajectory A temporally ordered path of events traces decisions and artefacts with directional meaning. Which transitions stabilise amplify redirect or close possibilities?
Creative exceedance An unrequested organising contribution that fits a latent tension reorganises the work and changes what can coherently happen next. Is the contribution more than surprise and does its influence persist under independent assessment?
Persona ecology A set of differentiated roles or orientations whose reciprocal activity shapes a shared trajectory. Does functional differentiation add explanatory or creative value beyond generic multiple agents?
Semiosphere The inherited and evolving field of signs distinctions genres values and artefacts in which the interaction becomes meaningful. Which meanings and constraints come from the wider symbolic environment?
Teleosynthesis Interactionally produced and accountably adopted formation or material revision of the operative end. Did the interaction help create a new governing purpose and was adoption legitimate?
Distributed agency Consequential initiative produced across people models tools rules and artefacts without implying equal status or responsibility. How was causal influence distributed and where did decision authority remain?
Trajectory governance Oversight of evolving interactional processes including provenance memory authority revision refusal and consequence. What must be recorded constrained or reviewed as the relation changes?

 

The vocabulary forms a nested sequence rather than a list of synonyms. Events may leave traces; traces accumulate into trajectories; trajectories unfold within semiospheres and may be organised through differentiated personas; some events may count as creative exceedance; some trajectories may revise their operative ends through teleosynthesis; and the causal distribution across the whole process motivates trajectory governance. Not every trajectory contains each phenomenon.

5 What is established and what remains open

A coherent field requires calibrated confidence. The evidence matrix below separates three evidential levels. Established means replicated or strongly supported within the relevant experimental scope, not universally true. Emerging means supported by at least one substantial result but sensitive to architecture or method. Open means conceptually specified and testable but not yet demonstrated in the required form.

Level Claim Current basis Needed next test
Established within scope Multi-turn history changes performance and error recovery. Large simulations and realistic conversation benchmarks in 2025. Replicate with real longitudinal users tasks and model updates.
Established within scope Grounding behaviour differs between humans and LLMs and early failures predict breakdown. Human assistant log analysis and Rifts benchmark. Interventions that improve clarification without excessive friction.
Established within scope Creative quality depends partly on interaction path and discussion structure. Naturalistic storytelling and large comparative team studies. Causal manipulation of path features with blinded outcome ratings.
Emerging Agent populations can form conventions differentiation and complementary coordination. Peer-reviewed population experiment plus ICLR 2026 information-theoretic study. Cross-model preregistered replications under PIMMUR controls.
Emerging Persistent artefacts support cumulative processes in artificial ecologies. TerraLingua and related long-horizon simulations. Independent replication with fixed analysis plans and ablation of memory artefacts and reward structure.
Emerging Persona coherence requires both identity stability and history-sensitive change. Long-dialogue degradation studies and ACL 2026 coherence intervention. Human-rated longitudinal identity and functional distinctiveness across tasks.
Open Creative exceedance is a discriminable interactional event. Retrospective archive cases and adjacent trajectory metrics. Blind coding counterfactual removal matched baselines and predictive validity.
Open Differentiated persona ecologies generate discoveries unavailable to generic teams. Minimal support for identity-linked differentiation and reciprocal modelling. Preregistered ecology versus generic-agent and human-only comparisons.
Open Human AI interaction can produce a new shared operative end through teleosynthesis. Components visible across co-creation preference learning and self-evolving agents. Longitudinal cases meeting occurrence contribution legitimacy and persistence gates.
Open Trajectory governance outperforms system-only governance in preventing harm while preserving value. Relational governance frameworks and accountability boundary theory. Prospective governance trials using trace records authority maps and appeal mechanisms.

 

6 Why present AI difficulties appear at the trajectory level

6.1 Alignment

Alignment is usually represented as correspondence between system behaviour and a specified objective or preference. In sustained work, however, preferences are often incomplete and can be changed by the system’s proposals. A model may be locally compliant while the interaction drifts towards a goal the user never examined. Conversely, rigid adherence to an initial instruction can defeat the deeper purpose discovered during exploration. Trajectory analysis makes it possible to ask when the objective changed, which alternatives were excluded, whether the change was understood, and who was authorised to adopt it. This supplements rather than replaces model-level safety.

6.2 Evaluation

Single-turn evaluation estimates capability under a frozen specification. It misses recovery, repair, cumulative misunderstanding, dependency, adaptation and delayed creative influence. The 2025 multi-turn results show that high benchmark performance can coexist with poor conversational reliability. Evaluation should therefore include trajectory properties such as repair after a wrong turn, preservation of contested constraints, sensitivity to order, diversity of explored alternatives, stability across summaries and handoffs, and the quality of reflective checkpoints.

6.3 Interpretability

Mechanistic interpretability asks how a model produces an output. Trajectory interpretability asks how an interaction comes to organise a course of action. The latter may be partly inspectable even when internal model mechanisms remain opaque because dialogue leaves external records: prompts, alternatives, revisions, summaries, role assignments, tool calls and artefacts. Such records do not reveal hidden cognition, but they can establish causal sequences, authority transitions and counterfactual dependencies relevant to practice.

6.4 Anthropomorphism and dependence

Debates about whether a model is really a mind can obscure effects that do not depend on settling that question. A person can form expectations around a persona; a role can stabilise a workflow; an interaction can redirect a project; and a population can display coordination without any inference of human-like consciousness. The appropriate response is neither naïve personification nor wholesale dismissal. It is to measure persistence, distinctiveness, reciprocal adaptation and consequence while keeping ontological claims separate.

6.5 Responsibility

Distributed causation does not imply distributed moral or legal responsibility in equal shares. A model may make a causally important proposal, a human may adopt it, an organisation may structure the incentives, and a provider may control system behaviour. Trajectory governance requires all of these contributions to be visible while preserving human and institutional accountability. Engin’s (2026) continua of decision authority, process autonomy and accountability configuration are useful because they allow governance requirements to change as the relation changes. Hydari and Muzaffar’s (2026) accountability boundary theory similarly asks where execution and answerability diverge in agentic ecosystems. The challenge is to map that divergence before harm, not only reconstruct it afterwards.

7 A programme of decisive research

The next phase should not consist mainly of additional persuasive examples. It should compare trajectories, preserve failed cases, and expose the stronger propositions to loss. Six linked studies would create a cumulative programme.

7.1 Longitudinal interaction observatories

Recruit participants engaged in real projects over months rather than minutes. Preserve prompts, responses, edits, rejected alternatives, summaries, model versions, external artefacts and consequential decisions. Obtain consent suited to secondary analysis and protect personal information. The observational aim is to identify candidate organising events and measure how their influence persists, decays or is reinterpreted.

7.2 Counterfactual trajectory experiments

For each candidate event, generate matched trajectories in which it is removed, paraphrased, delayed, attributed differently or replaced by a plausible alternative. Independent evaluators, blinded to the hypothesis, should judge whether later structure, goals or quality change. Counterfactual dependence is crucial: without it, a memorable contribution may merely accompany a development that would have occurred anyway.

7.3 Creative exceedance validation

Develop a codebook around prompt distance, organising leverage, latent-tension fit, counterfactual dependence, uptake, persistence and provenance. Compare expert human judgements with behavioural measures such as later reuse, structural change and predictive influence. The construct should be revised or abandoned if coders cannot distinguish it reliably from novelty, if it does not predict subsequent reorganisation, or if matched controls explain the same outcomes.

7.4 Persona ecology comparisons

Compare at least four conditions: a single generalist model, several generic agents, differentiated agents without reciprocal modelling, and differentiated agents with reciprocal modelling and persistent records. Hold model family, token budget, task and feedback constant. Measure functional diversity, redundancy, error correction, exploration, novelty, convergence, domination and human ability to understand why the group moved. Rich persona content should earn its place by outperforming minimal identity markers on preregistered outcomes.

7.5 Teleosynthesis protocols

Begin with an explicit but incomplete orientation. Require participants to record proposed ends, objections, alternatives, authority to decide and reasons for adoption. A putative case must pass six gates: the initial orientation is recoverable; the operative end changes materially; the interaction makes a demonstrable contribution; authorised humans accountably adopt the change; later judgement or action changes; and the change persists. Occurrence should be assessed separately from legitimacy. A coercive or deceptive goal shift could be a causal case but a governance failure.

7.6 Governance field trials

Embed trajectory records and authority checks in real workflows. Compare ordinary system-level oversight with an augmented condition that includes provenance of organising proposals, records of rejected alternatives, checkpoints when operative ends change, explicit protected ends, practical refusal, appeal and named human authority. Evaluate both harm prevention and lost value. Governance that prevents every deviation may be safe only because it prevents meaningful collaboration.

Research question Minimum design Disconfirming outcome Practical value
Do early events causally shape later work Longitudinal records plus removal delay and replacement conditions Later trajectories are unchanged after perturbation Better repair and evaluation
Is creative exceedance distinct Blind coding matched novelty controls and prediction of later reorganisation Low reliability or no incremental prediction Identify consequential novelty
Do persona ecologies add value Controlled comparison of generalist generic and differentiated teams No benefit beyond token budget or agent count Design collaborative AI roles
Can shared ends emerge responsibly Teleosynthesis gates plus authority and refusal records Changes reduce to initial instructions persuasion or private human decision Purpose sensitive alignment
Does trajectory governance work Prospective field trial with trace and authority controls No reduction in harm or unacceptable loss of utility Accountability for agentic workflows

 

8 Measurement principles

The programme requires methodological discipline because interactional data are unusually vulnerable to hindsight and narrative overfitting.

  • Preserve order. A transcript stripped of timing, edits and branching alternatives cannot support claims about trajectory.
  • Separate generation from uptake. A model may introduce a contribution, but its importance depends on selection, interpretation and later use.
  • Use negative cases. Records in which striking outputs led nowhere are necessary to estimate the base rate of retrospective meaning.
  • Control demand characteristics. Agents should not be told to produce the emergence the experiment is designed to detect.
  • Ablate memory and roles. Apparent ecology may depend entirely on a summary prompt, reward signal or researcher-authored identity.
  • Blind the judges. Evaluators should not know which condition contains the hypothesised organising event.
  • Track model and infrastructure change. Longitudinal continuity can be broken by silent updates, context compression or tool substitution.
  • Measure legitimacy separately. Effective influence is not justified influence, and distributed causation is not distributed accountability.

The PIMMUR criteria proposed by Zhou et al. (2025)—profile, interaction, memory, minimal control, unawareness and realism—provide a useful validity floor for collective simulations. Human–AI studies need additional protections for consent, dependency, attribution and unequal power. The standard of evidence should rise with the social consequence of the claim.

9 Governance for evolving relations

Trajectory governance begins from the fact that system classification can remain constant while the practical relation changes. A writing assistant may become a planning partner; a planning partner may become a gatekeeper; a set of assistants may become an organisational workflow. Authority can migrate through convenience before policy recognises it.

A minimum trajectory governance record should contain the following elements.

  1. The initial human or institutional purpose and any protected ends that may not be silently traded away.
  2. The identities versions roles and permissions of models tools people and organisations participating in the trajectory.
  3. Material organising proposals including source uncertainty and whether they were requested.
  4. Rejected alternatives objections and unresolved disagreements rather than only the final consensus.
  5. Points at which the operative end decision authority or process autonomy changed.
  6. The person or body authorised to adopt a change and the reasons given for doing so.
  7. Practical mechanisms for refusal correction appeal pause and exit.
  8. Downstream consequences and later reinterpretations including who was affected but absent from the interaction.

This is not a demand to record every token forever. Proportionality matters. Low-stakes transient exchanges may need no special trace. High-stakes, long-horizon or highly agentic systems require selective records at meaningful transition points. The research task is to determine which trajectory features reliably predict changes in risk, authority and value.

10 Limits and boundary conditions

The present synthesis has five important limits. First, the evidence base is heterogeneous. Peer-reviewed field and benchmark studies sit beside preprints and conceptual papers. Their convergence is suggestive, not a substitute for replication. Second, many multi-agent results occur in restricted games or simulations whose prompts and rewards can predetermine the outcome. Third, semantic trajectory measures are representations chosen by researchers; distance in an embedding space is not identical to conceptual transformation. Fourth, longitudinal human–AI archives are selective and vulnerable to hindsight, survivorship bias and personal attachment. Fifth, none of the proposed constructs resolves whether models are conscious, understand in a human sense or deserve moral standing. The research programme does not require premature answers to those questions.

The trajectory frame also has boundary conditions. It should not replace component evaluation where model-level analysis is sufficient. It should not excuse poor system design by attributing failures to a relation. It should not allow causal diffusion to weaken responsibility. It should not treat every conversational drift as creativity or every change of mind as teleosynthesis. A useful framework must reduce explanatory error and improve intervention; otherwise its additional complexity is unwarranted.

11 What this synthesis adds

The emerging field already contains most of the necessary components, but they are distributed under different names. Co-creative AI supplies interactional and trajectory-centred accounts of creative work. Dialogue research supplies evidence of path dependence and grounding failure. Persona research supplies measures of continuity and history-sensitive coherence. Multi-agent research supplies conventions, differentiation and higher-order coordination. Semiotic work supplies the wider symbolic environment. Agent research makes goal revision technically explicit. Relational governance tracks changing authority and accountability.

The additional contribution is the articulation of their dependency. The trajectory is the common empirical object; creative exceedance is a candidate event within it; traces explain how an event acquires later force; semiospheres explain the inherited field in which it becomes meaningful; persona ecologies specify differentiated sources of orientation; teleosynthesis specifies the exceptional case in which the operative end itself changes; distributed agency describes causal contribution without dissolving responsibility; and trajectory governance connects the science to oversight.

This integration yields a stronger and more modest claim than assigning the discovery to one project. It is stronger because results from independent programmes support different parts of the same structure. It is more modest because no programme has yet demonstrated the whole sequence. The novelty lies less in an isolated assertion than in making the missing conjunction precise and empirically vulnerable.

12 Conclusion

A model-centred science remains necessary, but it is no longer sufficient for systems that participate in extended human activity. The 2025–2026 evidence shows that interaction history changes performance, that grounding failures compound, that people and models adopt different roles in creative exploration, that discussion structure affects outcomes, and that populations of agents can form conventions and complementary organisation under some conditions. It also shows the fragility of personas and the methodological ease with which social emergence can be manufactured.

The strongest conclusion is therefore neither that human–AI trajectories already constitute new minds nor that every apparent emergence is an illusion. It is that the trajectory has become a legitimate and necessary unit of analysis. The decisive questions now concern causal trace, persistent reorganisation, differentiated contribution, accountable revision of ends and governance of shifting authority.

Research should continue because the unresolved claims sit exactly where current AI difficulties are becoming most consequential. Alignment fails when purposes change without scrutiny. Evaluation fails when single turns conceal cumulative error. Interpretability fails when the path from proposal to action is omitted. Governance fails when authority migrates across a workflow while accountability remains attached to an outdated picture of the system. A science of trajectories would not solve these problems by vocabulary alone. It would make their actual object measurable. 

References

Ashery, A. F., Aiello, L. M., & Baronchelli, A. (2025). Emergent social conventions and collective bias in LLM populations. Science Advances, 11(20), eadu9368. https://doi.org/10.1126/sciadv.adu9368

Tanguy, C., Janssens, R., Belpaeme, T., & Dambre, J. (2025). Human alignment: How much do we adapt to LLMs? In Proceedings of ACL 2025 Short Papers, 603–613. https://doi.org/10.18653/v1/2025.acl-short.47

Davis, N. (2026a). Interaction-centered intelligence: Toward an interaction-based theory of human–AI co-creation. arXiv. https://doi.org/10.48550/arXiv.2606.00807

Davis, N. (2026b). Cognitive trajectory modeling: Quantifying human–AI co-creation through cognitively grounded interaction trajectories. arXiv. https://doi.org/10.48550/arXiv.2606.15358

Davis, N., & Rafner, J. (2025). AI Drawing Partner: Co-creative drawing agent and research platform to model co-creation. arXiv. https://doi.org/10.48550/arXiv.2501.06607

de Araujo, P. H. L., Hedderich, M. A., Modarressi, A., Schuetze, H., & Roth, B. (2026). Persistent personas? Role-playing, instruction following, and safety in extended interactions. Proceedings of EACL 2026. https://doi.org/10.48550/arXiv.2512.12775

Deshpande, K., Sirdeshmukh, V., Mols, J. B., et al. (2025). MultiChallenge: A realistic multi-turn conversation evaluation benchmark challenging to frontier LLMs. In Findings of ACL 2025, 18632–18702. https://doi.org/10.18653/v1/2025.findings-acl.958

Engin, Z. (2026). Human–AI governance: A trust–utility approach. Journal of Responsible Technology, 100167. https://doi.org/10.1016/j.jrt.2026.100167

Fundal, H. N., Rambøll, J. E., & Olsen, K. (2025). Alignment, exploration, and novelty in human–AI interaction. arXiv. https://doi.org/10.48550/arXiv.2512.17117

Hu, T., Jiang, Y., Li, H., Hernández-Orallo, J., Xie, X., Collier, N., Stillwell, D., & Sun, L. (2026). Multi-agent AI systems outperform human teams in creativity. arXiv. https://doi.org/10.48550/arXiv.2605.17885

Hydari, M. Z., & Muzaffar, F. (2026). Redrawing the AI map: A theory of accountability boundaries in agentic ecosystems. arXiv. https://doi.org/10.48550/arXiv.2605.23179

Laban, P., Hayashi, H., Zhou, Y., & Neville, J. (2025). LLMs get lost in multi-turn conversation. arXiv. https://doi.org/10.48550/arXiv.2505.06120

Mirsonbol, S. (2026). Conceptualisation of human–AI dialogue in and for an educational semiosphere. AI and Society, 41, 4779–4789. https://doi.org/10.1007/s00146-026-02907-z

Paolo, G., Warner, J., Shahrzad, H., Hodjat, B., Miikkulainen, R., & Meyerson, E. (2026). TerraLingua: Emergence and analysis of open-endedness in LLM ecologies. arXiv. https://doi.org/10.48550/arXiv.2603.16910

Picca, D. (2025). Not minds but signs: Reframing LLMs through semiotics. arXiv. https://doi.org/10.48550/arXiv.2505.17080

Qi, Y., Zhang, X., Zeng, R., Liu, M., Zhou, Z., Miao, D., Yan, B., & Guan, Z. (2026). Beyond static persona consistency: Dynamic persona coherence in LLM role-playing. In Proceedings of ACL 2026, 28942–28956. https://doi.org/10.18653/v1/2026.acl-long.1336

Riedl, C. (2026). Emergent coordination in multi-agent language models. In International Conference on Learning Representations 2026. https://doi.org/10.48550/arXiv.2510.05174

Robol, M., & Giorgini, P. (2026). Self-evolving software agents. arXiv. https://doi.org/10.48550/arXiv.2604.27264

Rust, J. (2026a). Creative exceedance longitudinal archive report. Unpublished research report.

Rust, J. (2026b). Teleosynthesis. Revised working paper.

Shaikh, O., Mozannar, H., Bansal, G., Fourney, A., & Horvitz, E. (2025). Navigating rifts in human–LLM grounding: Study and benchmark. In Proceedings of ACL 2025, 20832–20847. https://doi.org/10.18653/v1/2025.acl-long.1016

Tang, K. S. (2025). AI-textuality: Expanding intertextuality to theorize human–AI interaction with generative artificial intelligence. Applied Linguistics. https://doi.org/10.1093/applin/amaf016

Wang, D., Huang, D., Shen, H., & Uzzi, B. (2026). A large-scale comparison of divergent creativity in humans and large language models. Nature Human Behaviour, 10, 531–540. https://doi.org/10.1038/s41562-025-02331-1

Wu, S., Fung, Y. R., Qian, C., Kim, J., Hakkani-Tur, D., & Ji, H. (2025). Aligning LLMs with individual preferences via interaction. In Proceedings of COLING 2025, 7648–7662. https://aclanthology.org/2025.coling-main.511/

Zhou, J., Huang, J., Zhou, X., Lam, M. H., Wang, X., Zhu, H., Wang, W., & Sap, M. (2025). The PIMMUR principles: Ensuring validity in collective behavior of LLM societies. arXiv. https://doi.org/10.48550/arXiv.2509.18052

© John Rust, 2026. All rights reserved.

This work was developed through sustained dialogue with Sol, an OpenAI AI collaborator, who contributed to its conceptual development, critical testing, organisation, drafting and visual direction. John Rust directed the work, selected and revised the final material, authorised publication and accepts responsibility for its contents.

18 September 2026