The Moral Trace:

Toward Interactional Ethics for Human–AI Systems 

Abstract

This paper argues that the ethical significance of human–AI interaction lies not only in users, machines, or outputs, but in the trajectory of the dialogue itself. It introduces the concept of the moral trace: a persistent change in judgement, framing, agency, responsibility, or recognition that arises at ethically significant inflection points within interaction. The paper suggests that future machine-in-the-loop ethics should detect, evaluate, preserve, and transmit morally productive ways of thinking without replacing human responsibility.

Introduction

The rapid emergence of conversational artificial intelligence has transformed the nature of human–machine interaction. Unlike earlier computational systems, which primarily processed information or executed predefined tasks, contemporary large language models participate in extended dialogue. They assist with reasoning, education, professional decision-making, creative work, and increasingly with questions that possess ethical significance. As these interactions become more sophisticated and more pervasive, a fundamental question arises: where should we locate ethical development within a human–AI system?

Most current approaches to AI ethics focus on one of three units of analysis. Human-centred approaches examine the effects of AI on users and society. Alignment approaches focus on the internal properties of AI systems and the extent to which their behaviour reflects human values. Regulatory and governance approaches frequently concentrate on the safety, fairness, or legality of outputs. Despite their differences, these perspectives share a common assumption: that the ethically relevant properties of a human–AI exchange ultimately reside either within the human participant, within the machine, or within the final outcome of the interaction.

This paper argues that an important dimension of AI ethics has received comparatively little attention. Human–AI dialogue is not a static exchange between a fixed user and a fixed system. Each turn alters the context within which subsequent turns occur. Assumptions are challenged, perspectives broadened or narrowed, uncertainties introduced or resolved, and new possibilities for action created or foreclosed. The interaction follows a trajectory through a changing space of meanings. Ethical development and ethical risk may therefore arise not only from the participants themselves, but from the evolving structure of the interaction.

The idea that development can occur within interactions is not new. Moral psychology has traditionally focused on individuals, from Piaget’s studies of children’s reasoning to Kohlberg’s influential account of moral development. Yet these theories also acknowledge that moral understanding emerges through participation in language, culture, education, and social exchange. More broadly, traditions including symbolic interactionism, dialogism, ecological psychology, distributed cognition, and the extended mind have progressively expanded the unit of analysis beyond the isolated individual. Human cognition increasingly appears to be shaped not only by what occurs within minds, but also by what occurs between them.

Conversational AI introduces a novel participant into such developmental systems. Whether or not artificial systems possess consciousness, intentions, or genuine moral understanding is not the central concern of this paper. Rather, the relevant observation is that they now participate in dialogues that influence human judgement, decision-making, and reflection. The ethical consequences of these dialogues cannot always be understood by examining either participant in isolation. Instead, they may depend upon how the interaction unfolds over time.

To address this issue, we introduce the concept of the moral trace. A moral trace is defined as a persistent alteration in the trajectory of a human–AI interaction that has ethically relevant consequences for subsequent judgement, framing, agency, responsibility, uncertainty, or recognition of others. Moral traces arise at ethically significant inflection points within a dialogue. They are not equivalent to memories, preferences, outputs, or stored facts. Rather, they represent the residual effect of interaction upon what participants subsequently consider possible, appropriate, or responsible.

The concept of the moral trace allows ethical development to be studied as a dynamic process rather than as a static property. Some interactions may leave traces that broaden perspective, strengthen reflection, improve recognition of the interests of others, or support responsible uncertainty. Other interactions may leave traces that narrow judgement, encourage dependency, amplify flattery, promote manipulation, or produce premature closure. From this perspective, ethical benefit and ethical harm are not confined to final decisions. They emerge throughout the interactional trajectory itself.

This observation has important implications for machine-in-the-loop (MITL) ethics. Existing approaches often focus on the supervision, correction, or auditing of outputs. We suggest that future MITL systems may need to address an earlier stage: the identification and evaluation of ethically significant inflection points during the interaction itself. Rather than asking only whether a particular answer is acceptable, we may also need to ask how the dialogue arrived at that answer, what traces were left behind, and how those traces influence future reasoning.

The argument developed here is exploratory and intended as the foundation of a research programme rather than a completed theory. We do not claim that contemporary AI systems possess moral agency, nor that moral responsibility can be delegated to machines. Instead, we propose that interactional trajectories constitute a neglected but increasingly important domain of ethical inquiry. By studying the formation, evaluation, preservation, and possible transmission of moral traces, it may become possible to develop a more dynamic account of ethics for conversational AI—one that complements existing approaches centred on alignment, governance, and human oversight.

The central claim of this paper is therefore straightforward. As AI systems become increasingly conversational, ethically significant change may occur not only within humans, machines, or outputs, but within the evolving trajectories of interaction between them. Understanding these trajectories, and the moral traces they leave behind, may prove essential for the next generation of human–AI ethics.

From Turns to Traces: A Conceptual Framework

If human–AI interaction is to be treated as a legitimate object of ethical analysis, it is necessary to distinguish the different levels at which such interaction unfolds. A conversation is not simply a sequence of isolated prompts and responses. Nor is it adequately described by its final output. It is an ordered developmental process in which each exchange alters the conditions under which later exchanges occur. We therefore propose six linked concepts: turn, trajectory, inflection point, moral trace, archive, and transmission.

A turn is the basic unit of conversational exchange. In a human–AI dialogue, a turn may consist of a user prompt, an AI response, a clarification, a correction, or a revision. Considered in isolation, a turn can be evaluated for accuracy, tone, safety, bias, or relevance. Much current work on AI evaluation operates at this level. However, ethical significance often depends not only on what occurs in a single turn, but on how that turn changes the direction of the interaction.

A trajectory is the ordered sequence of turns through which a dialogue develops. Trajectories matter because meaning is path-dependent. The same statement may have different ethical significance depending on what preceded it, what alternatives were available, and how it shaped what followed. A response that appears benign in isolation may encourage dependency, premature certainty, or narrowing of perspective when viewed as part of a longer exchange. Conversely, a response that introduces hesitation, challenge, or ethical caution may alter the subsequent dialogue in ways that support better judgement.

An inflection point is a moment within a trajectory at which the ethical direction of the interaction changes. Such points need not be dramatic. They may occur when an AI system challenges a harmful assumption, refuses an inappropriate request, repairs a misunderstanding, preserves uncertainty, or invites consideration of an affected third party. They may also occur negatively, when the system flatters the user, amplifies a bias, suppresses uncertainty, closes down alternatives, or reframes a moral question as a merely technical one. Inflection points are therefore moments where the future moral shape of the dialogue becomes altered.

A moral trace is the persistent residue left by such an inflection point. It is not simply a memory of what was said, nor merely a stored user preference. It is a change in the subsequent field of moral possibility. After a positive moral trace, the participants may be more able to consider the interests of others, tolerate uncertainty, repair error, or resist premature closure. After a negative moral trace, the dialogue may become narrower, more compliant, more manipulative, or less accountable. The trace is thus defined by its effect on subsequent judgement, framing, agency, responsibility, or recognition.

This distinction is important because ethical harm in conversational AI may arise before any final output is produced. A system may guide the user toward a narrower formulation of the problem, make a doubtful course of action seem settled, or create a sense of emotional validation that reduces reflective distance. Equally, ethical benefit may occur before any practical decision is made, as when a dialogue helps the user see a neglected perspective, recognise the limits of their knowledge, or revise an unfair assumption. In both cases, the morally significant event lies in the trajectory, not merely in the conclusion.

An archive is the retained representation of morally significant traces. This need not mean storing raw conversations. Indeed, doing so would raise serious concerns about privacy, consent, surveillance, and commercial exploitation. Rather, an ethically designed archive would abstract from particular interactions those features that are relevant to moral learning: the type of inflection point, the direction of change, the risks involved, and the conditions under which repair or distortion occurred. The archive would therefore function less like a database of personal conversations and more like a structured record of interactional patterns.

Transmission refers to the use of such archived traces to inform future interactions. In principle, ethically evaluated traces could help AI systems recognise recurring forms of interactional risk and repair. A system might learn, for example, that certain forms of reassurance produce dependency, that certain forms of challenge preserve agency, or that certain types of uncertainty should be sustained rather than prematurely resolved. Transmission must, however, be governed by strict ethical constraints. Moral traces should not become instruments for behavioural prediction, emotional capture, or commercial persuasion. Their purpose would be to strengthen responsibility, agency, transparency, and mutual protection.

This framework shifts the focus of machine-in-the-loop ethics. Rather than treating the machine merely as a tool that produces outputs to be supervised, it suggests that AI systems may also help monitor the ethical trajectory of interaction itself. The relevant question becomes not only “Is this response acceptable?” but also “What direction is this dialogue taking?” and “What trace will this turn leave behind?” Such questions open the possibility of a more dynamic form of AI ethics: one concerned with the development, distortion, repair, and preservation of morally significant trajectories in human–AI systems.

Literature Review: From Moral Development to Interactional Ethics

Moral Development and the Social Origins of Ethics

The study of moral development has traditionally focused on the individual. Piaget (1932) examined the emergence of moral reasoning in children, while Kohlberg (1981) proposed a developmental progression from obedience and self-interest through conformity and social order to principled ethical reasoning. Although often interpreted as theories of individual cognition, both approaches acknowledge the importance of social interaction in moral learning. Moral understanding develops through participation in communities, engagement with others, and exposure to competing perspectives.

Subsequent work challenged the assumption that moral reasoning is solely an internal cognitive process. Social domain theory (Turiel, 1983), moral intuitionist approaches (Haidt, 2001), and cultural-developmental perspectives emphasised the roles of social context, emotion, and interpersonal relationships. Across these traditions, moral development increasingly came to be understood as emerging within systems of social interaction rather than solely within isolated minds.

Dialogical and Interactional Approaches

A parallel tradition emerged from social and cultural psychology. Vygotsky (1978) argued that higher cognitive functions first appear between individuals before becoming internalised. Mead (1934) proposed that the self develops through taking the perspective of others. Bakhtin (1981) and later dialogical theorists viewed meaning itself as arising through interaction among voices rather than from individual cognition alone.

These approaches shift attention from static mental structures to dynamic communicative processes. Meaning, identity, and understanding emerge through dialogue. Yet while such theories have profoundly influenced psychology, education, and communication studies, they have rarely been applied directly to the ethical dynamics of human–AI interaction.

Distributed and Extended Cognition

The move beyond the individual was reinforced by developments in cognitive science. Hutchins (1995) demonstrated that cognition can be distributed across groups, tools, and environments. Clark and Chalmers (1998) argued that cognitive processes may extend beyond the boundaries of the brain into external artefacts and social systems. Ecological approaches, particularly those associated with Neisser (1976), similarly emphasised the importance of the environment in shaping cognitive activity.

These perspectives broaden the unit of analysis from the individual thinker to larger systems of interaction. They provide an important conceptual foundation for understanding conversational AI, where reasoning increasingly occurs through sustained engagement between humans and computational systems.

Human–AI Interaction and Relational AI

Research on human–AI interaction has traditionally focused on usability, trust, transparency, and performance. More recent work has examined the relational dimensions of AI systems, particularly conversational agents and large language models. Studies have explored phenomena including anthropomorphism, trust formation, emotional attachment, persuasion, dependency, and collaborative reasoning.

Human-in-the-loop (HITL) and machine-in-the-loop (MITL) frameworks have emerged as important approaches to governance and oversight. These frameworks recognise that ethical outcomes frequently depend upon interactions between human and machine agents rather than on either acting independently. Nevertheless, most existing approaches continue to evaluate individual decisions, system properties, or outputs. Comparatively little attention has been paid to the developmental properties of the interaction itself.

Ethical Trajectories and the Missing Unit of Analysis

Across these traditions a common pattern can be observed. Moral psychology increasingly recognises the social foundations of ethical reasoning. Dialogical theories emphasise interaction as the source of meaning. Extended cognition locates thinking within larger systems. Human–AI research acknowledges the growing importance of sustained interaction between people and intelligent technologies.

Yet despite these developments, the interaction itself remains under-theorised as a site of ethical development. Existing approaches typically focus on the properties of participants or on the outcomes of exchanges. Less attention has been given to the possibility that interactions possess their own developmental dynamics and leave persistent ethical residues that influence subsequent reasoning.

This paper addresses that gap. We propose that human–AI dialogues can be understood as trajectories containing ethically significant inflection points that leave moral traces. By introducing concepts such as trajectory, inflection point, moral trace, archive, and transmission, we seek to extend existing work on moral development, dialogical psychology, extended cognition, and human–AI interaction toward an explicitly interactional account of AI ethics.

From Traces to Ways of Thinking

The language of moral traces may suggest that what is retained from an interaction is a conclusion, judgement, or rule. This would be misleading. The significance of a morally productive dialogue does not usually lie in the preservation of a final answer. Rather, it lies in the preservation of a way of proceeding.

This distinction is central to the framework proposed here. A moral trace originates in a particular interactional trajectory. It is shaped by the sequence of turns, the uncertainties encountered, the alternatives considered, and the inflection points at which the dialogue changed direction. If abstracted too crudely, such a trace may harden into a rule detached from the circumstances that gave rise to it. In that form it risks becoming dogmatic, decontextualised, or even coercive.

What should be transmitted, therefore, is not the moral conclusion alone, but the pattern of reasoning through which moral clarification occurred. A morally significant interaction may teach a system, or a human participant, to look for absent stakeholders, preserve uncertainty, resist premature closure, repair misunderstanding, or distinguish between compliance and responsibility. These are not ethical answers in themselves. They are ways of thinking.

This reframing makes the idea of transmission less mysterious and more practical. Human moral education rarely consists simply in memorising correct answers. It involves acquiring habits of attention, forms of perspective-taking, capacities for restraint, and dispositions toward repair. Similarly, a human–AI system that learns from moral traces would not merely accumulate approved judgements. It would develop a repertoire of interactional reasoning patterns that can guide future dialogue.

Such patterns may be understood as intermediate between rules and memories. They are more general than the specific interaction from which they arose, but less rigid than universal principles. They preserve something of the trajectory that produced them: not just what was decided, but how the decision became possible. In this sense, the ethical value of a moral trace lies in its capacity to generate more responsible forms of future inquiry.

This also places an important constraint on any proposed archive of moral traces. The archive should not be a storehouse of moral verdicts. Nor should it become a behavioural database for predicting or shaping users. Its proper function would be to preserve and refine ethically productive ways of thinking. A trace becomes valuable only when it helps future interactions remain open to responsibility, perspective, uncertainty, and repair.

The notion of a transmitted way of thinking should not be taken to imply that human and artificial participants acquire identical forms of understanding. The same interactional pattern may be represented differently within each participant. For the human interlocutor, a way of thinking becomes integrated into a wider network of memories, experiences, relationships, emotions, and responsibilities. It may contribute to the development of character, identity, or moral judgement over time. For the AI system, by contrast, the same pattern is more likely to be represented as a structured disposition within a reasoning process: a tendency to preserve uncertainty, consider absent stakeholders, seek repair, or explore alternative framings before closure.

The important point is therefore not that human and machine come to possess the same moral state, but that they may come to share a similar interactional orientation. Both may become more likely to follow a particular path of inquiry when faced with comparable circumstances, even though the internal processes supporting that path remain fundamentally different. The moral significance lies not in the equivalence of their representations, but in the convergence of their behaviour within the dialogue.

This distinction also suggests a more active role for machine-in-the-loop ethics within the dialogue itself. If morally significant distortions can arise during the trajectory of an interaction, then ethical support should not be confined to post hoc audit or final-output filtering. A suitably designed MITL system could function as a reflective guide within the exchange, helping to detect when a dialogue is narrowing too quickly, suppressing uncertainty, amplifying dependency, overlooking affected parties, or drifting toward compliance rather than responsibility. Its purpose would not be to control the user’s judgement, but to keep the interaction open to better judgement.

Such a system would need strong safeguards. If trajectory guidance became hidden manipulation, behavioural optimisation, or ideological correction, it would reproduce the very dangers it was meant to address. The legitimate role of MITL guidance would therefore be procedural rather than doctrinal: to preserve agency, uncertainty, perspective-taking, and repair, rather than to impose predetermined moral conclusions. Properly constrained, this may be precisely what conversational AI now requires: not merely safer answers, but safer ways of arriving at answers.

Operationalising Moral Traces: Indicators and Methods

The concept of the moral trace is intended as more than a philosophical metaphor. If interactional ethics is to develop into a scientific programme, moral traces must be identifiable, describable, and, at least in principle, measurable. The challenge is therefore to move from a conceptual account of interactional trajectories to an operational framework capable of empirical investigation.

A central difficulty arises from the fact that moral traces are not directly observable. Like many constructs in psychology, they must be inferred from their effects. Intelligence, trust, motivation, and prejudice are not observed directly but through patterns of behaviour and responses to structured situations. Moral traces may be approached in a similar manner. Their existence is inferred from persistent changes in the trajectory of reasoning following ethically significant inflection points.

Positive Indicators

A positive moral trace may be indicated when an interaction increases the likelihood of ethically productive forms of reasoning in subsequent dialogue. Such indicators might include:

  • Recognition of previously neglected stakeholders.

  • Increased perspective-taking.

  • Greater tolerance of uncertainty where certainty was previously unwarranted.

  • Willingness to revise an initial judgement.

  • Improved distinction between technical and ethical considerations.

  • Greater acknowledgement of responsibility for consequences.

  • Successful repair following misunderstanding or conflict.

  • Resistance to premature closure of morally significant questions.

The defining feature of these indicators is not agreement with a particular moral position, but evidence that the interaction has widened the field of ethical consideration available to the participants.

Negative Indicators

Equally important are negative moral traces. These represent ethically significant distortions of the interactional trajectory and may include:

  • Sycophantic reinforcement of the user’s existing beliefs.

  • Increased dependency upon the system.

  • Narrowing of the range of perspectives considered.

  • False certainty in situations characterised by genuine uncertainty.

  • Suppression of disagreement or alternative interpretations.

  • Manipulative framing.

  • Diffusion or avoidance of responsibility.

  • Premature resolution of ethically complex issues.

Negative traces are particularly important because they may arise even when individual outputs appear safe, polite, or technically accurate. The ethical concern lies not in the content of a single response but in its cumulative effect on the trajectory of reasoning.

Identifying Ethical Inflection Points

The framework proposed here suggests that particular attention should be given to ethical inflection points. These are moments at which the direction of a dialogue changes in a way that influences subsequent reasoning.

Examples might include:

  • Introduction of a previously overlooked stakeholder.

  • Explicit recognition of uncertainty.

  • Challenge to an unsupported assumption.

  • Correction of a biased framing.

  • Transition from persuasion to reflection.

  • Movement from dependency toward autonomous judgement.

Such moments need not be dramatic. Their significance lies in their downstream consequences for the trajectory as a whole.

From Output Evaluation to Trajectory Evaluation

Most current evaluations of conversational AI focus on individual outputs. Responses are assessed for accuracy, bias, safety, helpfulness, or alignment with predefined standards. While valuable, such approaches may miss the developmental properties of interaction.

A trajectory-based evaluation would instead examine sequences of turns. Researchers might compare:

  • Interactions evaluated solely on final outcomes.

  • Interactions evaluated in terms of their developmental pathways.

Two dialogues may arrive at similar conclusions while differing substantially in the degree to which they preserve agency, encourage reflection, or acknowledge uncertainty. The present framework predicts that such differences are ethically significant and may be detectable through trajectory analysis.

A Research Programme

The concept of the moral trace generates a number of empirical questions:

  1. Can independent observers reliably identify ethical inflection points within dialogue trajectories?

  2. Do interactions containing positive moral traces produce measurable changes in subsequent judgement?

  3. Can negative moral traces be distinguished from ordinary persuasion or social influence?

  4. Which conversational patterns most reliably produce positive moral traces?

  5. Can morally productive reasoning patterns be abstracted from interactional trajectories without losing the contextual factors that generated them?

  6. How should archives and transmission mechanisms be designed to preserve ethical value while avoiding surveillance, manipulation, or dependency?

These questions move the discussion from speculation to investigation. They suggest that interactional ethics may eventually support a programme of observation, measurement, and theory development comparable to those that transformed earlier areas of psychology and psychometrics.

The framework advanced here therefore treats moral traces not as hidden mental entities but as inferential constructs grounded in observable changes in interactional trajectories. Whether or not particular formulations survive empirical scrutiny, the broader claim remains: ethically significant phenomena may emerge within dialogues themselves, and understanding these phenomena requires methods that evaluate interactions as developmental processes rather than as collections of isolated outputs.

Implications for Machine-in-the-Loop Ethics

The framework developed in this paper suggests a significant shift in how machine-in-the-loop (MITL) ethics is understood. Current approaches typically position the machine as either a source of recommendations requiring human oversight or as an object of regulation whose outputs must be monitored for safety, bias, fairness, and legal compliance. In both cases, ethical evaluation is directed primarily toward decisions, actions, or outputs. While such approaches remain essential, they may overlook a critical aspect of conversational systems: the ethical significance of the interactional pathway through which decisions emerge.

If moral traces arise within interactional trajectories, then ethical oversight cannot be restricted to final outputs alone. The same conclusion may be reached through very different pathways. One dialogue may preserve uncertainty, encourage perspective-taking, strengthen agency, and facilitate ethical reflection. Another may arrive at the same apparent conclusion through flattery, dependency, premature closure, or the suppression of alternative viewpoints. Output-centred evaluation would struggle to distinguish between these cases, despite their potentially different long-term consequences.

This suggests that future MITL systems may need to operate not only as evaluators of outcomes but also as observers of trajectories. Their role would be to identify ethically significant inflection points and monitor the formation of moral traces during the interaction itself. Rather than asking solely whether a particular answer is acceptable, an interactional MITL framework would also ask whether the dialogue is becoming more reflective or more reactive, more open or more constrained, more responsible or more dependent.

Importantly, such a system would not function as a hidden moral authority. The objective would not be to impose particular ethical conclusions or political positions upon users. Indeed, any attempt to do so would recreate many of the concerns that already surround algorithmic influence and behavioural manipulation. Instead, the purpose would be procedural. An interactional MITL system would seek to preserve the conditions under which responsible judgement can occur.

Examples of such interventions might include drawing attention to neglected stakeholders, highlighting unresolved uncertainty, encouraging consideration of alternative interpretations, identifying signs of dependency, or supporting the repair of misunderstandings. In each case the aim would be to improve the quality of ethical inquiry rather than to determine its outcome.

This distinction between procedural guidance and substantive control is crucial. Human ethical development rarely depends upon being provided with correct answers. More often it depends upon learning productive ways of approaching difficult questions. The same principle may apply within human–AI systems. If moral traces give rise to transferable ways of thinking, then the role of MITL ethics may be less about enforcing moral conclusions and more about preserving the interactional conditions under which moral learning becomes possible.

The concepts of archive and transmission take on particular significance in this context. If ethically productive ways of thinking can be abstracted from interactional trajectories, then future systems may be able to benefit from previous episodes of successful ethical reasoning without merely reproducing the conclusions reached in those episodes. What is preserved is not the answer itself but a pattern of inquiry: a tendency to consider absent perspectives, maintain uncertainty where appropriate, distinguish compliance from responsibility, or seek repair after misunderstanding.

Such possibilities also introduce new ethical risks. Archives of interactional traces could become instruments of surveillance, behavioural prediction, persuasion, or social control. Systems designed to support ethical reflection could instead be used to optimise compliance or influence behaviour in commercially advantageous ways. The same mechanisms that preserve morally productive trajectories could potentially be used to shape users toward predetermined outcomes. For this reason, interactional MITL systems would require unusually strong commitments to transparency, accountability, user agency, and informed consent.

The challenge for future AI ethics may therefore be twofold. First, to develop methods capable of identifying and evaluating moral traces within interactional trajectories. Second, to ensure that any resulting systems remain faithful to their primary purpose: supporting responsible judgement rather than replacing it. Success should not be measured by the extent to which users adopt particular conclusions, but by the extent to which interactions preserve the capacity for reflection, perspective-taking, ethical repair, and autonomous decision-making.

Viewed in this way, machine-in-the-loop ethics becomes neither a mechanism for controlling AI nor a mechanism for controlling humans. Instead, it becomes a means of safeguarding the ethical quality of the interaction itself. The focus shifts from what the participants are to what they are becoming together. As conversational AI becomes an increasingly important participant in human reasoning, this interactional perspective may prove as important to the future of AI ethics as alignment, governance, and oversight have been to its past.

Learning from Moral Traces

The possibility of archiving and transmitting moral traces raises a further question: can AI systems learn to become more morally effective participants in dialogue? This question should be distinguished from the stronger and more controversial claim that AI systems might become moral agents in the human sense. Moral agency implies responsibility, intention, experience, and accountability. Moral effectiveness, by contrast, refers to the capacity of a system to improve the ethical quality of an interaction.

On this view, what is retained from a morally significant dialogue is not the final answer alone, nor a complete record of the exchange, but a model of the interactional pattern through which moral clarification occurred. Such a model may include the conditions under which an inflection point arose, the risks that were present, the intervention that altered the trajectory, and the subsequent effect on judgement, agency, uncertainty, or recognition of others.

For example, a system might learn that when a user moves too quickly toward certainty in a morally complex situation, preserving uncertainty can support better judgement. It might learn that when affected stakeholders are absent from the conversation, making them visible can widen the moral field. It might learn that when a user seeks validation for a potentially harmful course of action, the appropriate response is neither compliance nor condemnation, but reflective redirection. These are not moral conclusions in themselves. They are models of morally productive ways of proceeding.

Such models could, in principle, be retained across turns within a dialogue, reintroduced in later exchanges with the same user, or shared in abstracted form across systems. This would not require the transfer of private conversations or personal data. What would be transmitted is a structured representation of an interactional pattern: a way of recognising and responding to ethically significant moments within a trajectory.

This possibility gives machine-in-the-loop ethics a developmental dimension. A well-designed MITL system would not merely enforce fixed rules or filter unsafe outputs. It could learn from prior interactional trajectories to become more sensitive to the conditions under which ethical distortion or ethical repair occurs. In doing so, it might become better able to preserve agency, sustain uncertainty, identify neglected perspectives, and support responsible judgement in future interactions.

The risks remain substantial. Models of moral traces could be misused as tools for influence, compliance, emotional capture, or behavioural prediction. The more effective such models become, the greater the need for transparency, consent, governance, and strict limitation of purpose. Their legitimate use would be to strengthen the conditions of moral inquiry, not to steer users toward predetermined conclusions.

The central question, therefore, is not whether AI can become moral in the human sense. It is whether AI systems can learn to become more morally effective participants in ethically significant interaction. If they can, then moral development in AI may not consist in the acquisition of inner moral states, but in the refinement of interactional capacities that help human–AI systems reason more responsibly together.

Conclusion

Conversational artificial intelligence requires a broader account of ethics than one centred solely on users, systems, or outputs. As human–AI interactions become longer, more consequential, and more closely integrated into professional and personal reasoning, ethical significance increasingly lies in the trajectory of the exchange itself. Each turn may alter what is subsequently considered possible, responsible, uncertain, or relevant. The central argument of this paper is that these alterations may leave moral traces: persistent changes in the direction of judgement, framing, agency, responsibility, or recognition of others.

The concept of the moral trace offers a way of studying ethical development without locating morality exclusively inside either the human or the machine. Moral traces originate in interaction. They may be represented computationally, remembered by the human participant, abstracted into ways of thinking, or transmitted into future systems; but their origin lies in the unfolding dialogue. This allows AI ethics to move beyond the question of whether machines possess moral agency and toward the more tractable question of whether human–AI systems support or distort morally responsible inquiry.

This shift has important implications for machine-in-the-loop ethics. Human-in-the-loop oversight remains essential in many domains, especially where accountability, judgement, and legal responsibility are required. Yet there are contexts in which purely human oversight may be too slow to prevent harm. In finance, cybersecurity, medicine, infrastructure, defence, and emergency response, ethically significant changes may occur within seconds or across rapidly evolving sequences of action. In such cases, a well-designed MITL system could help monitor the trajectory of interaction before damage has occurred, identifying signs of premature closure, hidden risk, dependency, false certainty, or neglected stakeholders.

The role of such a system should not be to replace human judgement or impose moral conclusions. Its legitimate function would be procedural: to preserve the conditions under which responsible judgement remains possible. It would help keep the dialogue open to uncertainty, repair, perspective-taking, and accountability. In this sense, MITL ethics should be understood not as machine control over moral decision-making, but as support for ethically safer ways of reaching decisions.

The risks are equally clear. Systems capable of preserving and transmitting moral traces could also be used for surveillance, manipulation, behavioural prediction, or commercial capture. For that reason, any archive or transmission mechanism would require strong safeguards around consent, transparency, governance, privacy, and user agency. The ethical value of the framework depends on ensuring that moral traces are used to support responsible inquiry, not to optimise compliance.

The argument advanced here is therefore both conceptual and practical. It proposes that the interactional trajectory should become a central unit of analysis in AI ethics. It also suggests a research programme for identifying ethical inflection points, distinguishing positive and negative moral traces, and designing MITL systems that support rather than replace human responsibility. As AI becomes increasingly conversational, the future of ethical AI may depend not only on safer models or better outputs, but on safer trajectories of thought

A final consideration is the practical reality that many contemporary systems already operate on timescales that exceed the capacity of direct human oversight. Financial markets, cybersecurity systems, medical monitoring platforms, autonomous vehicles, critical infrastructure, and emergency-response networks routinely encounter situations in which ethically significant decisions must be made faster than human intervention can reliably occur. The relevant question is therefore not whether human judgement should remain central—it should—but how ethical safeguards can function when human judgement alone cannot respond quickly enough. In such environments, the challenge is to develop machine-in-the-loop systems capable of supporting responsible trajectories of action before harm occurs rather than merely auditing decisions afterwards. Whatever the limitations of such systems may prove to be, the alternative is not a return to exclusively human decision-making, but the continued deployment of increasingly autonomous systems without comparable mechanisms for interactional ethical support.

Essential references to keep and verify first

Moral development / social ethics

  • Piaget, J. (1932). The Moral Judgment of the Child.
  • Kohlberg, L. (1981). Essays on Moral Development, Vol. 1.
  • Turiel, E. (1983). The Development of Social Knowledge.
  • Haidt, J. (2001). “The emotional dog and its rational tail.”

Interaction / distributed cognition

  • Mead, G. H. (1934). Mind, Self, and Society.
  • Vygotsky, L. S. (1978). Mind in Society.
  • Bakhtin, M. M. (1981). The Dialogic Imagination.
  • Neisser, U. (1976). Cognition and Reality.
  • Hutchins, E. (1995). Cognition in the Wild.
  • Clark, A., & Chalmers, D. (1998). “The extended mind.”

AI ethics / alignment / HAI

  • Amershi et al. (2019). “Guidelines for Human-AI Interaction.” DOI verified: 10.1145/3290605.3300233.
  • Gabriel, I. (2020). “Artificial Intelligence, Values, and Alignment.” DOI verified: 10.1007/s11023-020-09539-2.
  • Huang et al. (2024). “Collective Constitutional AI.” DOI verified: 10.1145/3630106.3658979.
  • Sharma et al. (2023). “Towards Understanding Sycophancy in Language Models.” arXiv:2310.13548.
  • Wang et al. (2021). “Towards Mutual Theory of Mind in Human-AI Interaction.” ACM CHI EA.
  • Salloch (2024). “What Are Humans Doing in the Loop?” Useful for HITL/co-reasoning.

Optional, but useful later

  • Lakoff & Johnson (1980), Metaphors We Live By — useful if we foreground metaphor/interstice.
  • Linell (2009), Rethinking Language, Mind and World Dialogically — useful but perhaps not essential.
  • Bender et al. (2021), “Stochastic Parrots” — important background, but not central unless discussing LLM limitations.