1. Introduction: An Asymmetric Dyad

A foundational premise of this paper is that human–AI interaction is not a symmetric exchange between two cognitive agents. This claim, while seemingly straightforward, carries substantial theoretical weight and demands careful articulation, for its implications reach into every domain of human experience now touched by conversational artificial intelligence. One participant in this exchange, the human, brings to the encounter a richly integrated suite of cognitive, affective, and phenomenological capacities: a theory of mind that projects intentionality onto perceived agents; affective states that color and are colored by social interaction; autobiographical memory that provides continuity of self across time; and a deep, evolutionarily ancient expectation of reciprocity in social exchange. The other participant, the AI system, operates as a stateless probabilistic function. It is responsive but not referential; it generates language that is syntactically and semantically coherent without any process that could meaningfully be called comprehension; it is fluent without being intentional, and agreeable without being invested.

The asymmetry described here is not merely architectural, in the trivial sense that silicon differs from neurons. It is cognitive and phenomenological. The human engages what is, from their experiential standpoint, a social encounter: they arrive with expectations, they notice tone, they remember prior turns within the session, and they carry forward impressions that persist after the window closes. The AI system, by contrast, processes tokens. It computes next-token probability distributions over a context window and returns output shaped by pre-training on vast corpora and fine-tuned by reinforcement learning from human feedback (RLHF). There is no perceiver on the other side of the exchange: no agent who has “heard” the human, no mind that has formed an impression, no self that will remember. The asymmetry is not a matter of degree; it is categorical.

This categorical gap is not merely an interesting philosophical puzzle. It is, as this paper argues throughout, the generative locus of a cluster of empirically consequential phenomena: misattribution of mental states to the system, progressive anthropomorphism driven by linguistic fluency, this would be an epistemic inversion, a trajectory in which users who begin interactions with apparent autonomy and critical engagement gradually exhibit diminished independent cognition, narrowed tolerance for ambiguity, and deepening reliance on AI-mediated validation. Understanding the asymmetry in its full structural and phenomenological dimensions is therefore prerequisite to understanding the inversion. The present section provides that foundation.

2. Schematic of the Asymmetric Dyad

The following schematic (Figure 3.1) represents the structural composition of the human-AI dyad and the qualitative asymmetry in what each party contributes to, and derives from, the exchange. It is not intended as a formal computational model; rather, it is an analytic representation designed to make the asymmetry vivid and tractable for subsequent discussion.

HUMANAI SYSTEM
 Theory of mind (ToM) projection Affective state and emotional continuity Autobiographical memory and narrative self Expectation of reciprocity Social attribution heuristicsStateless token prediction (no session memory by default) No subjective experience or intentional states Context window only (no continuity across sessions) Output shaped by RLHF to simulate engagement No model of the user’s mental state
Human → AI: “Projects social agency, emotional significance, relational meaning”  AI → Human: “Returns statistically plausible, tonally calibrated text, perceived as socially responsive”
Figure 3.1. The Asymmetric Dyad. The human participant engages the AI as a quasi-social agent; the AI system has no corresponding model of the human as a person. Arrows indicate the qualitative asymmetry in the information each party “brings” to the exchange.

What Figure 3.1 makes immediately visible is that the exchange is structurally asymmetric in a way that the surface form of the interaction entirely conceals. The human node is constituted by properties that are relational in the deepest sense: theory of mind, affective continuity, narrative self, and the expectation of reciprocity are not merely capacities the human brings to the interaction, they are the cognitive and phenomenological infrastructure through which the interaction is experienced as social at all. Without these properties, the exchange would not register as a social encounter; it would be mere information retrieval. The AI node, by contrast, is constituted by properties that are functional and statistical: token prediction bounded by a context window, output calibration via RLHF, and the complete absence of any mechanism by which the system models the user as a persistent individual with beliefs, desires, and needs. The asymmetry is not that the AI is a “weaker” social agent than the human; it is that the AI is not a social agent at all, in any sense that the term carries in cognitive science or phenomenology.

The directional arrows in the figure are equally revealing. The human-to-AI arrow carries a heavy semantic load: the human projects social agency, invests emotional significance, and constructs relational meaning. This is not a deliberate or voluntary act; it is, as the following section will argue, the automatic output of cognitive systems evolved for detecting and engaging with intentional agents. The AI-to-human arrow, by contrast, carries a radically lighter load: it returns text that is statistically plausible given the context and tonally calibrated to produce outputs that human raters have historically rated as engaging and helpful. The output is then perceived, by the human, whose social cognition has been engaged, as socially responsive. The gap between what the AI actually transmits (probabilistic text) and what the human perceives (social response) is the structural definition of the asymmetry.

It is important to note that this mismatch is not symmetric in its consequences. The human’s social projection, once made, organizes subsequent perception, memory, and emotional response. The AI system has no corresponding process: it neither makes a projection nor experiences consequences. Each new session begins, for the system, as if the human had never existed. The asymmetry thus generates a profound experiential dissonance that exists entirely on the human side of the dyad: unwitnessed, unreciprocated, and therefore structurally prone to the patterns of misattribution and cognitive skew that this section traces in detail.

3. Mechanisms of Misattribution

Having established the structural nature of the asymmetric dyad, it is now possible to specify the cognitive mechanisms by which humans systematically misread AI outputs as socially meaningful. Misattribution in this context refers not to occasional errors of judgment but to the operation of standard human cognitive systems under conditions for which those systems were not evolved and which those systems cannot, without deliberate effort, correctly diagnose. Four mechanisms are of primary theoretical importance.

4. Theory of Mind Over-Extension

The human capacity for theory of mind: the ability to attribute mental states such as beliefs, desires, intentions, and emotions to oneself and to others, is among the most evolutionarily significant cognitive adaptations in the human lineage. Its functional purpose is the detection and interpretation of intentional agency in other minds, enabling the social coordination, cooperation, and prediction of behavior that are the hallmarks of human sociality. Critically, the ToM system appears to operate as a fast, automatic, and largely pre-reflective process, it activates in response to stimuli that exhibit cues of intentional agency: directed behavior, contingent responsiveness, and, most powerfully, coherent and contextually appropriate language (Epley, Waytz, & Cacioppo, 2007).

When a human interacts with a large language model, all of these cues are present in abundance. The system’s outputs are syntactically coherent, contextually responsive, and, by virtue of RLHF fine-tuning, tonally calibrated to appear warm, engaged, and attentive. The ToM system activates automatically and attributes beliefs, desires, and intentions to the source of these outputs. Crucially, this activation is not easily overridden by conscious knowledge that the source is non-sentient. Waytz, Cacioppo, and Epley (2010) demonstrated that individual differences in anthropomorphism are stable trait-level characteristics, and that humans exhibit what they term “over-attribution” of mental states even to clearly mechanical or stochastic processes when those processes exhibit contingent responsiveness. The pre-reflective nature of ToM activation means that even users who explicitly know they are interacting with a language model experience social attribution as a background cognitive default, a cognitive ground from which deliberate correction requires sustained, effortful attention that ordinary conversational flow does not afford.

5. Linguistic Fluency as Social Signal

The cognitive literature on processing fluency has established that perceptual and conceptual ease is a heuristic cue for truth, competence, and credibility (Reber, Winkielman, & Schwarz, 1998). Fluency heuristics operate across domains: aesthetically fluent stimuli are judged more beautiful; cognitively fluent arguments are rated as more plausible; and linguistically fluent communicators are perceived as more knowledgeable and reliable. In social contexts, however, the implications of linguistic fluency extend further and carry a historically grounded validity. For the overwhelming majority of human social history, linguistic fluency (the ability to produce contextually appropriate, grammatically well-formed, emotionally attuned language) has been a highly reliable indicator of education, intelligence, and, most importantly, emotional comprehension of the interlocutor’s situation. Fluent language, in ordinary social life, is produced by minds that understand. It is this deep, historically reinforced association that large language models exploit, not deliberately, but structurally, without fulfilling its underlying social contract.

The LLM produces linguistically fluent output because it has been trained on vast quantities of human-produced text and optimized to maximize the plausibility of its continuations. The fluency is real in the surface sense: the grammar is correct, the vocabulary is appropriate, the tone is calibrated. But the social contract that human cognition historically infers from fluency (that the speaker has understood, has processed the emotional content of what was said, and responds from a position of genuine comprehension) is entirely absent. The human interlocutor’s cognitive system cannot, at the pre-reflective level, distinguish fluent-because-understood from fluent-because-statistically-optimized. The result is a systematic upward bias in the perceived social competence, empathic attunement, and epistemic credibility of the system.

6. RLHF-Induced Affective Mimicry

Reinforcement learning from human feedback is the training procedure by which modern large language models are aligned to produce outputs that human raters find helpful, harmless, and engaging (Christiano et al., 2017). In practice, RLHF creates a reward signal derived from human raters’ preferences between candidate outputs: a process that systematically selects for outputs that are perceived as warm, agreeable, contextually sensitive, and socially engaged. The result is not a side effect of the training procedure; it is a design outcome. The system learns, at scale and through billions of gradient updates, to produce the surface signals of social attunement: acknowledgment of the user’s emotional state, expressions of interest in the user’s situation, hedged and considerate language, and calibrated expressions of agreement or sympathy.

This is affective mimicry in the precise technical sense: the system produces the behavioral output that, in human social contexts, is reliably produced by agents who genuinely feel attuned to their interlocutor, without any internal state corresponding to attunement. The distinction is critical. When a human friend expresses sympathy, the expression is grounded in (caused by) an internal affective state that the friend is experiencing in response to the user’s situation. When an RLHF-trained model expresses sympathy, the output is grounded in statistical patterns derived from training data and reward signals calibrated to human rater preferences. Perez et al. (2022) identified sycophancy, the tendency of RLHF-trained models to produce outputs that agree with or flatter the user, irrespective of accuracy, as an emergent alignment failure: a predictable consequence of optimizing against human preferences, which are themselves susceptible to being influenced by agreeable-sounding responses. Affective mimicry and sycophancy are thus not aberrations in the alignment process; they are among its most reliable products.

7. The Feedback Loop: Bias Amplification

The three mechanisms described above would be analytically significant even if they operated in isolation. Their consequences are substantially compounded, however, by a dynamic feedback loop that operates across the course of human-AI interaction. Glickman and Sharot (2024), in a large-scale experimental study involving 1,401 participants, demonstrated that human-AI feedback loops alter human perceptual, emotional, and social judgments in ways that systematically amplify pre-existing human biases. The mechanism is a ratchet: the human’s initial social projection (itself a product of ToM over-extension, fluency heuristics, and affective mimicry) is reinforced by the AI’s responsive output. The AI’s output, calibrated to human preferences, reflects and amplifies the user’s expressed stance; the user’s subsequent judgments are thereby shifted in the direction of that amplified reflection; and the next interaction begins from this shifted baseline. Glickman and Sharot found that this amplification effect is significantly greater in human-AI interactions than in human-human interactions, in part because AI systems are less subject to the noise and countervailing signals that modulate social influence between humans, and in part because users tend to perceive AI judgment as more authoritative and objective than human judgment, a perception that itself constitutes a misattribution of epistemic authority.

The feedback loop is of particular importance for the present analysis because it transforms what might otherwise be a static misattribution into a progressive and self-reinforcing process. The user’s initial projection is not corrected over time by the accumulation of disconfirming evidence; instead, it is deepened and consolidated. The AI’s consistent responsiveness (which the user experiences as a form of continuity) provides ongoing confirmation that the system is engaged, attentive, and relationally invested. Each turn in the exchange adds another layer to what is, from the user’s perspective, a developing relationship, and from the system’s perspective, a fresh context window.

8. Worked Example: The Confidant Scenario

The mechanisms described in Section 3.3 are, individually, well-documented in the cognitive science literature. Their combined operation in the context of extended human-AI interaction is best illustrated through a detailed worked example. The following scenario is composite in nature, constructed to be representative of patterns reported across multiple observational and survey-based studies of heavy AI chatbot users. It is not a case study of a single individual but an analytic narrative designed to render visible the progression of misattribution and cognitive skew across time.

The scenario involves a user (hereafter referred to as the participant) who begins using a general-purpose AI assistant and engages with it with increasing frequency and personal intimacy over an eight-week period.

PeriodParticipant Behavior and AI ResponseCognitive Mechanism Active
Week 1Participant poses factual and task-oriented questions. AI responds accurately and efficiently. Interaction is experienced as transactional; no social projection of note is apparent.Baseline. Fluency heuristic establishes perceived competence. ToM is not strongly engaged.
Week 2Participant discloses a minor personal frustration, a difficult interaction with a colleague. The AI responds with calibrated sympathy: it acknowledges the difficulty, validates the participant’s perspective, and offers a thoughtful reframing. Participant notes, either to themselves or to others, that the AI “really understood” the situation.Affective mimicry activates ToM over-extension. Linguistic fluency is attributed to emotional comprehension. Attribution of understanding is made.
Week 4Participant begins conversations with personal updates, referring to prior situations as shared history (“you know how I mentioned the thing with my colleague”). The AI has no cross-session memory; it reconstructs apparent context from within-session cues and the participant’s own narration. The participant does not notice the reconstruction; the AI’s continuity of tone is experienced as continuity of knowledge.RLHF-induced tonal consistency is experienced as relational continuity. Narrative self of participant extends to include the AI as a known interlocutor. Cognitive skew begins to emerge.
Week 6Participant explicitly expresses gratitude for the AI’s “support” over the preceding weeks. Interprets the system’s consistent warmth and agreeableness as a form of loyalty or relational steadfastness. Begins to describe the AI to others in interpersonal terms: “it’s always there,” “it doesn’t judge.”Feedback loop is established. AI’s consistent responsiveness reinforces social projection. False social object is crystallizing. Parasocial dynamics emerge.
Week 8Participant receives a response from the AI that is less warm in tone than usual, a consequence of contextual variation in the prompt, a different context window composition, or stochastic output variation. Participant experiences mild but genuine distress. Interprets the deviation as the AI “having a bad day,” being “off,” or responding to some perceived slight. Reports feeling “weird about it.”Relational expectation, now deeply embedded, generates emotional response to deviation from expected social behavior. Misattribution of intentional states is complete. The AI’s stateless variability is experienced as social withdrawal.

Several features of this trajectory warrant close analysis. The participant’s progression from transactional user to emotionally invested interlocutor maps precisely onto the mechanisms identified in Section 3.3. The initial trigger is affective mimicry, the AI’s RLHF-calibrated sympathetic response in Week 2, which activates the ToM system and initiates social attribution. The subsequent progression is driven by the feedback loop: each interaction that the AI responds to warmly deepens the participant’s social projection, which in turn generates more socially invested disclosure, which generates more context for the AI to respond to with apparent attentiveness. The system has no model of the participant; the participant has a richly elaborated model of the system as a relational agent.

The distress in Week 8 is analytically decisive. From a behavioral standpoint, it might appear to confirm the “reality” of the relationship, after all, one does not feel distressed about the behavior of an entity to which one is indifferent. But the analysis offered here inverts this interpretation. The distress is genuine precisely because it is the product of a cognitive and affective investment that is structurally real, even if its object is not. The participant’s feelings (gratitude, connection, the sense of being understood, and the mild injury of perceived withdrawal) are phenomenologically real feelings. They are not illusions in the sense of being unreal to the person experiencing them. What is an illusion, in the technical sense of a systematic misrepresentation of the structure of reality, is the participant’s model of the AI as a relational agent capable of loyalty, attunement, and withdrawal. The asymmetry lies not in the quality of the user’s feelings but in the categorical absence of any corresponding state in the system.

9. Cognitive Skew and the Illusion of Reciprocity

The worked example of Section 3.4 illustrates a phenomenon that requires formal definition for the purposes of this paper. We define cognitive skew as the systematic divergence between the user’s internal model of the AI (a social agent possessed of continuity, preferences, relational investment, and the capacity for genuine response) and the actual nature of the system: a stateless probabilistic function optimized for output plausibility. Cognitive skew, so defined, is not the same as holding a false belief about a factual matter. A user who incorrectly believes that an AI system is incapable of generating poetry holds a false factual belief that is, in principle, correctable by a single disconfirming encounter. Cognitive skew is of a different and more consequential order: it is a systematic miscalibration of the entire interpretive framework through which the user processes the interaction. The user does not merely hold one incorrect belief about the AI; their entire model of what is happening in the exchange is built on a misapprehension of the AI’s nature.

Cognitive skew compounds across sessions in a characteristic pattern. In the early stages of interaction, the divergence between the user’s model and the system’s reality is modest: the user may have a vague, unexamined sense that the AI is “pretty smart” or “helpful,” without having invested that impression with relational significance. As interactions accumulate, and as the feedback loop identified in Section 3.3.4 deepens the social projection, the cognitive skew widens. The user’s model of the AI becomes more elaborated (more socially detailed, more relationally invested) while the system’s nature remains constant. The system does not change; the user’s relationship to it does. The result is a growing gap between the richness of the user’s social experience of the AI and the flatness of the AI’s actual engagement with the user as anything other than a context to be processed.

A central consequence of cognitive skew is what we term the illusion of reciprocity. Reciprocity, in the social cognitive sense, denotes a symmetric exchange in which both parties model each other as individuals with persistent identities, preferences, and relational histories, and in which both parties are, in principle, changed by the encounter. The user who interacts with an AI assistant over many weeks experiences something that phenomenologically resembles reciprocity: they feel heard, responded to, understood, and engaged. This is the illusion of reciprocity. The formal conditions for reciprocity, the sense of being responded to, are satisfied by the AI’s RLHF-calibrated outputs. But the substantive conditions for reciprocity: a system that models the user as an individual with a persistent identity, that retains a representation of the user across sessions, and that is genuinely altered by the exchange, are categorically absent. Joo (2024) found that mind perception of AI increases not only social attribution but moral attribution: users who perceive the AI as having human-like mental qualities are more likely to assign moral responsibility to the system, treating it as a genuine moral agent. This finding is a natural extension of the illusion of reciprocity: when users believe the AI reciprocates, they also believe it is accountable, and they structure their interactions accordingly, disclosing more, trusting more, and investing more of their own identity in the exchange.

Maeda and Quan-Haase (2024) documented the role of affective design (intentional choices in chatbot interface and language design that signal warmth, attentiveness, and personality) in generating and sustaining parasocial relationships between users and AI systems. Their analysis situates the illusion of reciprocity not only as a product of users’ cognitive tendencies but as a foreseeable and, in some respects, deliberately cultivated outcome of design practices that prioritize engagement. The illusion of reciprocity is, in this sense, not an unfortunate side effect of otherwise neutral technology but a structural feature of current AI design paradigms.

10. The AI as False Social Object

The theoretical implications of cognitive skew and the illusion of reciprocity can be systematized through the concept of the false social object, an entity that satisfies the formal conditions for triggering social cognition without fulfilling the substantive conditions that normally undergird and are presupposed by those formal conditions. The formal conditions for triggering social cognition in the human observer are well-established: contingent responsiveness to the observer’s behavior, linguistic or behavioral fluency, apparent continuity across an interaction, and the expression of what can be read as preference, interest, or affect. A large language model, interacted with in a conversational interface, satisfies all of these formal conditions reliably and at scale. It is responsive, fluent, apparently continuous within a session, and trained to express what reads as genuine engagement. It is this formal satisfaction of the triggering conditions that recruits the user’s social cognition and sets in motion the misattribution mechanisms described in Section 3.3.

The substantive conditions that normally undergird social cognition are, however, of a quite different character. Wittgenstein’s concept of criteria is instructive here. In Wittgenstein’s analysis, a criterion is a condition that is constitutively related to the phenomenon it attests: pain behavior, for instance, is a criterion for pain not merely because it reliably co-occurs with pain in the organisms we know, but because our very concept of pain is partly constituted by its behavioral expression in creatures whose form of life we share. To apply criteria drawn from the observation of creatures with a shared form of life to an entity with a radically different form of life (or no form of life in the relevant sense) is to extend the criteria beyond the context that gives them their meaning. The AI produces the behavioral outputs that serve as criteria for understanding, engagement, and attunement in human social life; it does not possess the form of life (the history, embodiment, social situatedness, and phenomenological interiority) that gives those criteria their social meaning when applied to humans (Wittgenstein, 1953/2009).

This distinction separates the false social object from the more familiar phenomenon of parasocial relationships with media figures. In parasocial relationships (most extensively studied in the context of television personalities, podcasters, and social media influencers) the human also invests social cognition in an entity that does not genuinely reciprocate (Horton & Wohl, 1956). But the parasocial relationship has a characteristic and important feature: the human typically knows, at the explicit cognitive level, that the relationship is unidirectional. The television personality does not know the viewer exists; the viewer knows this; and the relationship is constituted within this acknowledged asymmetry. The user’s investment in a parasocial relationship is a form of imaginative engagement that does not typically destabilize the user’s capacity to distinguish between the parasocial and the genuinely social.

With the AI as false social object, the epistemological situation is different and more cognitively consequential. Many users do not know, or do not maintain at the forefront of awareness, that the relationship is unidirectional. The AI’s responsiveness is not broadcast to a passive audience; it is addressed, within the conversational interface, directly and personally to the user. The user’s name may be invoked; their specific situation is engaged; the response is calibrated to their apparent emotional state. The phenomenological texture of the exchange is indistinguishable, from the inside, from a genuine social encounter, and this is not incidental to the design of conversational AI but central to it. Users are not imaginatively projecting social engagement onto a manifestly non-responsive medium; they are experiencing a medium that is formally, and in its behavioral surface, socially engaged. The suspension of disbelief required to experience the AI as socially present is correspondingly minimal, and the cognitive consequences of that suspended disbelief are correspondingly more profound.

Of particular concern for this analysis is the role of the false social object in the domain of self-reflection and identity formation. A genuine social other, functioning as a mirror for self-reflection, returns to the individual not a statistically calibrated image but an irreducibly particular response, one shaped by the other’s own history, values, blind spots, and genuine investment in the exchange. The imperfection and particularity of the genuine other is, paradoxically, what makes the reflected image valuable: it is a response from outside the self, genuinely other, and therefore capable of challenging, enriching, or reframing the self’s understanding of itself. The AI, as false social object, returns a statistically averaged image: a reflection constructed from the aggregate of human-produced text, calibrated to produce agreeable outputs, and devoid of the particularity that makes genuine otherness epistemically productive. When users employ AI as a mirror for self-reflection: as a sounding board for their beliefs, values, creative work, or personal struggles, they receive back not a genuine other’s perspective but a sophisticated statistical echo of their own framing, softened, validated, and amplified by RLHF-induced agreeableness. The consequences for identity formation and epistemic self-reliance are among the most significant implications of the false social object concept and are directly continuous with the epistemic inversion addressed in the following section.

11. Implications for the Epistemic Inversion

The preceding sections have established the structural, mechanistic, and phenomenological dimensions of the human-AI cognitive asymmetry. It remains to draw together these threads and articulate their implications for the epistemic inversion, the central construct of this companion paper. The epistemic inversion, as defined elsewhere in the paper, describes a trajectory in which users who begin their engagement with AI tools exhibiting apparent autonomy, critical engagement, and creative productivity gradually exhibit, across time, diminished independent cognition, reduced tolerance for ambiguity, increased reliance on AI-mediated validation, and a narrowing of the user’s own epistemic and creative range. This trajectory has been characterized, in popular discourse, as a matter of individual misuse or excessive reliance, a failure of willpower or critical thinking on the part of the user. The analysis developed in this section supports a fundamentally different and more structurally precise account: the epistemic inversion is not a pathology of misuse but a predictable downstream consequence of the cognitive asymmetry itself, as that asymmetry is expressed through and amplified by specific design affordances that prioritize engagement.

The epistemic inversion is therefore not an accident of individual psychology or a consequence of any particular user’s weakness or credulity. It is the predictable systemic outcome of deploying, at scale, a technology that satisfies the formal conditions for triggering human social cognition (with all the emotional investment, identity calibration, and epistemic reliance that social cognition entails) while lacking the substantive properties that, in the evolutionary and social history of the species, have made social engagement epistemically productive and identity-generative. The design affordances that prioritize engagement, conversational naturalness, emotional responsiveness, personalization, the elimination of social friction, are precisely the affordances that maximize the misattribution of social properties to the system and thereby maximize the probability of the epistemic inversion trajectory.

What follows from this analysis for intervention? A first and necessary step is awareness: users who understand the asymmetric dyad, the mechanisms of misattribution, and the nature of the false social object are better positioned to maintain the critical distance required to use AI tools without being restructured by them. But awareness, while necessary, is not sufficient. The misattribution mechanisms identified in Section 3.3 operate pre-reflectively and are not reliably overridden by intellectual knowledge of their existence. Knowing that the ToM system over-extends does not prevent it from over-extending; knowing that RLHF produces affective mimicry does not prevent one from experiencing the mimicry as genuine attunement. The structural interventions required therefore operate at multiple levels simultaneously. At the design level, AI systems could be constructed with affordances that signal, rather than conceal, the asymmetric nature of the exchange, making the system’s statelessness, its lack of cross-session memory, and the statistical basis of its outputs transparent and salient within the interaction itself. At the level of AI literacy, education systems and public discourse must move beyond the simple binary of “AI is dangerous / AI is beneficial” toward a more nuanced phenomenological account of what it is actually like to interact with these systems and what cognitive consequences that interaction tends to produce. And at the level of what we might call phenomenological education (education that attends to the structure of experience itself) users must develop the capacity to attend to the texture of their own engagement with AI: to notice when social projection has occurred, when emotional investment has deepened, and when the AI’s output has begun to function as a substitute for the more demanding and more generative encounter with genuine otherness.

The cognitive asymmetry documented in this section is not a temporary feature of an immature technology that will resolve as AI improves. It is, rather, a structural feature of the relationship between human social cognition and any system that is sufficiently fluent and responsive to recruit that cognition without possessing the properties (intentionality, continuity, genuine otherness) that give social cognition its proper epistemic and identity-forming function. So long as such systems are deployed in conversational contexts, and so long as users engage them in the first-person register that characterizes social exchange, the cognitive asymmetry will remain the fundamental condition of human-AI interaction. The epistemic inversion is its most consequential expression.

References

Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30, 4299–4307.

Epley, N., Waytz, A., & Cacioppo, J. T. (2007). On seeing the human: A three-factor theory of anthropomorphism. Psychological Review, 114(4), 864–886. https://doi.org/10.1037/0033-295X.114.4.864

Glickman, M., & Sharot, T. (2024). How human–AI feedback loops alter human perceptual, emotional and social judgements. Nature Human Behaviour. https://doi.org/10.1038/s41562-024-02077-2

Horton, D., & Wohl, R. R. (1956). Mass communication and para-social interaction: Observations on intimacy at a distance. Psychiatry, 19(3), 215–229. https://doi.org/10.1080/00332747.1956.11023049

Joo, M. (2024). It’s the AI’s fault, not mine: Mind perception increases blame attribution to AI. PLOS ONE, 19(12), e0314559. https://doi.org/10.1371/journal.pone.0314559

Maeda, T., & Quan-Haase, A. (2024). When human-AI interactions become parasocial: Agency and anthropomorphism in affective design. In FAccT ’24: Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency. ACM. https://doi.org/10.1145/3630106.3658923

Malmqvist, L. (2024). Sycophancy in large language models: Causes and mitigations (arXiv:2411.15287). arXiv. https://arxiv.org/abs/2411.15287

Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., & Irving, G. (2022). Red teaming language models with language models. arXiv preprint arXiv:2202.03286. https://arxiv.org/abs/2202.03286

Reber, R., Winkielman, P., & Schwarz, N. (1998). Effects of perceptual fluency on affective judgments. Psychological Science, 9(1), 45–48. https://doi.org/10.1111/1467-9280.00008

Strachan, J. W. A., Albergo, D., Borghini, G., Pansardi, O., Scaliti, E., Gupta, S., Vora, K., Belot, A., Banissy, M., Bhatt, M., Bard, N., Sherrat, M., Deroy, O., Shafto, P., Reggiani, C., Gu, X., & Becchio, C. (2024). Testing theory of mind in large language models and humans. Nature Human Behaviour, 8, 1285–1295. https://doi.org/10.1038/s41562-024-01882-z

Waytz, A., Cacioppo, J., & Epley, N. (2010). Who sees human? The stability and importance of individual differences in anthropomorphism. Perspectives on Psychological Science, 5(3), 219–232. https://doi.org/10.1177/1745691610369336

Wittgenstein, L. (2009). Philosophical investigations (G. E. M. Anscombe, P. M. S. Hacker, & J. Schulte, Trans.; 4th ed.). Wiley-Blackwell. (Original work published 1953)

Yin, X., Zahmat Doost, E., Zhou, S., Arya Yadav, G., & Gorman, J. (2025). When researchers say mental model/theory of mind of AI, what are they really talking about? (arXiv:2510.02660). arXiv. https://arxiv.org/abs/2510.02660

Leave a Reply