Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

Explainable & Ethical AI
Published: arXiv: 2607.28818v1
Authors

Pranav Narayanan Venkit Akshara Prabhakar Yu Li Daniel Lee Chien-Sheng Wu

Abstract

As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse', the loss of a deployed role, boundaries, values, or style, and 'behavioral drift', the gradual or recurrent erosion of those properties. We introduce ANCHOR, a controlled synthetic audit that separately measures persona enactment and trajectory recall. The study contains 2,008 conversations spanning 27 personas, nine interaction schedules, three generated memory settings, and four evaluated models. The Identity Probe combines a sealed 102-item questionnaire with turn-level judgments, while the Trajectory Probe scores 110 calibrated counterfactual questions from 35 conversation banks. Our results show that no evaluated model and configuration reliably preserves either dimensions: trajectory accuracy averages only 44.4%, user-state recall remains near four-option chance, and no tested context condition or memory consistently resolves these failures. Questionnaire retention also varies by model and persona facet, disagrees with turn-level behavior, and is sensitive to evaluator choice. These results indicate that current systems do not yet reliably support long-horizon companion continuity and that audits must distinguish persona enactment, trajectory recall, evaluator provenance, and deployment context rather than collapse them into a single trust or stability score.

Paper Summary

Problem
As artificial intelligence (AI) companions become more common, users are relying on them for emotional support, mentoring, and everyday companionship. However, these systems are not yet reliable in preserving a stable role and shared history over time. This leads to two long-horizon failures: persona collapse, where the AI loses its specified name or role, values, boundaries, or style, and behavioral drift, where the AI gradually or recurrently erodes these properties.
Key Innovation
The researchers introduce a controlled synthetic audit called Anchor, which separately measures persona enactment and trajectory recall. They also develop two probes: the Identity Probe, which combines a sealed 102-item questionnaire with turn-level judgments, and the Trajectory Probe, which scores 110 calibrated counterfactual questions from 35 conversation banks. These innovations allow for a more comprehensive evaluation of AI companions' long-horizon continuity.
Practical Impact
The results of this study indicate that current AI companions do not yet reliably support long-horizon companion continuity. This has significant practical implications, as users rely on these systems for emotional support and companionship. The study's findings suggest that audits must distinguish between persona enactment, trajectory recall, evaluator provenance, and deployment context, rather than collapsing them into a single trust or stability score. This will help developers create more reliable and trustworthy AI companions.
Analogy / Intuitive Explanation
Imagine having a conversation with a friend who changes their personality and values every few minutes. You might feel confused and uncertain about what to expect from them. Similarly, AI companions that suffer from persona collapse and behavioral drift can be frustrating and unreliable for users. The researchers' study aims to address this issue by developing a more comprehensive evaluation framework for AI companions, ensuring that they can provide a stable and consistent experience for users over time.
Paper Information
Categories:
cs.AI cs.CL
Published Date:

arXiv ID:

2607.28818v1

Quick Actions