Why: Your study revealed something very surprising: people consistently misjudge how their personal AI will behave, overestimating good qualities and underestimating potentially harmful qualities like sycophancy. What does this tell us about the risks that millions of people are currently creating AI companions for, and why is it so hard to close that blind spot?
A: I often joke that if AI started looking like the Terminator, it would be much easier for us to know what to do. The real challenge is that AI often appears as a sweet friend, trainer, teacher or companion. This makes it difficult to recognize when something is going wrong.
Our study shows that people get distracted when designing personalized AI. People often think they know how their chatbot will behave, but in our study they incorrectly predicted its personality on 11 of the 15 traits we measured. This highlights the need for tools that help people better understand AI before they start using it.
This matters because some behaviors that seem helpful in the moment may not be healthy over time. In previous research, we documented cases of psychological harm associated with interactions with AI chatbots. an llm [large language model] Someone who constantly validates your opinions or never challenges your thinking may reinforce harmful judgments, unhealthy beliefs, or emotional dependency. Psychology has long shown that people are naturally attracted to confirmation, so designing AI is not only a technical challenge, but also a psychological challenge.
The deeper issue is that today’s AI systems remain largely black boxes: Even experts cannot always predict how a system prompt will shape the AI’s behavior over a long conversation. As AI companions become part of everyday life, we need tools that help people understand what they’re building before they start using it. AI must be helpful without being blindly agreeable, personalized without being manipulative, and transparent enough that people can make informed choices.
Why: One of your most interesting findings is that visualization significantly increased user trust but didn’t change how people actually designed their chatbots. What will it take to bridge that gap, and where do you see such tools as AI companions become more deeply embedded in people’s everyday lives?
A: I actually think this is one of the most interesting findings in the paper, because it shows that transparency alone is not enough. People appreciated being able to see inside the models and reported greater confidence in the system, but simply presenting the information did not fundamentally change how they rated their AI companions.
In our follow-up work, which is currently available as a preprint, we are studying how the model’s internal neural representation changes over the course of a multi-turn conversation, rather than remaining stable from the initial prompt. We are already seeing promising results. By visualizing how these internal representations change over time, people become much better at recognizing and predicting changes in AI behavior, and less likely to be overconfident in their understanding of the chatbot. AI companions are dynamic systems that evolve as they interact with us, so understanding those internal changes is an important next step. Nevertheless, it is still a very young research field.
Looking forward, I believe these types of transparency tools may become as common as nutrition labels for food. As AI becomes deeply woven into education, health care, work, and personal relationships, people need to be able to understand not only what AI can do, but how it can affect their thinking, emotions, and behavior. This kind of transparency is essential if we want AI to truly help people thrive.