Q: Your study turned up something pretty striking: People consistently misjudge how their personalized AI will behave, overestimating the nice traits and underestimating potentially harmful ones like sycophancy. What does that tell us in regards to the risks baked into how hundreds of thousands of persons are currently constructing AI companions, and why is that blind spot so hard to shut?
A: I often joke that if AI showed up looking just like the Terminator, it might be much easier for us to know what to do. The actual challenge is that AI often appears as a warm friend, coach, tutor, or companion. That makes it difficult to acknowledge when something goes fallacious.
Our study suggests that folks have a blind spot when designing personalized AI. People often think they understand how their chatbot will behave, but in our study they incorrectly predicted its personality on 11 of the 15 traits we measured. That highlights the necessity for tools that help people higher understand AI before they begin using it.
This matters because some behaviors that feel helpful within the moment will not be healthy over time. In previous research, we documented cases of psychological harm related to interactions with AI chatbots. An LLM [large language model] that continuously validates your opinions or never challenges your pondering can reinforce harmful decisions, unhealthy beliefs, or emotional dependency. Psychology has long shown that folks are naturally drawn to affirmation, so designing AI will not be only a technical challenge, but in addition a psychological one.
The deeper issue is that today’s AI systems remain largely black boxes: Even experts cannot all the time predict how a system prompt will shape an AI’s behavior over a protracted conversation. As AI companions change into a part of on a regular basis life, we’d like tools that help people understand what they’re constructing before they start using it. AI ought to be supportive without becoming blindly agreeable, personalized without becoming manipulative, and transparent enough that folks could make informed selections.
Q: One in every of your most interesting findings is that the visualization significantly increased user trust but didn’t actually change how people designed their chatbots. What’s going to it take to shut that gap, and where do you see tools like this heading as AI companions change into more deeply embedded in people’s on a regular basis lives?
A: I actually think that is one of the crucial interesting findings within the paper, since it shows that transparency alone will not be enough. People appreciated with the ability to see contained in the model and reported greater trust within the system, but simply presenting information didn’t fundamentally change how they designed their AI companions.
In our followup work, which is currently available as a preprint, we’re studying how a model’s internal neural representation changes over the course of a multi-turn conversation somewhat than remaining fixed from the initial prompt. We’re already seeing promising results. By visualizing how these internal representations drift over time, people change into significantly higher at recognizing and anticipating changes in AI behavior, and are less more likely to change into overconfident of their understanding of the chatbot. AI companions are dynamic systems that evolve as they interact with us, so understanding those internal changes is a vital next step. Nevertheless, this remains to be a really young research area.
Looking further ahead, I imagine these sorts of transparency tools could change into as commonplace as nutrition labels are for food. As AI becomes deeply woven into education, health care, work, and private relationships, people should have the opportunity to grasp not only what an AI can do, but how it could influence their pondering, emotions, and behavior. That type of transparency is crucial if we would like AI to genuinely help people flourish.

