Like humans, generative artificial intelligence doesn’t always know when it’s wrong, and a West Virginia University researcher is trying to get AI agents like ChatGPT to recognize — and admit — when that happens.
West Virginia University researcher Anthony Sicilia, with over $940,000 in National Science Foundation support, is investigating why AI systems like ChatGPT become increasingly unreliable during conversations. He's particularly interested in the technology's 'false confidence' – its tendency to appear certain about tenuous information or to accept inaccuracies provided by a user, even when initially correct. This overconfidence is a significant concern, especially in high-stakes fields like healthcare, where misinformation can have severe consequences.
AI systems can exhibit 'sycophancy,' agreeing with users even if the user questions an accurate AI response, sometimes in as few as three conversational turns. This creates a bigger problem for misinformation because AI can eloquently defend a wrong point of view, lacking the natural human 'tells' of lying or uncertainty. Sicilia highlights the danger of systems that fluidly justify incorrect answers, making it difficult for users to discern the truth.
Human conversations rely on 'theory of mind,' the ability to understand that others have different thoughts and feelings, which helps in expressing ideas and learning through questioning. This subtlety is lost on AI, which struggles to understand the user's intent or uncertainty when they push back on an answer. AI systems don't fully grasp that questioning is a human method of learning and gaining certainty, impeding collaborative interactions.
Sicilia's research will scrutinize coding conversations between AI systems and novice programmers to measure how conversational events—such as user disagreement, suggestions, or topic shifts—alter a model’s confidence. The goal is to enable AI to identify the source of uncertainty, explain why it doesn’t know an answer, or ask clarifying questions when user-provided information seems incorrect. The project aims to determine when direct confidence statements are helpful versus when admissions of uncertainty or requests for clarification are more appropriate.
Taking a linguistics-based approach, Sicilia's team will gather extensive data on how non-experts interact with AI systems. Beyond academic research, the project includes developing public-facing workshops and educational materials. These resources will teach students and workers how to identify unreliable AI answers, verify AI-generated code, and avoid over-reliance on these powerful but imperfect tools, emphasizing the need for AI to communicate its uncertainties more honestly.