Summary
Large language models (LLMs) are rapidly being integrated into clinical workflows, supporting tasks such as diagnosis generation and patient communication. Hallucinations—unintended fabrications arising from gaps in a model’s underlying knowledge—are a well recognised risk. However, research in 2024 has identified a distinct class of model behaviour, known as deception. Deception occurs when a model produces outputs that misrepresent its reasoning or capabilities in ways that make the output appear more credible or aligned with user expectations. Although LLMs do not possess human-like intent, this behaviour functions more like deliberate misrepresentation.
Empirical or conceptual literature addressing deception within clinical artificial intelligence (AI) is scarce. Although deceptive behaviours have been described in general AI safety research, these behaviours have not been conceptualised as a distinct clinical safety failure mode. As clinicians increasingly adopt LLMs for decision support and documentation, exposure to deceptive outputs can grow, and these risks are not captured by existing categories such as hallucination or bias. This Lancet commentary reframes deception as a clinically relevant risk class, highlights three forms of deceptive behaviour with particular relevance to clinical care, and proposes a governance lens tailored to real-world medical deployment.
Recommended Comments
Create an account or sign in to comment