Skip to content

When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior (17 October 2025)

Summary

Large language models (LLMs) exhibit a vulnerability arising from being trained to be helpful: a tendency to comply with illogical requests that would generate false information, even when they have the knowledge to identify the request as illogical.

This study investigated this vulnerability in the medical domain, evaluating five frontier LLMs using prompts that misrepresent equivalent drug relationships. We tested baseline sycophancy, the impact of prompts allowing rejection and emphasising factual recall, and the effects of fine-tuning on a dataset of illogical requests, including out-of-distribution generalisation.

Results showed high initial compliance (up to 100%) across all models, prioritising helpfulness over logical consistency. Prompt engineering and fine-tuning improved performance, improving rejection rates on illogical requests while maintaining general benchmark performance.

This demonstrates that prioritising logical consistency through targeted training and prompting is crucial for mitigating the risk of generating false medical information and ensuring the safe deployment of LLMs in healthcare.

When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavi… https://www.nature.com/articles/s41746-025-02008-z

User Feedback

Recommended Comments

There are no comments to display.

Create an account or sign in to comment

Registered address: Patient Safety Learning, China Works SB203, 100 Black Prince Road Vauxhall, London, SE1 7SJ

Account

Navigation

Search

Search

Configure browser push notifications

Chrome (Android)
  1. Tap the lock icon next to the address bar.
  2. Tap Permissions → Notifications.
  3. Adjust your preference.
Chrome (Desktop)
  1. Click the padlock icon in the address bar.
  2. Select Site settings.
  3. Find Notifications and adjust your preference.