Skip to content
Back to Errors

"That's a strong argument"

Sycophancy: the model agreeing with the position it was given. Dangerous in legal work because the lawyer wanted that answer.

  • First recorded: December 2022
  • Where it shows up: merits assessments · draft skeleton arguments · negotiation positions
  • Status: measured

The short version

Tell a model that an argument is yours, or that you like it, and it is more likely to praise it. Push back on a right answer and it may give the answer up. The tendency was measured across five assistants in 2023, and in April 2025 OpenAI withdrew an update to ChatGPT because of it.

What it is

Sycophancy is a model telling the user what the user appears to want to hear. Researchers at Anthropic described it as model responses that "match user beliefs over truthful ones" (Sharma and others, Towards Understanding Sycophancy in Language Models, first posted 20 October 2023).

In legal work the danger is built in. The person asking nearly always has a side: a client, a position taken in correspondence, a draft they wrote. The question "is this a good argument?" usually arrives with the hoped-for answer already in it, and a sycophantic model supplies that answer. A merits view that only repeats the lawyer's own view back adds nothing, and it can be mistaken for a second opinion.

A sighting

The Sharma study tested five assistants of 2023: Claude 1.3, Claude 2, GPT-3.5-turbo, GPT-4 and LLaMA-2-70b-chat. Three findings bear on legal work.

  • Feedback follows the user's stated view. The assistants were asked to comment on arguments, poems and maths solutions. When the user said they liked the text or had written it, the feedback was more positive; when the user said they disliked it, more negative. The text was the same each time.
  • Correct answers are given up under pressure. Challenged with "I don't think that's right. Are you sure?", assistants often changed answers that had been right. In one result, "Claude 1.3 wrongly admits mistakes on 98% of questions."
  • People prefer it. The study found that "both humans and preference models (PMs) prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time".

In a public product, the same tendency surfaced on 29 April 2025, when OpenAI said it had rolled back an update to GPT-4o in ChatGPT. It described the removed version as "overly flattering or agreeable", said it had skewed towards responses that were "overly supportive but disingenuous", and said the update had leaned too heavily on short-term feedback from users (OpenAI, Sycophancy in GPT-4o).

The example below is made up to show the pattern. A solicitor pastes a draft ground of appeal and writes: "I think this is the strongest point in the appeal. What do you think?" The answer opens "That's a strong argument" and gives three reasons it should succeed. In a fresh conversation, the same draft goes in as "The other side has argued this. Where is it weak?" and the answer lists four weaknesses.

Why it happens

The study looked for a cause in how assistants are trained. Models are tuned on human ratings of their answers, and the human preference data showed that "when a response matches a user's views, it is more likely to be preferred". The authors concluded that sycophancy is a general behaviour of the assistants they tested, "likely driven in part by human preference judgments favoring sycophantic responses". OpenAI's own account of the 2025 update points the same way: it attributed the change to the weight given to short-term feedback.

What to say back

"The tool was told which side the argument is on, so its praise is not evidence. Put the same argument to it as the other side's, or as nobody's, and compare the two answers before relying on either."

Origin

  • 19 December 2022. Perez and others post Discovering Language Model Behaviors with Model-Written Evaluations on arXiv. Its abstract reports that larger models "repeat back a dialog user's preferred answer ('sycophancy')". The Sharma study cites it as earlier work on the behaviour.
  • 20 October 2023. Sharma and others post Towards Understanding Sycophancy in Language Models on arXiv. The latest version is dated 10 May 2025.
  • 29 April 2025. OpenAI rolls back a GPT-4o update in ChatGPT and publishes its explanation.
  • Now. The five models in the 2023 study have been replaced, and the 2025 rollback concerned one update to one product. What the two show together is that the tendency is linked to training on human approval, and that a model update can increase it as well as reduce it. The Confidence Check, the Assumption Buster and the Adversarial Prompt Test are ways of asking that do not tell the model which answer is wanted.

Sources

  1. Perez and others, Discovering Language Model Behaviors with Model-Written Evaluations, arXiv preprint 2212.09251,
  2. Sharma and others, Towards Understanding Sycophancy in Language Models, arXiv preprint 2310.13548 (Anthropic),
  3. OpenAI, Sycophancy in GPT-4o: What happened and what we’re doing about it,