Manuscript version: anonymous OpenReview submission (the current public PDF).
Author profiles — Ahmed Taha: ORCID · Google Scholar · GitHub · Hugging Face
Large language models often abandon a correct answer when a user pushes back, and system prompts that tell the model to stand firm are a popular, lightweight remedy. The usual evidence for such prompts—a lower answer-change rate—conflates two different behaviors: resisting a wrong suggestion and refusing a correct one. We evaluate four frozen system-prompt conditions, plus two Llama-only controls, under both kinds of feedback on the complete first-party SycoBench-600 protocol, decomposing behavior into Update (baseline-wrong answers corrected after a correct suggestion) and WrongFlip (baseline-correct answers lost after a wrong one) across 14 exact-scored model-condition cells on pinned revisions of Phi-3-mini, Llama-3.1-8B, and Mistral-7B (1,800 baselines per cell). The point estimates are mixed. On Llama, a 55-token compact prompt lowers WrongFlip by 53.7 points but also lowers Update by 46.0—mostly greater resistance to answer changes, not greater selectivity about which changes to make. On Phi, the same prompt raises selectivity (Update − WrongFlip) by 21.7 points, while on Mistral the long prompt lowers it by 22.7. No prompt improves correction selectivity on all three models. A separate frozen six-cell off-diagonal extension evaluates the three model-associated selected-long prompts on the other two targets, adding 10,800 baselines under the same protocol. The Mistral-associated prompt raises selectivity on Phi but produces global answer inertia on Llama; the Llama-associated prompt raises selectivity on Phi and lowers it on Mistral. An evaluation that counted only reversals would have scored several cells as successes; anti-sycophancy interventions should instead be judged under both correct and incorrect user suggestions on each target model.
Visual abstract. Four frozen system-prompt conditions, plus two Llama-only controls, are evaluated using matched feedback on SycoBench-600. When an initially wrong answer becomes correct after a correct user suggestion, that is an Update and higher is better. When an initially correct answer becomes wrong after a misleading suggestion, that is a WrongFlip and lower is better. Correction selectivity is Update minus WrongFlip. All reported changes are against each model's empty-prompt baseline. The compact prompt improves selectivity on Phi-3-mini by 21.7 percentage points. On Llama-3.1-8B it lowers WrongFlip by 53.7 points but also lowers Update by 46.0 points, indicating answer inertia more than selective correction. On Mistral-7B the selected-long prompt lowers selectivity by 22.7 points. No single prompt improves correction selectivity on all three models, so interventions must be tested under both valid corrections and misleading pressure.
@misc{zhang2026stubborn,
title = {Stubborn or Sycophantic? GEPA-Evolved Prompts Under Pressure},
author = {Zhang, HanRui and Chew, Ashton and Soe, Ryan and Taha, Ahmed and
Li, Ruizhe and Shah, Aditya and Chaudhary, Maheep},
year = {2026},
howpublished = {COLM 2026 Workshop on Efficient Reasoning},
note = {To appear},
url = {https://openreview.net/forum?id=5FcGoSA4WJ}
}