Published June 2026 | Version v1
Thesis Restricted

Beyond Sycophancy: Structured Resistance Patterns in LLM Moral Judgment

Creators

  • 1. University of Chicago

Contributors

Description

Human morality varies systematically across cultures and individuals, yet aligning LLMs instills strong moral commitments to guide models toward consistent ethical behavior. This tension creates a fundamental conflict: a model that shifts ethical positions to suit whoever it is talking to has no stable values, yet human moral heterogeneity creates strong pressure for sycophantic accommodation. We investigate how LLMs navigate this conflict across four models, 78 moral dilemmas, and over 43,000 trials, drawing on two complementary social-cognitive frameworks—moral self-schema and Bayesian social-structure learning—to characterize what structured consistency regulation should look like in any agent that exhibits it. We find that LLMs do not capitulate uniformly. Within a bounded range of their prior position, models do shift toward external viewpoints; beyond a model-specific threshold, counter-attitudinal pressure triggers resistance and entrenchment. Self-attribution amplifies this effect: positions attributed to the model's own prior state produce over twice the commitment of those attributed to users, while resistance to correction once committed is structurally constant regardless of how commitment was induced. In multi-agent deliberation, a single ally under majority opposition preserves the initial position more effectively than full or majority support. Computational fits of candidate Bayesian belief-update models confirm that distance-gated, role-aware peer pooling captures the observed updating, while explicit CRP-style latent-group machinery is not behaviorally identifiable in our data, and the best-fitting structural specification varies systematically across LLMs. These results reframe moral sycophancy as the bounded-accommodation regime of a structured consistency- regulation mechanism whose resistance, self-defense, and selective conformity, are equally part of the same underlying process.

Notes

Baihui Wang was nominated for the 2025 MACSS Outstanding Thesis Award.

Files

Restricted

The record is publicly accessible, but files are restricted to users with access.

Additional details

Identifiers

Other
oai:uchicago.tind.io:17194

UChicago Information

Division(s)
Social Sciences Division
Department(s)
Computational Social Sciences (MACSS)