Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models

Aryan Kasat; Smriti Singh; Vinija Jain; Aman Chadha

Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models

Aryan Kasat, Smriti Singh, Vinija Jain, Aman Chadha

Published: 01 Apr 2026, Last Modified: 25 Apr 2026ICLR 2026 Workshop LLM ReasoningEveryoneRevisionsBibTeXCC BY 4.0

Track: long paper (up to 10 pages)

Keywords: LLM Reasoning, LLM Alignment, LLM Moral Reasoning, Kohlberg's Theory of Development

TL;DR: LLMs show reasoning inconsistency when asked to reason on complex moral dilemmas

Abstract: Do large language models reason morally, or do they merely sound like they do? We investigate whether LLM responses to moral dilemmas exhibit genuine developmental progression through Kohlberg's stages of moral development, or whether alignment training instead produces reasoning-\emph{like} outputs that superficially resemble mature moral judgment without the underlying developmental trajectory. Using an LLM-as-judge scoring pipeline validated across three judge models, we classify more than 600 responses from 13 LLMs spanning a range of architectures, parameter scales, and training regimes across six classical moral dilemmas, and conduct ten complementary analyses to characterize the nature and internal coherence of the resulting patterns. Our results reveal a striking inversion: responses overwhelmingly correspond to post-conventional reasoning (Stages~5--6) regardless of model size, architecture, or prompting strategy, the effective inverse of human developmental norms, where Stage~4 dominates. Most strikingly, a subset of models exhibit \emph{moral decoupling}: systematic inconsistency between stated moral justification and action choice, a form of logical incoherence that persists across scale and prompting strategy and represents a direct reasoning consistency failure independent of rhetorical sophistication. Model scale carries a statistically significant but practically small effect ($F(2,229)=6.05$, $p=0.003$, $\eta^2=0.050$, $d=0.55$); training type has no significant independent main effect ($p=0.065$); and models exhibit near-robotic cross-dilemma consistency (ICC~$>$~0.90), producing logically indistinguishable responses across semantically distinct moral problems. We posit that these patterns constitute evidence for \emph{moral ventriloquism}: the acquisition, through alignment training, of the rhetorical conventions of mature moral reasoning without the underlying developmental trajectory those conventions are meant to represent.

Presenter: ~Smriti_Singh2

Format: Yes, the presenting author will attend in person if this work is accepted to the workshop.

Anonymization: This submission has been anonymized for double-blind review via the removal of identifying information such as names, affiliations, and identifying URLs.

Submission Number: 170

Loading