LLM Probing with Contrastive Eigenproblems: Improving Understanding and Applicability of CCS

Stefan F. Schouten; Peter Bloem

LLM Probing with Contrastive Eigenproblems: Improving Understanding and Applicability of CCS

Stefan F. Schouten, Peter Bloem

Published: 30 Sept 2025, Last Modified: 03 Nov 2025Mech Interp Workshop (NeurIPS 2025) PosterEveryoneRevisionsBibTeXCC BY 4.0

Keywords: Probing, Foundational work, Applications of interpretability

TL;DR: Proposes relative contrast consistency, and a method to solve for it by formulating it as an eigenproblem.

Abstract: Contrast-Consistent Search (CCS) is an unsupervised probing method able to test whether large language models represent binary features, such as truth, in their internal activations. While CCS has shown promise, its two-term objective has been only partially understood. In this work, we revisit CCS with the aim of clarifying its mechanisms and extending its applicability. We argue that what should be optimized for, is relative contrast consistency. Building on this insight, we reformulate CCS as an eigenproblem, yielding closed-form solutions with interpretable eigenvalues and natural extensions to multiple variables. We evaluate these approaches across a range of datasets, finding that they recover similar performance to CCS, while avoiding problems around sensitivity to random initialization. Our results suggest that relativizing contrast consistency not only improves our understanding of CCS but also opens pathways for broader probing and mechanistic interpretability methods.

Submission Number: 171

Loading