ReasoningShield: Safety Detection over Reasoning Traces of Large Reasoning Models

Changyi Li; Jiayi Wang; Xudong Pan; Geng Hong; Min Yang

ReasoningShield: Safety Detection over Reasoning Traces of Large Reasoning Models

Changyi Li, Jiayi Wang, Xudong Pan, Geng Hong, Min Yang

19 Sept 2025 (modified: 01 Feb 2026)ICLR 2026 Conference Withdrawn SubmissionEveryoneRevisionsBibTeXCC BY 4.0

Keywords: AI Safety, Large Reasoning Models, Content Safety Detection

Abstract: Large Reasoning Models (LRMs) leverage transparent reasoning traces, known as Chain-of-Thoughts (CoTs), to break down complex problems into intermediate steps and derive final answers. However, these reasoning traces introduce unique safety challenges: harmful content can be embedded in intermediate steps even when final answers appear benign. Existing moderation tools, designed to handle generated answers, struggle to effectively detect hidden risks within CoTs. To address these challenges, we introduce *ReasoningShield*, a lightweight yet robust framework for moderating CoTs in LRMs. Our key contributions include: (1) formalizing the task of CoT moderation with a multi-level taxonomy of 10 risk categories across 3 safety levels, (2) creating the first CoT moderation benchmark which contains 9.2K pairs of queries and reasoning traces, including a 7K-sample training set annotated via a human-AI framework and a rigorously curated 2.2K human-annotated test set, and (3) proposing a specialized framework tailored for complex reasoning tasks, which utilizes a structured stepwise analysis paradigm and a strategic two-stage training pipeline to capture risk propagation and boundary ambiguity. Experiments show that ReasoningShield achieves state-of-the-art performance, outperforming task-specific tools like LlamaGuard-4 by 35.6\% and general-purpose commercial models like GPT-4o by 15.8\% on benchmarks, while also generalizing effectively across diverse reasoning paradigms, tasks, and unseen scenarios. All resources are released at https://anonymous.4open.science/r/ReasoningShield.

Primary Area: alignment, fairness, safety, privacy, and societal considerations

Submission Number: 17273

Loading