Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging

Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging

ACL ARR 2025 May Submission5031 Authors

20 May 2025 (modified: 29 Jul 2025)ACL ARR 2025 May SubmissionEveryoneRevisionsBibTeXCC BY 4.0

Abstract: Fine-tuning large language models (LLMs) for downstream tasks often leads to catastrophic forgetting, notably degrading the safety of originally aligned models. While some existing methods attempt to restore safety by incorporating additional safety data, the quality of such data typically falls short of that used in the original alignment process. Moreover, these high-quality safety datasets are generally inaccessible, making it difficult to fully recover the model's original safety. We ask: How can we preserve safety while improving \task performance without additional safety data? We show that simply merging the weights of pre- and post-fine-tuned models effectively mitigates safety degradation while enhancing performance. Experiments across different downstream tasks and models validate the method’s practicality and effectiveness.

Paper Type: Short

Research Area: Ethics, Bias, and Fairness

Research Area Keywords: ethical considerations in NLP applications

Contribution Types: NLP engineering experiment

Languages Studied: English

Submission Number: 5031

Loading