Science Cast

Layered Unlearning for Adversarial Relearning

Timothy QianMay 15, 2025 2:38am

Views (2)
Comments (0)

Export Citation

Voice is AI-generated

Connected to paperThis paper is a preprint and has not been certified by peer review

Layered Unlearning for Adversarial Relearning

arXivPDFMay 14, 2025 12:00am

Authors

Timothy Qian, Vinith Suriyakumar, Ashia Wilson, Dylan Hadfield-Menell

Abstract

Our goal is to understand how post-training methods, such as fine-tuning, alignment, and unlearning, modify language model behavior and representations. We are particularly interested in the brittle nature of these modifications that makes them easy to bypass through prompt engineering or relearning. Recent results suggest that post-training induces shallow context-dependent ``circuits'' that suppress specific response patterns. This could be one explanation for the brittleness of post-training. To test this hypothesis, we design an unlearning algorithm, Layered Unlearning (LU), that creates distinct inhibitory mechanisms for a growing subset of the data. By unlearning the first $i$ folds while retaining the remaining $k - i$ at the $i$th of $k$ stages, LU limits the ability of relearning on a subset of data to recover the full dataset. We evaluate LU through a combination of synthetic and large language model (LLM) experiments. We find that LU improves robustness to adversarial relearning for several different unlearning methods. Our results contribute to the state-of-the-art of machine unlearning and provide insight into the effect of post-training updates.

TwitterandLinkedIn

0 comments

Add comment

Layered Unlearning for Adversarial Relearning

Layered Unlearning for Adversarial Relearning

AI-powered Paper ChatBeta

AI-powered Paper ChatBeta

0 comments