
Kevin Lu, Nicky Kriplani, Rohit Gandikota, Minh Pham, David Bau, Chinmay Hegde, Niv Cohen
NeurIPS 2025 (Poster)
This work analyzes the mechanisms of existing concept erasure methods using a suite of probing methods to assess how much knowledge is remaining in the model. Overall, we find that many of the tested methods simply steer outputs away from the underlying knowledge as opposed to actually removing it.
Kevin Lu, Nicky Kriplani, Rohit Gandikota, Minh Pham, David Bau, Chinmay Hegde, Niv Cohen
NeurIPS 2025 (Poster)
This work analyzes the mechanisms of existing concept erasure methods using a suite of probing methods to assess how much knowledge is remaining in the model. Overall, we find that many of the tested methods simply steer outputs away from the underlying knowledge as opposed to actually removing it.

Nicky Kriplani, Minh Pham, Gowthami Somepalli, Chinmay Hegde, Niv Cohen
arXiv preprint arXiv:2503.00592
This paper introduces SolidMark, a new evaluation method that provides per-image memorization scores to better detect when diffusion models have memorized specific training images. We use SolidMark to re-evaluate existing memorization mitigation techniques and demonstrate its ability to assess pixel-level memorization in generative models.
Nicky Kriplani, Minh Pham, Gowthami Somepalli, Chinmay Hegde, Niv Cohen
arXiv preprint arXiv:2503.00592
This paper introduces SolidMark, a new evaluation method that provides per-image memorization scores to better detect when diffusion models have memorized specific training images. We use SolidMark to re-evaluate existing memorization mitigation techniques and demonstrate its ability to assess pixel-level memorization in generative models.