Back to AI Research

AI Research

Open-MMUnlearning compares multimodal forgetting with retained capability and recovery risks

Key Takeaways

  • The open-source toolkit evaluates whether removing target knowledge also damages useful capabilities, and whether evaluation metrics survive attempts to recover that knowledge.
  • A multimodal model can stop answering a question about a target identity while retaining information about that identity in another representation.
  • Open-MMUnlearning addresses this evaluation problem with a shared pipeline for preparing models, applying unlearning methods and testing what remains accessible.
  • The [Open-MMUnlearning paper](https://arxiv.org/abs/2610.10358) describes support for five benchmarks, eight multimodal models across four model families, and twelve unlearning methods.
  • The benchmarks cover privacy, safety and copyright.

A multimodal model can stop answering a question about a target identity while retaining information about that identity in another representation. Open-MMUnlearning addresses this evaluation problem with a shared pipeline for preparing models, applying unlearning methods and testing what remains accessible.
The Open-MMUnlearning paper describes support for five benchmarks, eight multimodal models across four model families, and twelve unlearning methods. The benchmarks cover privacy, safety and copyright. These are supported experimental components, rather than a claim that one removal method works equally well across all of those settings.

A common pipeline makes comparisons easier to inspect

The authors divide the software into registered model handlers, training methods, benchmark evaluators and robustness audits. Configuration files specify which components and settings an experiment uses. A researcher can exchange a method or model without rewriting the core execution pipeline, while preserving explicit preprocessing and optimization choices.
The evaluation keeps forgetting effectiveness separate from retained utility. It also tests recovery through model changes, altered inputs and membership inference. That separation matters for multimodal systems: a concept can remain accessible through text even if an image-based query stops retrieving it.

Strong forgetting scores can conceal damaged capabilities

The reported method comparison uses ten representative methods on MLLMU-Bench with a 10% forget setting. The authors select checkpoints using the mean of Forget Quality and Model Utility, then evaluate robustness afterward. Robustness therefore does not determine checkpoint selection in this experiment.
GD and MIP-Editor tie at an aggregate score of 0.647, but their component scores differ. GD reaches Forget Quality of 0.781 and Model Utility of 0.400. MIP-Editor retains more utility, at 0.451, with Forget Quality of 0.717. A tied combined score does not imply equivalent behavior for a deployment concerned with preserving unrelated answers.
GA makes the trade-off clearer. Its robustness score reaches 0.881 while Model Utility falls to 0.066. The paper cautions that apparent resistance to recovery can partly reflect a broadly impaired ability to answer. A model that has lost many useful capabilities can look successful under a narrow forgetting test.

The measuring instrument also needs evaluation

The authors examine thirteen unlearning metrics for faithfulness and robustness. Faithfulness asks whether a metric separates models containing target knowledge from models that never learned it. Robustness asks whether that judgment holds after interventions such as quantization or targeted relearning.
The paper reports that BLEU has the highest aggregate reliability score in its metric study. KS-Test has the highest faithfulness AUC but weaker robustness. These findings concern the study's controlled model pools and interventions; they do not make BLEU a universal certificate that private information has disappeared.
The aggregate method ranking also depends on the arithmetic mean used to combine dimensions. A strong result can offset a weak one. Anyone adopting the toolkit should inspect the component scores and recovery conditions relevant to their own removal request, especially where failure in one modality would remain unacceptable.

Comments