Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

Researchers propose boundary-aware self-distillation to refine model safety refusal for specific harmful subsets within broader topics. The paper argues that current topic-level guards like LlamaGuard-3 fail to distinguish between benign and harmful prompts in complex domains like politics. The method aims to train models to refuse only targeted harms while maintaining utility for related benign queries.

Cover image for Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic