文章背景与核心概要
随着智能体(Agentic)AI 系统获得对用户敏感数据更广泛的访问权限,确保隐私安全已成为一个核心关注点。现有研究主要聚焦于“语境隐私”(Contextual Privacy),即模型基于通用社会规范来调节信息的披露;而本文则创新性地提出了“个性化隐私”(Personalized Privacy)的概念。该方法认识到可接受的信息披露边界具有主观性,且因不同用户的个体差异而异。
为解决这一问题,作者推出了 P3Bench(个性化隐私保护基准),这是一个用于测试大语言模型遵循特定、用户自定义披露政策能力的全新评估框架。研究发现,当前最先进的模型(如 Qwen2.5-7B 和 Gemma3-4B)在政策执行方面存在显著困难,表现出极高的“政策忽视”率。为此,作者提出了一种名为 Repair 的推理时干预方法,通过定向调控特定的注意力头,使模型的回复能够符合用户特定的隐私偏好。
通过注意力头干预实现大语言模型中的个性化隐私控制
作者: Junseok Kim, Nakyeong Yang, Kyomin Jung
发表时间: 2026年8月21日
会议/期刊: EMNLP 2026
arXiv ID: 2608.21209
摘要
As agentic AI systems gain broader access to sensitive user data, ensuring privacy has become a paramount concern. While existing research focuses on "contextual privacy"—where models regulate disclosure based on general social norms—this paper introduces the concept of "personalized privacy." This approach acknowledges that acceptable disclosure boundaries are subjective and vary between individual users.
随着智能体(Agentic)AI 系统获得对用户敏感数据更广泛的访问权限,确保隐私安全已成为一个核心关注点。现有研究主要聚焦于“语境隐私”(Contextual Privacy),即模型基于通用社会规范来调节信息的披露;而本文则创新性地提出了“个性化隐私”(Personalized Privacy)的概念。该方法认识到可接受的信息披露边界具有主观性,且因不同用户的个体差异而异。
The authors introduce P3Bench (Personalized Privacy Preservation Benchmark), a new evaluation framework that tests how well LLMs adhere to specific, user-defined disclosure policies. Their findings reveal that current state-of-the-art models (such as Qwen2.5-7B and Gemma3-4B) struggle significantly with policy enforcement, exhibiting high rates of "policy ignorance." To solve this, the authors propose Repair, an inference-time intervention method that targets specific attention heads to align model responses with user-specific privacy preferences.
作者推出了 P3Bench(个性化隐私保护基准),这是一个用于测试大语言模型遵循特定、用户自定义披露政策能力的全新评估框架。研究发现,当前最先进的模型(如 Qwen2.5-7B 和 Gemma3-4B)在政策执行方面存在显著困难,表现出极高的“政策忽视”率。为此,作者提出了一种名为 Repair 的推理时干预方法,通过定向调控特定的注意力头,使模型的回复能够符合用户特定的隐私偏好。
核心贡献
- Personalized Privacy Framework: A shift from universal privacy norms to user-centric, customizable disclosure policies.
- P3Bench: A novel benchmark designed to evaluate how LLMs handle personalized privacy constraints.
- Performance Analysis: Empirical evidence showing that prompt-based policy enforcement is unreliable, with failure rates reaching up to 74.28% in tested models.
- Repair Method: A robust, inference-time attention head intervention technique that significantly improves adherence to personalized privacy requirements.
- 个性化隐私框架: 从普适的隐私规范转向以用户为中心、可定制的信息披露政策。
- P3Bench: 一个旨在评估大语言模型如何处理个性化隐私约束的新型基准。
- 性能分析: 实证证据表明,基于提示词的政策执行不可靠,在测试模型中的失败率高达 74.28%。
- Repair 方法: 一种强健的推理时注意力头干预技术,显著提高了模型对个性化隐私要求的遵从度。
访问与资源

元数据
- Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
- DOI: https://doi.org/10.48550/arXiv.2608.21209
- 研究方向: 人工智能 (cs.AI);计算与语言 (cs.CL);机器学习 (cs.LG)
- DOI: https://doi.org/10.48550/arXiv.2608.21209