文章背景与核心概要
现有的个性化大语言模型(LLM)基准测试大多依赖于简化的文本画像或孤立的行为信号,这导致在评估跨领域行为个性化时存在明显盲区。为了填补这一空白,研究人员推出了 LUNAR——首个旨在评估大语言模型如何利用跨日常生活领域(衣、食、住、行)的纵向应用交互历史来进行个性化响应的基准测试。
为克服数据稀疏性和隐私挑战,LUNAR 采用了一种以真实世界行为模式为基础的多阶段粗到细合成流水线。对 19 个主流大语言模型的评估表明,单纯获取行为日志尚不足以实现强大的个性化。核心研究发现强调,证据选择、跨领域整合以及隐私控制仍然是未来个性化大语言模型发展面临的根本性挑战。
LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs
Summary
Existing personalized Large Language Model (LLM) benchmarks rely heavily on simplistic textual personas or isolated behavioral signals, leaving a gap in evaluating cross-domain behavioral personalization. To bridge this gap, researchers introduce LUNAR, the first benchmark designed to assess how LLMs personalize responses using longitudinal app interaction histories across universal daily-life domains (clothing, food, housing, and mobility).
To overcome data sparsity and privacy challenges, LUNAR employs a multi-stage coarse-to-fine synthesis pipeline anchored in real-world behavioral patterns. Evaluations on 19 mainstream LLMs reveal that raw access to behavioral logs is insufficient for robust personalization. Key findings highlight that evidence selection, cross-domain integration, and privacy control remain fundamental challenges for future personalized LLM development.
Metadata
- arXiv ID: arXiv:2608.05246 [cs.AI]
- Subjects: Artificial Intelligence (
cs.AI) - Submission Dates: Submitted on August 5, 2026; Last revised August 14, 2026 (v2)
- DOI: 10.48550/arXiv.2608.05246
Authors
- Jiahao Zhang
- Yongzhi Tong
- Zelin Fu
- Pengde Zhao
- Yanmei Jiang
- Feng Jiang
- Min Yang
Abstract
现有的大语言模型(LLM)个性化基准测试主要依赖于文本角色(personas)或孤立的行为信号,对跨领域行为个性化的评估相对有限。在这些场景中,模型的回答必须基于异构的日常生活活动。为了填补这一空白,我们推出了 LUNAR,这是首个用于评估 LLM 如何利用衣、食、住、行等通用日常生活领域中纵向应用交互历史来实现个性化响应的基准测试。
Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities. To address this gap, we introduce LUNAR, the first benchmark for evaluating how LLMs personalize responses from longitudinal app interaction histories across universal daily-life domains, including clothing, food, housing, and mobility.
为了支持可扩展的基准构建,同时减轻数据稀疏性和隐私顾虑,LUNAR 采用了一个以真实世界行为模式为基础的多阶段粗到细合成流水线。保真度分析表明,它比其他合成基准更能贴近真实的行为分布。对 19 个主流 LLM 的实验表明,访问行为日志是深度个性化的必要条件,但并非充分条件:既更多的上下文 nor 更大的模型都不能保证更好的性能;有效的个性化取决于跨领域的相关证据选择和整合。直接检索细粒度行为记录的表现持续优于压缩记忆,然而更强的个性化可能会以牺牲隐私保护为代价。这些发现将证据选择、跨领域整合和隐私控制确定为个性化 LLM 的核心挑战。
To support scalable benchmark construction while mitigating data sparsity and privacy concerns, LUNAR uses a multi-stage coarse-to-fine synthesis pipeline grounded in real-world behavioral patterns. Fidelity analyses show closer alignment with real behavioral distributions than other synthetic benchmarks. Experiments on 19 mainstream LLMs show that access to behavioral logs is necessary but not sufficient for deep personalization: neither more context nor larger models guarantees better performance; effective personalization depends on selecting and integrating relevant evidence across domains. Direct retrieval of fine-grained behavioral records consistently outperforms compressed memory, while stronger personalization can come at the cost of privacy protection. These findings identify evidence selection, cross-domain integration, and privacy control as key challenges for personalized LLMs.
Key Takeaways & Findings
-
跨领域整合的必要性: 访问全面的行为日志是必需的,但模型需要复杂的机制来跨不同领域选择和整合证据,而不是盲目地依赖上下文大小或模型规模。
- The Necessity of Cross-Domain Integration: Access to comprehensive behavioral logs is required, but models need sophisticated mechanisms to select and integrate evidence across diverse domains rather than relying blindly on context size or model scale.
-
检索与压缩: 在保持个性化质量方面,直接检索细粒度行为记录的表现始终优于压缩记忆表示。
- Retrieval vs. Compression: Direct retrieval of fine-grained behavioral records consistently outperforms compressed memory representations in maintaining personalization quality.
-
隐私权衡: 实现更深入的个性化可能会无意中损害用户隐私保护,这强调了未来 LLM 架构中对强大隐私控制框架的需求。
- The Privacy Trade-off: Achieving deeper personalization can inadvertently compromise user privacy protection, emphasizing the need for robust privacy control frameworks in future LLM architectures.
Links and Resources
- Full-Text Access: View PDF | HTML Version (Experimental)
- Citations & Metrics: Google Scholar | Semantic Scholar | NASA ADS