← Back to all articles
arXiv cs.AIAugust 17, 2026

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

Excerpt

arXiv:2608.05246v2 Announce Type: replace Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities. To address this gap, we introduce LUNAR, the first benchmark for evaluating how LLMs personalize responses from longitudinal app interaction histories across universal daily-life domains, including