arXiv cs.CLAugust 19, 2026
Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits
Excerpt
arXiv:2503.20182v2 Announce Type: replace Abstract: As large language models (LLMs) increasingly function as human-like assistants exhibiting human-like personality traits, understanding their behavioral characteristics becomes essential for responsible AI development. However, existing evaluation efforts, which often adapt human psychological assessments such as the Big Five Inventory (BFI), face two significant limitations. First, these approaches often lack reliability, as minor prompt variat