Large Language Model Psychometrics,A Systematic Review ofEvaluation. Validation. and EnhancementHaoran Ye,Jing,Jin', Yulhang Xie!.Xin Zhane-.GuoieSong 14sStateKeyLaboratory of General Artificial IntelligenceSZ0ZABN EIITOS]IAStZ80S0SZAIXIDSchool of Intellieence Science and Technology, Peking University2School of Psychological and Cognitive Sciences, Pcking University Key Laboaoy of Machine eceopio Mimisty of Educaion, Pdking Unvesity4pKI.Wihan Institute for Artificial Intelligencehryedstu,pku.edu.cngjsong@pku.edu. cnProjet Websice:htp ./ lm-syshometis.comLLMPsychometricsAbstractlanguagemodes(LLAMS) has oupaced tnatioaleak-specific benchmarks, and estabishing human-centered evaluatat with Psychometrics. the science of quantifving the intangible aspects of hs, and intelligence. This survey introduces and synts, which leverages psychometric ins trumehanceLLMs We sysemaicall explore theciples,broadening evaluation scopes, refininment of human-centered Al systems for socictal benefit. A curated repository of LM psychometicresources is available at htps /g ithutb com/valuceby te-ai/Awesome-IL.M-Psychometrics
Contents Introduction2Preliminarv and Methodological Foundations2.1Large LanguageModels2.2Psychometrics,23Pwvchomctic Evaluation of AlBecfore the Fra of LLMs....LLMPsychometrics: Definition, Scope, andTaxonomy3moo Psychomctris orBenchmartking Principles,4.1 Fundamental Difierences Between Psychometics and Al Benchmarking,.-..42Benchmarking with Psychometrics-Inspired PinciplesPsvchometrics for MeasuringPsvchologicalConstructsPersonalityTraits...Values5.1.3MoralityAtitudes& Opimions5.1.452MeasuningCognitiveConsructs=-...HenristicsandBiases...5.2.200000000000000000005.2.3PsychologyofLanguage..5.24Leaing andCognitive Cnoablties6 Psychometric Evaluation Methodology6.1TestFormat6.1.1StructunedTest6.12UnstroctunedTss6.2DataandTaskSources6.3PrompingStrategiesModelOutput and Scoring6.46.4.1Closed-EndedOutputandScoring64.20pen-Ended Output and Scoring+65InferenceParameters7PsvchometricValidation7.1Reliability and Consistency,7.2Validity7.2.1
7.2.27.23Criterion and Eeologpical Walvdity..7.3Standards and Recommendations;Psvchometrics for LLM Enhancement88.1TraitManipulation8.2Safety andAlignment.83Countive Enhancemen..Trends, Challenges, and Future Directions9.1PsychometricValidation92Fom Human ConstrucstoLLMConstnucs...+..93Perceivedvs.Aligned Traits.I194AmhmpomnphiaioncChleges95Exoandine Dimensions in Model Deploymen. ..oS9.7Fom Evaluaion to Enhancemen..10Conclusion
1Introduction"Whateverexis atallexist n some amount. To know it thoroughly involves knowing ts quantityas well as its quality."[ Thomdike, 1962,The advent of large language models (LLMs)represents a transformative breakthrough in Al. These systems exhibi'proficiency in naturalgapplications likencare ISingha,2024].Their increasinghow can we rigorously evaluate these Al systemsind-truth labels with human inputs.However, LLMs havetriggeredwhat traditional benchmarks cans, and cognitive biasesation have made staticnpromises the robustnesscomplex, intangible human psychologycal by t...