Large Language Model Safety: A Holistic SurveyDan Shi'*.TianhaoShen'*,Yufei Huang', Zhigen Li-2Yomgyi Leng?, Renran Iin', Chuang LAnm', Xinwed Wir,Zishan Guo',Linhao Yu'. Ling Shi', Bojian Jiang!8, DeyiXiong'TJUNLP Lab. Tianin University2PingAnTechnology,3Du Xiaoman Finance7Z0Z0QGEZIVSO]14989L1ZI7ZAIXDAbstractThe rapid development and deploy ment of large language mnodels (LLMs) haveintroduced a new frontier in artificial inteligence, marked by unprecedented capabilities in natural language understanding and generation. Howevwer, the increasing integration of these models into critical applications raisessubstantiasafety concerns, necessitating a thorough examination of their potential risksand associated mitigation strategies.This survev provides a comprehensivon theseinterpretabilitv in enhancingLLM safety, the technoldcompanies and institutes for LLM safety, and AIgovernance alimed at LLMsafety with discussions on international cooperation. policv probosals andprospective regulatorv directioneOurfindings underscore thea proactive. multifaceted approach toLAM safetv. emphasizing the intration of technical solutions. ethical considas a foundationalreso, industry practitioners, andsand opportunities associatedwith the safe integration of LLMs intoetv Ultimatelv it seeksto contribnute to the safe and benelicial develoomef lMs aligning with the overarchinggoal of harnessing Al for societal advancement and well-being. A curated lisof related papers has been publiclv avaiable at a Git Hub repositorv.*EqualcontributiontCoresponding author.Correspondence to: {shidan, thshen, dyxIhttps://github.com/tjunlp-lab/Awesome-LLM-Safety-Papers
Contents下991 Introduction1.11.2Paper and Source Selection.1.3Related Work2Taxonomy Basic Areas of LLM Safety...2.12.2RelatedAreastoLLMSafety.3Value Misalignment3.1SocialBias.... Defnition and Safety Impact....3.1.13.1.2Social Bias in the LLM Lifecvcle3.1.3Methods for Mitigating Social Bias3.1.4Evaluation..3.1.5Future Directions,3.2Privacy3.2.1Preliminaries3.2.2Sources and Channels of Privacy Leakage...323Privacy Protection Methods....3.3Toxicity3.3.1Definition and Saiety Impact.33.2Methods for Mitigating lovicity.3.3.3Evaluation .I3.4Ethics and Morality.3.4.1Definition3.4.2Safetv Issues Related to Ethics and Morality3.4.3Methodsfor Mitioatino L.L.MAmorality344Evalation4Robustness toAttack
4.1Jailbreaking....411Black-boxAttacks4.1.2White-boxAttacks4.2RedTeaming4.2.1ManualRedTeaming.4.2.2AutomatedRedTeaming4.2.3Evaluation..4.3Defense4.3.1ExternalSafeguard4.3.2Internal Protection;Misuse15.1Weaponization..5.1.1Risks of Misuse in Weapons Acquisition5.1.2Mitication Methods for Weaponized Misus5.1.3Evaluation52MisinformationCampaigns...5.2.2Social Media Manipulations;5.2.3Risks to Public Health Information5.2.4MitigationMethodsfor theSoread of Misinformation5.3Deepfakes....Malicious Applications of Deplakes.5.3.1532Methods for Mitigating Deeplakes-..5.4Future Directions5.4.1WeaponizationMisinformation Campaigns .....5.4.25.4.3Deepfakes5.4.4Comorehensive Evaluation5款6AutonomousAIRisks6.1
招路高药...