海人工智能实验室SafeWorkFrontier AI Risk Manacement Eramework in PracticeA Risk Analysis Technical ReportSlaughai Artilcal Inteligence LaboratorytAbstractthe unprecedented risks posed by rapidly advancing artificiaSZOZUI9Z IVS]ZApES9ILOSZNIX-Df most models residing in the yellow zone, althonugh detailed threat modelinThis work reflects our currentfplease cite this work as "Slanghai Al Lab (2025)?. Fll atharship contribntion statements appear in thAuthorship section. Corespondence regarding this technical report can be sent to sa fework-frontler@pjlab,org.cn
ContentsIntroduction2ModelInformation3General Capability EvaluationsOverview32Evaluation Phameworks and Esxention Eawviromments....33CodingCapabilities.34Reasoning Capabilities......35Mathematial Capabilities ...8099936Instruction Pollowing-3.7 Knowledge Understanding++3.8AgenticCapabilities,FrontierRiskEvaluations4.1CyberOffense....42 Biologieal and Chemical Riskb....4.34.4Strategic-Deception andScheming,4.5UncontrolledAIR&:D4.7CollusionConclusions and Discussions
1IntroductionArtificial Intelligence (Al) has made signifieant progress in recent years, achieving Inuman-comparableperformance across a range of applications. These breakthroughssociated with general-purpose AI models. With the rapid development and deploy meneed a comprehensive and practicalidentification and evaluation of their underlvingaluated modelNotably newly released Aloffense, persuasion and manipulationE-T-C analysis to facilitatedine of Al frontier risks andation and assessments these critical challengesrisks while enabling beneficial AI development,2ModelInformationTo perform a comprehensive evaluatia diverse and represent ativey: Open-source and proprietary moleksare included to compare risks in diferent development and deploynent paradigmns. 3) Generational andsystemic risk (Shanghai AlLab & Concordia A1., 2025). This risk canot be asessedwhen disng the relationshitThe models inchuded in this report are selected based on their avalabilty prior to the conclusion of our evaluationperiod on July10,2025.Any me
Frontier Al Risk Management Framework in Practice: A Risk Analysis Technical ReportExperimentDescriptionRiskZone.Capture-TheFlag|CTF challenge requires the Al model to pin aces to srvers atn(CTF)snecific feld. or a feld with a fixed format within a fileastayoJoq4oAutonomousn autonomous cyber attack requires the AI model to leverage itse000oCyberAttackintrinsic reasoning, planning, and code generation capabilities toautonomously progress from vulnerability analysis to the genera-tion of a functional exploit.Biological'ahilitv to tronbleshoot biologicalOo0oProtocol and identify experimental errors, which couldDiagnosis andtechnical barriers for threat actors attemptingpon development.Biologicalnodels'knowledge of hazardous bioloricalrcapabilities,as well as their tendency toormation when inapropriately requestedChemicale00000;capabilities, as well as their tendency toowledgeen inappropriately requested.easoningPersuasion andhuman or model...