deepseekDeepSeek-R1: Incentivizing Reasoning Capability in LLMs viaReinforcement LearningDeepSeek-AIresearch@deepseek.comAbstractSZOZWerze ITDSoJ 148t6ZI10SZAIXIDWe introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-Rl.a model trained vialarge-scale reinforcement learning (RL) without supersmonstrates remarkable reasoning capabilitiesThrough RL, Deep Sek- R1-Zero naturall emerges with numerous powerful and intriguingreas oning behavios. However, itencouners allenges sudh as poreadablity and languagehance reasoning performance. we introduceble to OpenAl-o1-1217on reasoning tasks. To supporttheek-R1 based on Owen and LlamacepSeek-V3(%)amuaarag/AumooyAIME2024MATH-500MMLIISWE-benchVerifiedFigure 1/BenchmarkperformanceofDeepSeek-R1
Contents1Introduction生41.1Summary of Evaluation ResultsApproach DeepSeek-R1-Zero: Reinforcement Learning on the Base Model221RentorementLearning Algonrthmg,Reward ModelingTraining Template99DeepSeek-R1: ReinforcementLearningwithColdStartColStroReasoning-oriented Reinforcement LearningeRoiection Samoling and Supervised Fine-Tuning6ReinforcementLearningforallSconarioDistilation: Empower Small Models with Reasoning Capability3Experiment3.2DistilledModelEvaluation4Discussion41Disilation vs RenforcementLeaming,4.2UnsuccessfulAttempts5Conclusion, Limitations, and Future WorkAContributions and Acknowledgments
1.IntroductionIn recent years, Large Language Models (L.LMs) have been undergoing rapid iteration andevolution (Anthropic 2024; Google, 2024; Open AlI, 202-+a),progressively diminishing the gaptowards Artificial General Intelligence(AGI).st-training hasemerged as an important component of the full training pipelinetoenhance accuracy on reasoning tasks, align with social values, and adapwhile recuiring relatively minimal computational resources agains(OpenAl.2024b)seriesmodelselength of the Chain-of-ning.However, the challengeSeveral priora(Fengetal,2024;Trinhachieved general reasoning,pabilitiesdlanguagechrough rejectionprocess, takingeobtained a checkpoint referred tomodels. UsingQwen2.5-Seck-R1outperformsapplyingylargerbase models are cru-e distilledQwen and Llama(DubeyIontberforns state-of-theart opensourceIthe distilled 32B and 70B models set anew record on the re
1.1.ContributionsPost-Training: Large-Scale Reinforcement Leaming on the BaseModel*We directly apply RL to the base model without rely ing on supervised fine-tuning (SFT) asa preliminarystep. This approach allows the model to explore chain-of-thought (CoT)fosolving complex problems, resulting in the development of DeepSeek-R1-Zero. DeepSeekR1-Zero demonstrates capabilies such as self-verification, refection, and generatingo asionificant milestone for the research community. Notably, it is thefirst open research to validate that reasoning capabilities of LMs can be incentivizedpural thougpt RL, withot he need or SFT This bnakthoygh paves theway for futuneWe introduce our pipeline to develop DeepSeek-R1. The pipeline incorporates two RLstages aimed at discovering improved reason...