LLM Post-Training: A Deep Dive into ReasoningLarge Language ModelsKomal Kumar;, Tajamul Ashraf*, Omkar Thawakar, Rao Muhammad Anwer, Hisham CholaklMubarak Shah, Ming-Hsuan Yang, Phillip H.S. Torr, Fahad Shahbaz Khan, Salman Khan,pstract-l arge Language Models (ILMa) have transformed the natural language processing landscape and rought to life diverseadSZ0Z9018Z【10SOIAIZEICZOSZADdobgies,anaying theiroeinrefining L.Ms, and infereance-time reasonin, and outine fture reachrections.We alsoprovideapublic repoard Modeling, Test-time ScalingIntroductionMontemporArY Large LanguageAlgorlithmiccategorization/ remarkable canabilitiesencompassing noAlgorithms&QwenansweringLLMs00ningf post-training approadhes for LMs). categorized into Fine-tuming, Reintoarcement Leana, and Test-time Scaling methods. We sumarize the lkeychminues used in recent LLA models, such as GPT-4 (30)LaMA3.3[13], and Deepseek R1 {40]ts while still stumbling on relativelymbolic reasoning that manipu-LMs operate in an implicit andanner 50,42, 51). For the scope of this work,
non-stationary objectives., ensuring responses remainat and aligned with userexpectationsxtends beyondconyast action spaces,of-domain generalization, underscoringthe trade-obetween specificity and versatilityand ea viron mental impact, especially when search techminrather than during training (8Ensuring accessibility and feasiblity is esential to maintain improving performance but risking overitting.high compute costs, and reduced generalization.Test-time scaling enhances the adaptabilityb) Reinforcement Learning in LLMs: In conventional RL of LLMs by dynamnically adjusting computationalan agent interacts with a structured environment, takingete actions to transition between states while maximize rewards[731. RLdomainsuch as roboticsature well-defined state-and ckar objectivs 74,75).i.LAsdites1.1PriorSurveysInstead of a finite action set, LLMs select tokensRecent surveys on RL and LMs provide valuable insightfrom a vast vocabulary, and their evolving state comprises anbut often focus on specitic aspects, leaving key post-training
it[108.8].However,deMs struggle to maint ainance overlong sequences. Ad-necessitates a structured approachch naturally aligns withRLhere each tchim a Markov Decision Process(MDP)[109).In this sets rpresentsthe xypence of okens guneratedthe action a is the next token,and a reward R(s.,.tthe quality of the outpnt. An LLM'spolicy ae iotimized to maximize the expected return:J(Te#R(St.at)mt foctor that determines how stronglyinterconnected optimizationstrategiesniques,adaptability needed for refiing decision-making in interactiDIdirections in cong-term objectives, models can dynamically adjust theing astructured framework for real-world applicationsy:1uns cxhthit emergent abilties due to salewhile RL refines and aligns them for betterrcaBackgroundsoning and interactionThe LLMs haveI wing Maximun Likelod Eastimatin (MLE)2.1RLbasedSequentialReasoning106.3. 107l, wtich maximizes the probabili...