AlAgents Beyond ChatGPTLLMLLMLLMZhou(Jo)YuColumbia University&ArklexAl
Who supportsAlAgents?Bill GatesAgents arebringing about the biggestCurrent agentsare justthinwrappers aroundLLMs.revolution in computing since we went fromtyping commands to tapping on icons.AutoregressiveLLMscanAndrewNgneverreason orplan.Ithink Al agentic workflows will drivemassive AI progress this year.Auto-GPT'slimitationsin...revealthatitis farfrombeinga practicalSam Altman2025 is when agents will work.Slides adapted fromYuSu
What are Al Agents?Perception: Multimodal inputs includingtext, image, audio, video, touch, etcAgentSensorsPerceptsPlanning (InnerMonologue);ReasoningChain-of-Thought reasoning over tokensEnvironmentthat powered by LLMsInnerMonologueReflection: meta-reasoning in every stopActions: function/tool calling, embodiedactions.ActuatorsAdapted fromRussell &Norvig(2020)
Al Agent Deployment ConsiderationPHASE1RESEARCHPHASE2SCALINGPHASE3INNOVATINGLevel1Level2Level3Level4Level5"JustWannaChat""YourWorkAnsistant""Agent-as-a-Servicel"AutonomousAgentaHumanholdmyboerPROTOAGIAnLLMLIMsascorecompoLLMsascorecompcompletingvarloustasksSlide:AlexWang@ScaleA
18OverviewModel self-improvement with LL.Ms (Yu etal, NMACL.204, Outstanding paper,2. Eliciting stronger model ability via tree search (Yu et al,EMNLP20233. Al agent self-improvement via tree search (Yueta,ICLR 2025)
Background: In-Context Self-ImprovementInput:Q:Calculate(4*1)-(2*3)=?Interactive Demonstrations, NAACL2024, Outstandingpaper
2Background: In-Context Selt-ImprovementInput:Q:Calculate(4*1)-(2*3)=?1u8noqp-jo-ueupQ:Calculate1+2=?1dwo.d jous-majAns:3Q;Calculate(4*-1)+(2*3)=?O:Calculate...Ans:..Let'sthink stepbystep:Q;Calculate(4*1)-(2*3)=?Step1:(4*1)-(2*3)=4-6Step2:4-6=-2Ans:-2Ans:-2
3Background: In-Context Selt-ImprovementInput:Q;Calculate(4*1)-(2*3)=?Self-Improvement Prompting(Madaan, et al,2023)Step1:(4*1)-(2*3)=4-6Step2:4-6=-3Ans:-3Madaan.A et al (2023) Self-Refine: Iterative Refinement wth Self-Feedback
Background: In-Context Selt-ImprovementInput:Q:Calculate(4*1)-(2*3)=?Self-Improvement Prompting,(Madaan, et al,2023)Step1:(4*1)-(2*3)=4-6Step2:4-6=-3promptAns:-3feedbackIn step2 the part"4-6=-3"isincorrect. This is because...promptupdateStep1:(4*1)-(2*3)=4-6Step 2:4-6=-2Ans:-2Madaan.A et al (2023) Self-Refine: Iterative Refinement wth Self-Feedback
5Background: In-Context Selt-ImprovementInput:Q;Calculate(4*1)-(2*3)=?Self-Improvement Prompting,(Madaan, et al,2023)Step1:(4*1)-(2*3)=4-6Step2:4-6=-3promptAns:-3feedbackIn step2 the part"4-6=-3"isincorrect. This is because...promptpromptfeedbackupdateStep1:(4*1)-(2*3)=4-6Step 2:4-6=-2Ans:-2Madaan.A et al (2023) Self-Refine: Iterative Refinement wth Self-Feedback
Background: In-Context Self-ImprovementMultistepArithmeticAcc,=31.3Codex(175B)(%)KeJn23yv5LLaMa(7B)+2.0Problem 1 small LM can not self-improve via promptingoAcC.=16.8-5.25.1LogicalDeductionAcc.=81.0Codex(1758)(%)en3yv5noLLaMa(7B)-2.1Acc.=45.8-4.1+S1BackgroundMotivationApproachExperiments
7Background: In-Context Self-ImprovementMultistepArithmeticAcc,=31.3Codex(175B)(%)KeJn23yy5LLaMa(7B)+2.0Problem 1: small LM can not self-improve via prompt...