State ofAl: ChinaArtificial AnalysisQ22025Highlights ReportFull report available to PremiumAccess subscribers
Artificial Analysis is a leading, and independent Al benchmarking andinsights provider. We support engineers and companiesto understandAcapabilties and make aritical decisions about their Al strategy,Our data, insights and publications are grounded in our comprehensivebenchmarking of AI technologies and use cases. This includes everythingfrom hourly performance testing of lanquagemodel APls tomillions ofvotes In our crowd-sourcedarenasOur public website artificialanalysis. ai, is widely referenced by companiesleading innovation in Al. To discuss this report, our publications, or ourservices, please get in touch at contact(@artificialanalysis. a
China's leading Allabs are now closer than ever to US leaders, withthe lead decreasing from more than a year to less than three monthsUS& China:Frontier LanguageModelIntelligence, OverTimeCommentaryArtificial Analysis Inteligence Index incorporates 7 evaluations: MMLU-Pro, GPQA Diamand, Humanity's Last ExamTheperformancegapbetweenUSLiveCode Bench,SciCoda,AIME, MATH-500)andChinesefrontiermodelssincethe releaseofChatOPTin2022hasremalned persistent,butisasUnitedStatesChinareasoning(high).wAo3-miri(high).Open0528,May*25)2025)modelleadstheChineseAlxopul@3ue?nows;s/jnuyepunyo1.OpenAl60·re leasedbyUSAlLabs50-DeepSeekandAlibabahaveGPT-40,OpenAprimarlydriventheChineseGPT-4.Oper(Jan.25)frontler,while advancesintheUSQwen1.5ChatpSeckV330-GPT-3.5Turbo.OpenAlhipuAQ2V20-QwenChat78,110BAlbabe10-Now22Jan'23Mar23May23Sap 23Now 23Jin 24Mar 24May24Ju24Sop24Now24Jon25Mr25May25J25八ArtificialAnalysisReleaseDatemporable msultsSource:Artis Inteligence Index
The Chinese open weights frontier surpassed the US in November 2024with Alibaba's release of QwQ 32B Preview. R1 consolidated this leadUS& China:Open Weights FrontierLanguage Model Intelligence, Over TimeCommentaryArtificial Analysis Inteligence Index incorporates 7 evaluations: MMLU-Pro, GPQA Diamand, Humanity's Last Exam,TheChineseopenweightstrontierLvecCodeBnch, SciCode,AIME,MATH-500surpassedthe US in November2024withthe releaseofQwQ32BPreview (overtaking Meta'sLlama3.1@China405B)TheopenweightsleadershipofChineseAllabsisreflectiveofthe8approachofthe topChineseAlQwen2.588abstooftenreleasethewelghtsoft4058,theIrflagshipmodels.This4contra sts with the topUS Allabser49B,NVIwhich generallydo not release thewaights of theirleading models, a.g.,na33708RekaRash3OpenAl,Anthropic andGoogleChina's DeepSeekR1(JanuaryR2025)wasthefirstopenwelghts15reasoningmodeltobecompetitiveQwenChat14B,rith OpenAl'so1LLM67BMar'23May23Nov'23May'24JJ24DeepSeek'sR10528(May2025)isAArtificialAnalysisRelease Datemodel currently availablead an company dl aims and com parable resultsSource:Artinalysis IntetigenceIndex
Leading Chinese Al labs DeepSeek and Alibaba have steadilyreleased new models, with DeepSeek taking the lead in late 2024Leading Chinese AlLabs:Language Model Intelligence, Over TimeCommentaryAtificial Analysis Intelligence Ind ex incorporates 7 av alua tions: ...