电脑桌面
添加运营动脉到电脑桌面
安装后可以在桌面快捷访问

OpenAI:《OpenAIo1大模型》英文技术报告会员免费优质

OpenAI:《OpenAIo1大模型》英文技术报告_第1页
1/43
OpenAI:《OpenAIo1大模型》英文技术报告_第2页
2/43
OpenAI:《OpenAIo1大模型》英文技术报告_第3页
3/43
OpenAI:《OpenAIo1大模型》英文技术报告_第4页
4/43
OpenAI:《OpenAIo1大模型》英文技术报告_第5页
5/43
OpenAI:《OpenAIo1大模型》英文技术报告_第6页
6/43
OpenAI:《OpenAIo1大模型》英文技术报告_第7页
7/43
OpenAI:《OpenAIo1大模型》英文技术报告_第8页
8/43
OpenAI:《OpenAIo1大模型》英文技术报告_第9页
9/43
OpenAI:《OpenAIo1大模型》英文技术报告_第10页
10/43
OpenAI o1 System CardOpenAISept 12, 20241IntroductionThe o1 model series is trained with large-scale reinforcement learning to reason using chain ofthought. These advanced reasoning capabilities provide new avenues for improving the safety androbustness of our models. In particular, our models can reason about our safety policies in contextwhen responding to potentially unsafe prompts. This leads to state-of-the-art performance oncertain benchmarks for risks such as generating illicit advice, choosing stereotyped responses,and succumbing to known jailbreaks. Training models to incorporate a chain of thought beforeanswering has the potential to unlock substantial benefits, while also increasing potential risks thatstem from heightened intelligence. Our results underscore the need for building robust alignmentmethods, extensively stress-testing their efficacy, and maintaining meticulous risk managementprotocols. This report outlines the safety work carried out for the OpenAI o1-preview and OpenAIo1-mini models, including safety evaluations, external red teaming, and Preparedness Frameworkevaluations.2Model data and trainingThe o1 large language model family is trained with reinforcement learning to perform complexreasoning. o1 thinks before it answers—it can produce a long chain of thought before respondingto the user. OpenAI o1-preview is the early version of this model, while OpenAI o1-mini isa faster version of this model that is particularly effective at coding. Through training, themodels learn to refine their thinking process, try different strategies, and recognize their mistakes.Reasoning allows o1 models to follow specific guidelines and model policies we’ve set, ensuringthey act in line with our safety expectations. This means they are better at providing helpfulanswers and resisting attempts to bypass safety rules, to avoid producing unsafe or inappropriatecontent. o1-preview is state-of-the-art (SOTA) on various evaluations spanning coding, math,and known jailbreaks benchmarks [1, 2, 3, 4].The two models were pre-trained on diverse datasets, including a mix of publicly available data,proprietary data accessed through partnerships, and custom datasets developed in-house, whichcollectively contribute to the models’ robust reasoning and conversational capabilities.Select Public Data: Both models were trained on a variety of publicly available datasets,including web data and open-source datasets. Key components include reasoning data andscientific literature. This ensures that the models are well-versed in both general knowledgeand technical topics, enhancing their ability to perform complex reasoning tasks.1

Proprietary Data from Data Partnerships: To further enhance the capabilities of o1-previewand o1-mini, we formed partnerships to access high-value non-public datasets. These propri-etary data sources include paywalled content, specialized archives, and other domain-specificdatasets...

1、当您付费下载文档后,您只拥有了使用权限,并不意味着购买了版权,文档只能用于自身使用,不得用于其他商业用途(如 [转卖]进行直接盈利或[编辑后售卖]进行间接盈利)。
2、本站所有内容均由合作方或网友上传,本站不对文档的完整性、权威性及其观点立场正确性做任何保证或承诺!文档内容仅供研究参考,付费前请自行鉴别。
3、如文档内容存在违规,或者侵犯商业秘密、侵犯著作权等,请点击“违规举报”。

查找下载文件的指引

一、电脑端

- Windows 系统:按下键盘快捷键 `Ctrl + J`,即可打开下载列表。  

- Mac 系统:按下键盘快捷键 `⌘ + J`,即可打开下载列表。  

二、手机端

1. 打开手机浏览器,点击浏览器右下角的 “≡”(或“更多”)图标。  

2. 在弹出的菜单中找到并点击 “下载内容”(或类似选项),即可查看已下载的文件。  

提示:不同浏览器界面略有差异,若未找到“下载”入口,可尝试在浏览器设置中搜索“下载”关键词。


OpenAI:《OpenAIo1大模型》英文技术报告

确认删除?
会员
教程
收藏
足迹
联系
  • 站长微信
回到顶部