deepseek-r1.txt
2025
Daya Guo & the DeepSeek-R1 team
Reasoning learned through reinforcement learning
Showed in Nature that a large language model’s reasoning can be incentivised through pure reinforcement learning, without human-labelled reasoning examples. Self-reflection, verification and changing strategy emerged during training.
The people
- Daya Guo
- Dejian Yang
- Haowei Zhang
- Junxiao Song
- Peiyi Wang
- Qihao Zhu
- Runxin Xu
- Ruoyu Zhang
Filed under
Key work
DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning (opens in new tab)
Nature (DeepSeek-AI; first preprint January 2025), 2025