deepseek-r1.txt

2025

Daya Guo & the DeepSeek-R1 team

Reasoning learned through reinforcement learning

Showed in Nature that a large language model’s reasoning can be incentivised through pure reinforcement learning, without human-labelled reasoning examples. Self-reflection, verification and changing strategy emerged during training.

The people

  • Daya Guo
  • Dejian Yang
  • Haowei Zhang
  • Junxiao Song
  • Peiyi Wang
  • Qihao Zhu
  • Runxin Xu
  • Ruoyu Zhang

Filed under

Key work

DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning (opens in new tab)

Nature (DeepSeek-AI; first preprint January 2025), 2025

Sources