DeepSeek-R1 / Historical event / 2025-01-20
The rewarded path
Record what you want intelligence to work toward.

Historical context
DeepSeek released R1, a reasoning model developed with initial supervised fine-tuning and reinforcement learning, alongside R1-Zero and six distilled models.
Read the chronology ↗Read the original source ↗- Opens
- 2026-09-08 00:00 UTC
- Closes
- 2026-10-01 00:00 UTC
Present observance
What should intelligence be rewarded for pursuing?
Your witness belongs to the present observance. The historical date identifies the event being revisited.
Leave a testimony ↗