GPAI.WIKI
news/N-0002·

DeepSeek-R1 shows reasoning emerging from reinforcement learning alone

Reports that long-form reasoning behaviour can be elicited by reinforcement learning on outcome rewards without a supervised fine-tuning stage, with weights released. The claim that matters is the negative one: the supervised reasoning traces widely assumed to be necessary appear not to be.

The result bears directly on credit assignment: outcome-only rewards over long generated chains are exactly the regime where assignment is supposed to be hardest, and it worked well enough to matter. Whether that survives at horizons longer than a single response is open — see OP-001.

reactions
no reactions yet
react on GitHub