From raw model to assistantLesson 5 of 10

Thinking longer: reasoning and test-time compute

Reward only correct answers and the worked answer wins: P(correct) 0.905 after 20 updates, average length up from 12.6 to 22.1 tokens.

Locked

Unlock the rest of Deep Learning: from one neuron to a transformer.

  • Every lesson, every resource, unlocked instantly.
  • Track progress and pick up where you left off.
  • Free preview lessons stay readable from the outline.

From

BDT 0

Enrol now