From raw model to assistantLesson 3 of 10

Learning from preferences, and reward hacking

A reward model that likes politeness scores a padded wrong answer 1.56. After ten updates the engine flags the hack: reward up, correctness down.

Locked

Unlock the rest of Deep Learning: from one neuron to a transformer.

  • Every lesson, every resource, unlocked instantly.
  • Track progress and pick up where you left off.
  • Free preview lessons stay readable from the outline.

From

BDT 0

Enrol now