Many languages, many sensesLesson 4 of 8

Long context and mixture of experts

Attention links grow with the square of the context: 8 tokens make 64, a million make 10¹². Experts let a huge model run like a small one.

Locked

Unlock the rest of Deep Learning: from one neuron to a transformer.

  • Every lesson, every resource, unlocked instantly.
  • Track progress and pick up where you left off.
  • Free preview lessons stay readable from the outline.

From

BDT 0

Enrol now