Memory and attentionLesson 4 of 9

Scaled dot-product attention

Match, scale by √d, weigh, mix. Dividing by √4 = 2 keeps long vectors from turning softmax into a hard pick.

Locked

Unlock the rest of Deep Learning: from one neuron to a transformer.

  • Every lesson, every resource, unlocked instantly.
  • Track progress and pick up where you left off.
  • Free preview lessons stay readable from the outline.

From

BDT 0

Enrol now