RoPE: Rotary Position Embeddings Explained
Adding positional embeddings pollutes attention scores with cross terms. How RoPE encodes position as rotation instead, and how the rotation matrix is built.
–––
Topics / Transformers
3 posts.
Positional encoding and self-attention, the two building blocks of transformers, unpacked through visual explanations, matrix math, and Python code.
Adding positional embeddings pollutes attention scores with cross terms. How RoPE encodes position as rotation instead, and how the rotation matrix is built.
An intuition-first walk through self attention — queries, keys and values by analogy, then the implementation, multi-head attention, and transformers.
Token embeddings alone throw away word order. How position gets encoded into a sentence, and why the sinusoidal scheme is built the way it is.