RoPE: Rotary Position Embeddings Explained
Adding positional embeddings pollutes attention scores with cross terms. How RoPE encodes position as rotation instead, and how the rotation matrix is built.
–––
Topics / Embeddings
3 posts.
How token IDs become meaningful vectors and how positional encoding adds word-order information, with training examples and diagrams.
Adding positional embeddings pollutes attention scores with cross terms. How RoPE encodes position as rotation instead, and how the rotation matrix is built.
Token embeddings alone throw away word order. How position gets encoded into a sentence, and why the sinusoidal scheme is built the way it is.
How tokens turn into meaningful vectors: the core idea behind embeddings, how the training data is prepared, and a working implementation.