RoPE: Rotary Position Embeddings Explained
Adding positional embeddings pollutes attention scores with cross terms. How RoPE encodes position as rotation instead, and how the rotation matrix is built.
–––
Topics / Attention
2 posts.
Queries, keys, values, and scaled dot-product attention — the intuition first, then matrix diagrams and a PyTorch implementation.
Adding positional embeddings pollutes attention scores with cross terms. How RoPE encodes position as rotation instead, and how the rotation matrix is built.
An intuition-first walk through self attention — queries, keys and values by analogy, then the implementation, multi-head attention, and transformers.