Topics / Transformers

Transformers

3 posts.

Positional encoding and self-attention, the two building blocks of transformers, unpacked through visual explanations, matrix math, and Python code.

·13 min

Self Attention

An intuition-first walk through self attention — queries, keys and values by analogy, then the implementation, multi-head attention, and transformers.

·11 min

Positional Embeddings

Token embeddings alone throw away word order. How position gets encoded into a sentence, and why the sinusoidal scheme is built the way it is.