#attention

3 articles about attention

llmSeptember 27, 2026

Episode 1: Attention Is Just 4 Mathematical Operations

Behind Transformers and the attention mechanism there are only four simple operations: matrix multiplication, transposition, scaling and softmax. Here they are, in Rust.

llmSeptember 27, 2026

Episode 2: Q, K, V — Why 3 Matrices Instead of One?

Query, Key, Value: why real models project each word into three different spaces, and what that separation of roles changes for attention.

llmSeptember 27, 2026

Episode 3: Embeddings — Absolute at the Start, Contextual at the End

Is an embedding the word on its own or the word in context? Both, at different moments. Plus the three engineering details around attention: multi-head, masking and position encoding.