AI TECHNOLOGY

AI Technology ClassroomLesson 2 / 5

From Tokens and Attention to the KV Cache

A story-led walkthrough of tokens, Query-Key matching, Value mixing, the two memory walls, and what KV cache saves—and still has to read—during autoregressive decoding.

13 min read

Suppose you read this sentence:

Mia raises a cat. John raises a dog. The dog chews a bone. The cat eats fish. What does Mia’s pet eat?

You probably answer “fish” without much effort. You know that Mia’s pet is the cat, and that the cat eats fish. A language model reaches the same answer through a much less familiar path: the sentence is split into tokens, converted into vectors, passed through many Transformer layers, and finally turned into a ranked list of possible next tokens.

This lesson follows that path one step at a time. The goal is not to memorize every matrix operation. It is to see what information is created, what changes from layer to layer, and why the KV cache eventually becomes a hardware problem.

Why use this sentence?

The useful clue appears before the question. The model cannot look into future text for the answer; it must work with the visible context on the left. The sentence also contains a distraction: John, the dog, and the bone. A useful attention mechanism must connect Mia to the cat and then connect the cat to fish without being pulled toward the wrong animal.

This is a two-hop relationship. It gives us enough structure to explain why a Transformer needs more than a dictionary lookup or one simple match.

01 | First, split the text into tokens

A model does not receive the sentence as one block of meaning. A tokenizer cuts it into smaller units and assigns each unit a token id. Depending on the tokenizer, a token may be a word, part of a word, punctuation, or a short character sequence.

The token ids below are only a teaching example. Real ids depend on the tokenizer used by the model.

Mia#4821raises#1930cat#7715.#11John#5602raises#1930dog#7208.#12

A token id is an index, not a meaning. The number assigned to “cat” does not contain a small definition of a cat. It simply tells the model which learned embedding vector to retrieve.

Position must enter the calculation too. The two appearances of “raises” may use the same token embedding, but they occur at different places and play different roles. Models add position information through methods such as positional embeddings or RoPE, so repeated tokens are not treated as interchangeable copies.

02 | A token id becomes a vector

An embedding table turns each token id into a vector: a long list of learned numbers. Those numbers are the starting coordinates the model uses for computation.

The first embedding does not yet contain everything the token will mean in this sentence. As the vector moves through the Transformer, it absorbs information from neighboring tokens and from relationships discovered in earlier layers. Engineers often call this evolving representation a hidden state.

One token id starts with one token embedding.

After the model reads the context, the hidden state at that position is no longer just the original lookup vector.

This distinction matters later. The cache does not store a fixed dictionary entry for “cat” or “Apple.” It stores projections derived from a context-sensitive hidden state at a particular layer.

03 | The same input produces Query, Key, and Value

Inside a self-attention layer, the hidden state at each position is projected into three different vectors: Query, Key, and Value.

Q

Query: what am I looking for?

The current position uses its Query to ask which earlier information matters now.

K

Key: when should I be found?

A Key lets later Queries measure how relevant this position may be.

V

Value: what can I contribute?

A Value carries the content that will be mixed into the receiving position.

These descriptions are analogies, not literal labels stored by the model. A Key is not a category such as “animal,” and a Value is not a sentence written in plain language. They are continuous vectors learned for numerical matching and mixing.

Query, Key, and Value also change by layer. Each layer receives a different hidden state and applies its own projection matrices. There is no single K/V card that is written once at the beginning and passed unchanged through the entire model.

04 | Query compares with Keys and mixes Values

At the question position, the Query is compared with the Keys of every position it is allowed to see. After scaling, the causal mask, and softmax, those comparison scores become attention weights. The weights decide how much each Value contributes to the result.

Miahigh
cathighest
Johnlow
doglow
fishhigh

The bars are illustrative. In dense attention, visible positions are not normally divided into a few hard categories. Several positions may receive nonzero weight, just in different amounts. The causal mask only prevents the model from using tokens that have not appeared yet.

05 | Attention output is not the final answer

After the Values are mixed, the model has not directly produced the word “fish.” A better picture is that the current position has brought useful clues back from the context and written them into its working representation.

In engineering terms, the mixed result passes through an output projection and is combined with the previous representation through a residual connection. A Transformer block also contains normalization and an MLP, with exact ordering varying by architecture. Together these operations update the hidden state.

The attention output is therefore an intermediate result. It still has to pass through later blocks before the model can choose its next token.

06 | Why multiple layers help

“What does Mia’s pet eat?” requires more than one direct lookup. One layer can help the “cat” position absorb the relationship “raised by Mia.” A later layer can use that enriched representation to connect the pet in the question with the food “fish.”

Earlier layerFirst hop: Mia → cat

The cat position absorbs the relationship that Mia is its owner.

Later layerSecond hop: cat → fish

The question can now use the enriched cat representation to reach its food.

This is a conceptual picture of depth, not a recording of two specific layers.

Real models distribute information across many attention heads, MLPs, and hidden dimensions. The two hops do not have to line up neatly with two named layers.

07 | Which data actually chooses the next token?

After the final layer, the hidden state at the last position contains the context the model has assembled for this decoding step. An output layer converts that vector into logits, one score for every token in the vocabulary. Softmax and the decoding rule then turn those scores into the next-token choice.

In our example, we hope “fish” receives the strongest score. Once selected, that token is appended to the sequence, and the model begins another decoding step.

The other positions do not each print an answer. Their representations helped build the context, and their per-layer Keys and Values may be kept for the next step. The final position’s logits are the scores used to choose the next token at this moment.

08 | KV cache saves recomputation, not rereading

After “fish” is appended, the model must predict again. A wasteful implementation would run every old token through every layer again and recompute all of their Keys and Values.

The KV cache keeps those completed K/V tensors. The new token computes its own Q/K/V and uses its new Query to attend over the cached Keys and Values.

The cache removes repeated computation, but it does not remove memory traffic. Old K/V tensors still have to be read so the new Query can decide which context matters.

Work in the next decoding stepStill required?
Recompute old tokens’ K/VNo; read them from the cache
Compute the new token’s Q/K/VYes
Run the new token through all Transformer layersYes
Compare the new Query with cached KeysYes
Mix the relevant cached ValuesYes

Each round recomputes the attention weights, not the old tokens’ K/V. The Query is new, so its relationship with every Key must be evaluated again. The cached K/V tensors are the part we can reuse.

The notes are not a summary

The reading-notes analogy has one limit: ordinary notes are usually condensed, while a basic KV cache is not a summary of the prompt. It keeps a Key and a Value for every retained token, at every cached layer and KV head. Without quantization, eviction, or another compression method, its capacity grows roughly linearly with context length.

The approximate storage per token is:

2 × number of layers × number of KV heads × head dimension × bytes per element

For example, take 46 layers, a total K/V dimension of 4096, FP16 values, and no GQA or KV quantization. The result is about 0.72 MiB per token.

Context lengthApproximate KV cache size
8,000 tokens5.6 GiB
30,000 tokens21 GiB
80,000 tokens56 GiB

These are representative numbers, not a rule for every 27B model. The actual footprint depends on the layer count, number of KV heads, head dimension, numerical format, and model architecture. The example still gives useful scale: without GQA or quantization, a long-context KV cache can reach the same order of magnitude as the FP16 weights of a model with tens of billions of parameters.

Why cache K/V but not old Queries?

A Query is a question for one decoding step. Once an old position has asked its question and completed that step, its old Query is no longer needed by the new token.

Keys and Values have a different role. Future Queries still need the old Keys for matching and the old Values for retrieval. Each new step therefore creates a fresh Query while reusing the stored K/V history.

KV cache saves recomputation, not rereading.

This is one reason autoregressive decoding is often limited by memory bandwidth.

09 | There are two memory walls

Saying that LLM decoding is memory-bound can hide two different kinds of traffic.

The first wall is model weights. During one inference run, the weights are normally fixed, and different requests use the same model parameters. Each generated token still needs access to the relevant weights in every layer. Batching several requests can amortize part of that weight traffic across multiple sequences.

The second wall is the KV cache. It grows with the context and usually belongs to an individual sequence. More concurrent users and longer conversations mean more K/V data to store and read.

PropertyWeight wallKV cache wall
Main dataModel parametersK/V from previous tokens
Behavior during inferenceUsually fixedKeeps growing
Size depends onModel size and precisionContext length, concurrent sequences, and KV architecture
Across requestsOne model is sharedUsually sequence-specific; identical prefixes may be shared separately

The distinction changes what hardware can do. Traditional weight-stationary compute-in-memory works naturally with fixed weights because those weights can remain near or inside the memory array while they participate in computation. That approach does not automatically remove the capacity and read-traffic costs of a dynamic KV cache.

KV pressure needs other tools: HBM for bandwidth, MQA or GQA for fewer K/V heads, PagedAttention for memory management, and KV quantization for fewer bits per value. Both walls involve moving data, but they do not call for exactly the same solution.

10 | Why MQA and GQA reduce K/V

The second memory wall cannot be solved only by leaving data in place. Another option is to keep less K/V data for every token.

Standard multi-head attention gives each Query head its own K/V heads. Multi-Query Attention (MQA) lets many Query heads share one K/V set. Grouped-Query Attention (GQA) takes a middle path: groups of Query heads share a smaller number of K/V heads.

Fewer K/V heads shrink the cache and reduce the amount of K/V data read during autoregressive decoding. The change is not free; model designers still trade off quality, training method, throughput, and latency.

For a concrete scale, suppose an MHA configuration uses 32 K/V heads and a GQA version uses 4, with every other factor unchanged. The K/V portion of the cache falls to roughly one eighth. The earlier 0.72 MiB-per-token example would become about 0.09 MiB per token under the same assumptions.

Three easy mistakes to make

A Key is not a discrete tag such as “animal” or “person.” It is a continuous vector used for soft matching with a Query. Distracting tokens are not normally forced to exactly zero unless masking or another mechanism removes them.

The K/V for one token also does not stay fixed across all layers. Consider “I took a bite of an Apple” and “I replaced my Apple.” The same token has been shaped by two different contexts: fruit in the first sentence, a company or product in the second. By the middle and later layers, their hidden states differ, so the projected K/V vectors differ too. The cache holds “Apple in this context,” not a dictionary entry for the word Apple.

Finally, an attention output is not a finished answer. Projection, residual connections, MLPs, and later layers still update the representation. Only after the final hidden state is converted into logits does the model obtain its next-token ranking.

Lesson recap

Text is split into tokens, token ids become vectors, and position information keeps repeated tokens distinct. Query compares with Keys, attention weights decide how Values are mixed, and each layer updates the hidden state. The last position’s final hidden state is converted into logits for the next token.

The next decoding step repeats the process. KV cache avoids recomputing old K/V, while MQA and GQA reduce how much K/V must be stored. The new token still needs computation, and the old context still has to be read.

KV cache saves recomputation, not rereading.

The experiment below reduces the mechanism to four two-dimensional vectors and calculates Q·K/√2, a causal mask, softmax and the weighted V sum. Choose position 1 with Q multiplier 0: the masked weights are 0.5, 0.5, 0, 0; without the mask they are all 0.25. This checks masking and mixing, not whether a model understands Mia, the cat and fish.

References

  1. Vaswani et al., “Attention Is All You Need”: Transformer architecture, scaled dot-product attention, multi-head attention, and residual connections.
  2. Su et al., “RoFormer: Enhanced Transformer with Rotary Position Embedding”: how RoPE introduces position information into self-attention.
  3. Shazeer, “Fast Transformer Decoding: One Write-Head is All You Need”: MQA and the K/V memory-bandwidth cost of incremental decoding.
  4. Ainslie et al., “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints”: GQA grouping and the quality-speed tradeoff.
  5. Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention”: how KV cache memory management affects serving throughput.
  6. Liu et al., “KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache”: KV cache quantization for lower capacity and read traffic.
  7. Cho et al., “AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization”: a hardware and quantization approach for dynamic KV cache rather than fixed weights alone.

The example token ids and attention weights are illustrative. Tokenizers, position encodings, Transformer blocks, and inference implementations vary across models.

Exercise: calculate attention weights

Choose a query position, change Q scaling, and toggle the causal mask. Step through dot products, softmax and the weighted V sum. Bar heights show the computed weights.

Four two-dimensional vectors and projection matrices are fixed classroom inputs. No learning, tokenizer, positional encoding or semantic inference. With the causal mask on, future positions have zero weight. Disabling it is a comparison, not a valid shortcut for autoregressive generation.

Conditions for this experiment

Learning guide

AI Technology Classroom

0 / 5

Prerequisites

  • Tokens, vectors, and basic matrix multiplication

What I learned

  • Explain how position information distinguishes repeated tokens
  • Trace Query-Key matching and Value mixing
  • Separate attention output from the final hidden state
  • Explain what KV cache saves and what it still needs to read

Key terms

Open glossary →

Further reading

Knowledge check

1. How can a model distinguish the same token at two positions?
2. What can KV cache reuse from old tokens?
3. What does KV cache not eliminate?
4. Why do MQA and GQA reduce decoding memory pressure?
5. Why does weight-stationary compute-in-memory more directly help model weights than KV cache?

Thanks for reading.

Take the concept with you, not just the terminology.

#AI Technology#LLM#Token#Embedding#Position Encoding#RoPE#Transformer#Attention#Query#Key#Value#KV Cache#MQA#GQA#Inference#Memory Bandwidth#Compute-in-Memory