Non-Prefix KV Cache Reuse in RAG
Changing the order of retrieved documents interrupts prefix cache reuse. This article examines selective recomputation, position alignment, CoinRAG's span reuse, and the conditions needed to compare them.
Retrieving the same document again does not make its cached KV reusable when the preceding context has changed.
The Boundary of Prefix Reuse
Retrieval-Augmented Generation (RAG) supplies retrieved documents to a model. During prefill, a Large Language Model (LLM) processes the input and computes Keys and Values. The KV cache stores these values so generation does not repeat the same computation.
Prefix caching in vLLM reuses the initial token sequence shared with an earlier request. Changing the retrieved documents or their order shortens that shared prefix. Identical documents can therefore yield less reusable computation.
The example below assumes document boundaries align with token block boundaries. S is the system instruction, A, B, and C are document chunks, and Q is the question.
| Previous request | New request | Shared prefix |
|---|---|---|
| S → A → B → Q | S → A → C → Q | S → A |
| S → A → B → Q | S → B → A → Q | S |
The cache key reflects this constraint. According to vLLM's design documentation, each block hash incorporates the preceding block hash as well as its own tokens. A document identifier alone does not establish a cache match.
Position and Context Mismatches
The non-prefix reuse studied in CacheBlend assembles separately precomputed chunk caches for each request. Encoding chunk B alone cannot capture the representation formed by reading chunk A before it. Differences persist particularly in later layers, which consume the preceding layers' attention outputs.
Positions also need alignment. Models using Rotary Position Embedding (RoPE) require cached Keys to be transformed for their new positions. A comparative study of chunk caching treats position alignment and missing cross-chunk attention as separate problems.
Moving positions alone therefore does not reconstruct the representations produced by full prefill. If A names a person and B refers to that person with a pronoun, evaluation must also test that connection. Successful position alignment and answer quality require separate checks.
Selective Recomputation
CacheBlend loads the cache and recomputes KV for selected important tokens. Those tokens attend to the current request context while others retain cached values, reducing computation. This does not guarantee that every token representation matches full prefill.
CacheClip uses an auxiliary model to select tokens and updates neighboring groups together. Its shared prefix reduces duplicated attention sinks, where early tokens attract attention irrespective of content.
| CacheClip component | Purpose |
|---|---|
| Query-specific auxiliary model scores | Select recomputation positions |
| Retain one shared prefix | Reduce duplicated attention sinks |
| Group tokens with a sliding window | Preserve local context during partial updates |
Both selection and recomputation take time. The fraction of reused tokens alone therefore misses any additional selection cost.
Position Alignment and Attention Fusion
LazyAttention targets the cost of materializing KV copies for positional adjustment. It applies positional encoding inside attention kernels, letting one physical cache serve different logical positions. This reduces relocation costs; it does not automatically recover semantic dependencies between documents.
Decoupled Attention Fusion (DAF) separates inter-document correction from question processing into distinct attention paths. It then combines their output states, using dense computation patterns compatible with FlashAttention kernels. Reducing token count and executing the remaining work through efficient kernels are separate design decisions.
The diagram summarizes DAF's computation paths. LazyAttention addresses position handling, whereas DAF changes how correction executes. Their acceleration factors cannot establish a ranking without matching the comparison conditions.
Token Span Reuse in CoinRAG
A nugget in CoinRAG is a contiguous token span within a source chunk. Slicing its KV from a precomputed full-chunk cache retains preceding context that isolated span encoding would lose.
Online processing retrieves chunks, selects nuggets within them, and aligns positions. The paper also uses fine-tuning adapted to this composition, so its headline performance cannot be attributed to cache slicing alone.
Conditions for Performance Comparisons
The linked primary sources were checked for paper identifiers and descriptions as of 2026-09. Measurement boundaries and baselines differ across papers. Those conditions need alignment before speedup factors can be compared.
| Study | Reported result | Relevant conditions |
|---|---|---|
| CacheClip Table 6 | Prefill: 5.641 s → 1.695 s | 16K input, 20% recomputation, Qwen2.5-14B-Instruct, L20 Graphics Processing Unit (GPU) and auxiliary model on a Central Processing Unit (CPU) |
| LazyAttention | 1.37× improvement in first-token latency | Against Block-Attention under a skewed document request distribution |
Time-to-First-Token (TTFT) measures the interval from a request to its first output token. A prefill-only measurement may omit retrieval or queueing. The table therefore cannot directly predict user-visible latency.
F1 is the harmonic mean of precision and recall computed from token overlap between the reference and generated answer. The chunk caching comparison uses an adjusted F1 that evaluates only questions where full prefill receives a nonzero score.
Otherwise, questions the model cannot answer can mask approximation losses in the average. Evaluation-set differences therefore also matter when comparing quality-preservation claims.
Measurements Before Adoption
The following table proposes checks for comparing these techniques in a service. Hold the model, questions, and retrieval results fixed while changing the cache path. Measure cache hits and misses separately to identify the source of a difference.
| Observation | Decision it informs |
|---|---|
| Repeated chunks versus prefix hits | Is non-prefix reuse needed? |
| Retrieval, queueing, KV loading, correction, and generation time | Is the targeted computation the bottleneck? |
| Accuracy on questions connecting evidence across documents | Is the cross-chunk context loss acceptable? |
| Cache rebuilding after source or model changes | Can precomputation costs be recovered? |
For example, improving prefill alone cannot resolve a workload dominated by retrieval time. Conversely, repeated chunks appearing in different orders justify an experiment. The gap between chunk recurrence and prefix hit rates provides that evidence.
Cache validity also belongs in the experiment. Version source text, tokenizers, and model weights so incompatible KV states are not mixed, and compare against full prefill within one version. This article does not reproduce the papers' serving benchmarks; adoption requires measurements in the target service.
Summary
Retrieving the same document and reusing the same KV have different requirements. Non-prefix reuse must address both positional alignment and context loss, while correction itself consumes computation. Decide whether to adopt it by measuring chunk recurrence, total request latency, and answer quality across documents together.