Timing Leaks in Shared Prefix Caches
Prefix cache hits can become visible in time to first token. Trust-group isolation and PrefixWall's selective isolation differ in their mechanisms and protection boundaries.
Sharing caches across users saves computation, but timing differences from cache hits can reveal whether a guessed input matches another request.
Input traces in response timing
A Large Language Model (LLM) computes attention Keys and Values (KV) for input tokens during prefill. A prefix cache reuses those values when the beginning of an input matches a previous request. Reducing the remaining prefill computation can change Time to First Token (TTFT).
An attacker can send guessed inputs and observe this timing difference. When users share a cache, the difference can indicate whether a guessed prefix matches another user's request. This is a timing side channel: execution time can leak information even when the response body does not contain the secret.
This path reconstructs the threat model in the PrefixWall paper. As of September 2026, this article uses the arXiv v2 dated May 20, 2026, and the vLLM v0.15.0 design documentation. Cache configuration details are limited to the documented version.
Cache keys and trust groups
The vLLM v0.15.0 design documentation describes block hashes that include the parent block hash, current tokens, and additional keys. Changes to earlier input therefore affect whether later blocks can be reused. A request's cache_salt enters the first block hash and consequently affects subsequent blocks.
Only requests with the same salt can share KV blocks. This field does not automatically identify users or organizational boundaries. The service must decide which requests belong to the same trust group.
| Control | What it determines | What it does not establish automatically |
|---|---|---|
| Tokens and parent block hash | Whether input prefixes match | Trust relationships between users |
| cache_salt | Set of requests allowed to reuse KV | Authentication and group assignment |
| Separation of input content | Placement of common and user-specific input | Cache isolation between groups |
Permission to share a public instruction is separate from permission to share personal input. Cache-key design therefore requires decisions about both shared content and the requests allowed to share it. The existing LLM serving article covers KV cache performance.
Selective isolation through ownership
A cache entry is a stored item whose ownership is tracked. PrefixWall records the user who first creates an entry as its OwnerID. An AttackFlag marks reuse of that entry across users. Ownership and this flag jointly control reuse of subsequent entries.
In Section 3.2, a different user hitting an unflagged entry reuses the current entry while setting its flag. The mechanism does not immediately reject that first hit. When a non-owner accesses a flagged entry, it also checks ownership of the next entry.
| Situation | Action | Information retained for later decisions |
|---|---|---|
| A cache miss creates an entry | Record the requester as owner | Start with the flag cleared |
| Another user hits an unflagged entry | Reuse the current entry and flag it | History of reuse across users |
| A non-owner of a flagged entry owns the next entry | Continue to the next entry | Ownership condition is satisfied |
| A non-owner of a flagged entry also does not own the next entry | Stop subsequent reuse and recompute | Restrict hit signals from the following region |
The final row recomputes the region that follows the current entry. Treating a flag as a policy that removes every user's cache hit would misrepresent the mechanism. The protection provided by selective isolation also depends on this history of reuse.
Conditions for enabling isolation
The paper separates the Detector, which restricts reuse, from the Activator, which turns isolation on or off. The Activator separates recent TTFT observations into hits and misses and compares their distributions. When those distributions overlap too little, it treats them as distinguishable by timing and enables isolation.
Section 3.3 estimates this overlap using Kernel Density Estimation (KDE). Isolation turns off when overlap is above the threshold and on when it is below. Metadata updates for ownership and flags continue even while isolation is disabled.
| Component or state | Decision or action | Scope to retain |
|---|---|---|
| Detector | Restrict subsequent reuse using ownership and flags | Protection has exceptions |
| Activator | Switch isolation based on hit and miss timing distributions | Depends on the threshold and measurement window |
| Isolation disabled | Allow reuse while continuing metadata updates | Preserve history for later activation |
Lowering the threshold can disable isolation while timing differences remain distinguishable. Load can also change the distributions, so cache efficiency alone cannot determine the threshold. Section 4 discusses these conditions alongside its protection guarantees.
Protection boundaries and performance comparisons
PrefixWall does not guarantee protection against every input guess. The guarantee in Section 4 assumes a known, nonempty prefix. It also requires either an existing flag on that prefix or a first guess that differs from the secret.
| Attack condition | Protection described in the paper |
|---|---|
| Attack on the first cache entry | Excluded from protection |
| Correct first guess after an unflagged prefix | Excluded from protection |
| Known prefix with an existing flag or an incorrect first guess | Restricts subsequent guessing signals within the stated threat model |
Performance results also need to retain the defense state used in the experiment. The main evaluation in Sections 5.1 and 5.2 keeps the Detector active throughout. Its results cannot be described as measurements of the Activator switching isolation for each case.
| Evaluation item | Reported condition or result | Comparison limit |
|---|---|---|
| Serving environment | vLLM 0.8.5 V1, A100 40GB | Distinct from the v0.15.0 documentation used here |
| Main evaluation request rate | One request per second | Detector always active |
| Cache reuse | Up to 70% higher than per-user isolation | Neither an average nor a 70-percentage-point increase |
| Latency | Up to 30% lower than per-user isolation | Not a guarantee across all loads |
These figures are reported by the paper and are not reproduced in this article. The baseline prepends a unique token for each user, so these results are not a direct measurement against vLLM's cache_salt. Section 5.3 examines the Activator threshold in a separate sensitivity experiment.
Defining the sharing boundary
Operations should first define which inputs can enter a shared cache and which requests may use that cache. Evaluating performance afterward makes both the intended sharing benefit and the exposure being reduced explicit. The following checks derive from the protection boundaries in the documentation and paper.
| Review item | What to inspect |
|---|---|
| Shared input | Whether common instructions include user-specific information |
| Request trust group | How authenticated identities map to salt assignment rules |
| Observable response timing | Whether other users can compare hit differences through repeated queries |
| Protection exclusions | Whether first-entry or correct-first-guess exceptions overlap with protected content |
| Performance conditions | Isolation method, request rate, and Detector and Activator states |
Salt-based isolation partitions the set of requests allowed to share in advance. PrefixWall selects isolation based on reuse history, so its adoption questions and protection assumptions differ. Preserving more cache reuse in the paper's evaluation does not by itself establish that a service's confidentiality requirements are met.
Summary
A shared prefix cache can expose reduced computation through differences in response timing. A salt partitions requests that can share, while PrefixWall restricts subsequent reuse through ownership and reuse history. Selective isolation does not protect against first-entry attacks or correct first guesses after unflagged prefixes. Operations should define sharing boundaries and protection exclusions before comparing performance under the same defense state.