Enterprise assistants repeatedly send policies, tool definitions and long reference documents. Prompt caching may reduce the computation and cost of these repeated inputs, but a hit rate alone can make a billing discount look like additional service capacity. This article focuses on reusing computational state for exact prefixes and proposes an acceptance method that checks effectiveness, cost and access boundaries together. All numerical examples are constructed illustrations, not IDENIFE measurements or provider prices. Technical sources were checked on October 11, 2026.
1. Identify the work being reused
Prompt caching and semantic answer caching solve different problems. The former reuses intermediate state for a processed prefix while still generating an answer for the current input. The latter may return an earlier answer directly, requiring separate decisions about question equivalence, freshness and access rights. Combining them under one cache-hit metric hides different correctness risks.
vLLM automatic prefix caching primarily saves shared-prefix processing during prefill; it does not directly remove the subsequent work of decoding new tokens.[1] Long reference material with short answers is therefore often a more promising test case than short questions requiring lengthy reports. An earlier first token does not imply a proportional reduction in completion time. Reduced queueing across a shared service is a possible indirect benefit that needs separate measurement.
Acceptance should distinguish three levels. Repeated text in application logs establishes an opportunity. Cache reads reported by the service establish actual hits. Faster and cheaper requests of acceptable quality under the target load establish business value. None substitutes for the others. If KV-cache quantization is also enabled, treat precision and prefix reuse as separate experimental variables. The method for validating low-bit LLM quantization can help establish independent baselines.
2. Treat the prompt as a versioned prefix structure
A support assistant might place stable role constraints and tool definitions first, followed by authorized, versioned reference material, and finally the current question and necessary dynamic state. A changing timestamp or request identifier at the beginning can prevent reuse of substantial material that follows. However, ordering can change model behaviour. Reordering therefore requires a new quality evaluation, rather than optimization solely for cacheable length.
vLLM identifies blocks using the parent-block hash, tokens in the current block and additional identity information. Only full blocks are cached, and allocation pressure can evict cached blocks.[2] A first half that looks similar is consequently insufficient. Fix the model, tokenizer, chat template, tool schemas and document version during acceptance, and record a fingerprint of the serialized prefix. Preserve the corresponding media identity for multimodal inputs.
We recommend a prefix inventory documenting the owner, version, effective scope, permitted sharing group and expected update frequency of each segment. Public rules and customer-specific documents must not acquire a common permission scope merely because their text is identical. For changing business data, freshness comes before reuse. A high hit rate against obsolete inventory or approval rules is a failed optimization.
During a version transition, define when traffic using the old version must end and measure the new version's cold-start window. Frequent releases, load-balancing changes and many distinct long prefixes can repeatedly rebuild the cache. Observe continuous business periods, rather than repeatedly submitting one request to the same process and reporting an impressive but unsustainable result.
3. Separate request hits, token coverage and billing
Record at least request hit rate H and token-read coverage R. H is the number of requests with positive cache reads divided by all requests. R is the sum of cache-read tokens divided by the sum of total input tokens. Ten requests can all reuse a very short beginning, producing H of 100% but low R. Aggregate by summing before dividing; an average of per-request coverage ratios gives short and long requests equal weight.
Interpret usage fields by interface. OpenAI documents total input and cached_tokens and cache_write_tokens details. Claude separately reports cache_read_input_tokens, cache_creation_input_tokens and uncached input_tokens.[3][4] Do not mistake the latter interface's input_tokens for the whole input, or add detail fields again to an existing total. Preserve raw usage and the applicable billing version before converting them to a common internal representation.
A hosted API's read count is provider usage evidence, not direct visibility into physical reuse of each GPU block. A self-hosted service can add block-hit, actual-prefill-token and queueing metrics, but must identify where they are counted and their denominators. Cross-platform comparisons require consistent definitions of full input, cache-eligible prefix and actual reads before comparing hit rates.
Stratify production statistics by task type, prefix version and tenant. Retain costs for successful, failed, retried and cancelled requests separately. Counting only the final successful attempt hides repeated cache writes and retry costs. Logs containing cache keys, source content or identities should follow existing access controls so observability does not create unnecessary copies of sensitive information.
4. Use four workload groups to test first-token and tail latency
The first group measures a cold cache. In an isolated test environment, use a new prefix known not to have been reused, or a clearing mechanism genuinely supported by the deployment. Record how coldness was established. Do not indiscriminately clear a production shared cache, or assume that a request sent for the first time by a test script has never appeared anywhere in the system.
The second group measures a stable warm cache: complete effective warming, then reuse the prefix while changing the final question. The third measures realistic mixed traffic, replaying prefix frequency, input and output lengths and arrival intervals, including one-off requests. The fourth tests invalidation and recovery: update the document version, exceed the lifetime or introduce capacity pressure in a self-hosted environment, then inspect latency and recovery when hits fall.
Client first-token latency can be defined from request transmission to receipt of the first non-empty response-content fragment. Total latency ends when the complete response arrives. Streaming metadata, heartbeats and empty fragments are not response content. For interfaces exposing separate reasoning events, state whether those count. Record server queueing and prefill time separately so network or gateway variation is not attributed to caching.
Keep the model, sampling parameters, input set, output limits and arrival pattern fixed between control and caching groups. Interleave experiments where practical to reduce temporal bias. Do not subtract an idle warm-cache result from a peak-load uncached result. Report throughput, error rate and P50, P95 and P99 latency, along with the number of requests in each group. Small samples do not justify presenting a few extreme observations as a stable P99 estimate.
Acceptance thresholds should follow business budgets. For an interactive assistant with a strict first-response requirement, improved warm first-token latency, no regression in mixed-traffic tails and cold requests still meeting timeouts can be separate conditions. Actual time thresholds require measurement in the relevant setting. Return any capacity claim to enterprise LLM service-capacity validation and repeat the same arrival load; do not multiply a single-request speedup into a concurrency promise.
5. Include writes in the economic calculation
Consider an illustrative calculation only. The shared prefix P contains 8,000 tokens and each dynamic suffix S contains 500. Ordinary input costs one unit per thousand tokens, the cache-write multiplier w is 1.4 and the read multiplier r is 0.2. After one write, nine further requests fully reuse the prefix within a valid reuse window. Output length and price are held constant; output, network and operational costs are excluded for now.
Without caching, ten requests cost 10 × (8 + 0.5) = 85 input units. With caching, the cost is 8 × 1.4 + 9 × 8 × 0.2 + 10 × 0.5 = 30.6, a 64% reduction in the input component. Here the write multiplier is the total rate for those tokens, not a surcharge to which ordinary input is added again. If every request arrives too late and writes anew, the same assumptions yield 10 × (8 × 1.4 + 0.5) = 117, exceeding 85.
In the simplified window with one write and every later request hitting, the number of subsequent hits must exceed (w − 1) / (1 − r) for the prefix to cost less than ordinary input throughout. This assumes 0 ≤ r < 1 and ignores other costs. In practice, settle actual read, write and uncached tokens for each request, then include output, failure and retry expenses. Partial hits or mixed lifetimes invalidate this simplified expression.
Minimum eligible lengths, write prices, retention periods and invalidation conditions vary across models and platforms. Claude's documentation, for example, distinguishes cache writes, reads and different lifetimes.[4] Preserve the interface and price schedule applicable on the test date for procurement comparisons. For self-hosting, account for memory consumption, throughput changes and operations; an API discount multiplier is not a measure of GPU savings.
6. Make cache sharing subordinate to access boundaries
vLLM's security documentation states that a shared process does not provide full tenant isolation. cache_salt can restrict prefix reuse but is not itself an isolation boundary.[5] Acceleration must therefore never justify bypassing authorization. Verify identity and document access before inference; where strong tenant isolation is required, implement architectural boundaries such as dedicated instances.
A trusted gateway can inject an unpredictable cache salt based on authenticated tenant identity and overwrite any client-supplied value, preventing users from choosing another party's sharing scope. When document permissions change, update retrieval and business access controls. Salt rotation affects future reuse; it does not prove that all existing copies were deleted. Deletion requirements need their own implementation and evidence.
In an authorized test environment, construct two tenants with identical material that must not be shared. Verify distinct cache identities and rejection of unauthorized requests before inference. Then check that reuse works within a permitted scope. Similar latency is not proof of isolation; combine gateway configuration, identity mappings and server observations. For a hosted API, also verify current contractual and retention terms rather than importing a self-hosted framework's boundaries.
7. Deliver a reproducible record of cache value
We recommend six deliverables: a versioned prefix inventory; requests covering cold, warm, mixed and invalidation scenarios; raw usage reconciled against billing; stratified latency and error results; quality regression results; and sharing-boundary and rollback records. Together they establish what could be reused, what actually was reused, what it cost, whether users benefited, whether answers remained reliable and how failures can be recovered.
An initial rollout should retain a way to disable the cache strategy. For provider mechanisms that are on by default and cannot be disabled, at least make prefix reordering, explicit cache settings and application-routing changes reversible. Observe the relationship between hits, write cost and tail latency after releases, with alerts tied to business budgets. If costs rise without latency gains, investigate repeated writes, reuse windows and workload distribution before extending retention.
The evidence for adopting prompt caching is lower total cost per effective business request under the same quality and permission constraints, while service experience meets requirements at the target load. Hit rate without request structure, cost decomposition and cold-state behaviour is insufficient for procurement or service commitments. This reflects the acceptance approach emphasized in IDENIFE's technical analysis: establish reproducible evidence before crediting optimization gains in a solution.
