All Guides
aiaiprompt-cachingperformancecostobservabilityopenai

Diagnosing Prompt Cache Misses in Production

Turn weak prompt-cache reuse into an evidence-backed investigation of prefixes, request settings, routing, cost, and latency.

Ryan VerWey
2026-10-01
8 min read

Prompt caching can lower input cost and time to first token when requests reuse a stable prefix. It can also fail quietly: the application still receives a valid response, but more input is processed from scratch than the team expected.

OpenAI made Prompt Cache Diagnostics generally available in the Responses API on September 8, 2026. For supported GPT-5.6-and-later models, an application can compare a request with a recent completed response and receive a classified explanation when expected prefix reuse did not occur. The feature complements the broader prompt-caching controls, including implicit and explicit breakpoints, cache lifetime, usage fields, and separate cache accounting.

The diagnostic result is evidence, not an automatic optimization. A cache miss can be intentional because the model, tools, service tier, output schema, reasoning settings, or conversation content changed. The goal is to distinguish useful changes from accidental churn, then confirm that a fix improves real cost or latency without weakening behavior.

Start With a Reuse Hypothesis

Do not begin by maximizing a cache-hit percentage. Define what should be reusable for one representative workload.

Record:

  • the model and endpoint
  • the stable developer instructions and reference material
  • tool names, descriptions, schemas, configuration, and ordering
  • output format, reasoning effort, verbosity, and service tier
  • the dynamic suffix, such as the current user request or tool result
  • expected request frequency and cache lifetime
  • the quality, latency, and cost measures that matter

Then state the hypothesis plainly: “These requests should share the same prefix through the end of the policy document, while the final user question changes.” That sentence gives the investigation a boundary. Without it, a lower cached-token count may be correct rather than defective.

Capture a Comparable Baseline

Choose a recent completed response from the same organization whose prefix the current request should reuse. Save its response ID with the request configuration and usage data. On the comparison request, set prompt_cache_options.comparison_response_id to that baseline ID.

from pathlib import Path
from openai import OpenAI

client = OpenAI()
stable_policy = Path("support-policy.txt").read_text()
stable_tools = [
    {
        "type": "function",
        "name": "lookup_account",
        "description": "Return the approved support summary for one account.",
        "parameters": {
            "type": "object",
            "properties": {"account_id": {"type": "string"}},
            "required": ["account_id"],
            "additionalProperties": False,
        },
    }
]

baseline = client.responses.create(
    model="gpt-6-astra",
    input=[
        {"role": "developer", "content": stable_policy},
        {"role": "user", "content": "Summarize account A."},
    ],
    tools=stable_tools,
)

comparison = client.responses.create(
    model="gpt-6-astra",
    input=[
        {"role": "developer", "content": stable_policy},
        {"role": "user", "content": "Summarize account B."},
    ],
    tools=stable_tools,
    prompt_cache_options={"comparison_response_id": baseline.id},
)

print(comparison.prompt_cache_diagnostics)
print(comparison.usage.input_tokens_details.cached_tokens)

The comparison ID asks for diagnostics only. It does not load the earlier conversation or force a cache lookup against only that response. Keep the actual conversation state explicit in the current request.

Compare like with like. A baseline from another model, organization, processing region, or materially different tool set is useful only when the difference itself is what you are testing.

Read Diagnostics and Usage Together

The diagnostic type describes the comparison:

  • cache_hit means no classified difference prevented reuse of the expected comparison prefix.
  • cache_miss includes a reason and an estimate of missed reusable tokens.
  • comparison_response_not_found means the short-lived diagnostic record is unavailable.
  • unavailable means the comparison was inconclusive or unsupported.

None of those values replaces usage accounting. Read usage.input_tokens_details.cached_tokens to measure reported cache reuse for the current response. A diagnostic hit can still include new uncached input after the reused prefix, and diagnostic token estimates can differ from billed usage fields.

Store enough metadata to reproduce the comparison without logging raw prompts unnecessarily:

EvidenceWhy it matters
Response IDs and timestampsProves which requests were compared and whether the record may have expired
Model and returned service tierDetects routing or configuration changes
Hash or version of stable instructionsFinds accidental prefix edits without copying sensitive content into logs
Tool-manifest versionDetects schema, description, configuration, or ordering drift
Cache mode and breakpointsShows which prefixes were eligible to be written or read
Cached and cache-write tokensMeasures reuse and write volume
Time to first token and total latencyConnects reuse to user-visible performance
Accepted-output resultPrevents a cheaper but worse response from looking successful

Trace the First Cache-Sensitive Difference

Prompt caching depends on exact prefix compatibility. The first meaningful difference can invalidate reuse from that point forward even when the visible user request looks similar.

Investigate in this order:

  1. Confirm the same model and compatible service tier handled both requests.
  2. Compare tool names, descriptions, schemas, configuration, and array order.
  3. Compare developer instructions and the ordering of input items.
  4. Check output schemas, reasoning effort, verbosity, parallel tool-call settings, and compaction behavior.
  5. Confirm any prompt cache key or separate-accounting boundary is stable for the intended group.
  6. Check whether the expected prefix met the model's minimum cacheable length and an eligible breakpoint.
  7. Check whether the entry could have expired or traffic could have crossed a processing boundary.

Fix one classified difference at a time, then repeat the comparison against the intended baseline. Diagnostics report the first classified reason, so another cause may appear only after the first is removed.

Stabilize Structure, Not Behavior

The safest optimization preserves request meaning while reducing accidental prefix churn.

Good candidates include:

  • keeping stable instructions and reference material before dynamic user content
  • versioning the tool manifest and preserving tool order
  • restricting tool use with supported selection controls instead of deleting and re-adding tool definitions
  • appending conversation turns rather than rewriting earlier items
  • placing an explicit breakpoint after reusable content on models that support it
  • keeping tenant or customer cache-accounting keys stable within the authorized boundary

Do not freeze a stale policy, suppress a required tool, keep the wrong model, or avoid necessary compaction merely to protect a cache hit. Behavior, safety, privacy, and correctness outrank reuse.

Prove the Economics

A cache write is not automatically a saving. Current OpenAI models can charge differently for uncached input, cache writes, and cache reads. The exact rates and supported controls can change, so use the current prompt-caching documentation rather than embedding a permanent price assumption in application logic.

Measure a representative sequence:

  1. One initial request that creates an eligible cached prefix.
  2. Several follow-up requests during the expected lifetime.
  3. The same workload with the proposed structural fix.
  4. A control workload that intentionally changes the prefix.

Compare cache-write tokens, cached tokens, uncached input tokens, total input cost, time to first token, output quality, and retry rate. Report cost per accepted output rather than cost per request. A change that improves reuse but increases failed or low-quality responses is not an optimization.

Respect Data and Tenant Boundaries

OpenAI documents that Prompt Cache Diagnostics can operate with Zero Data Retention because diagnostic records contain configuration metadata, token estimates, and hashes rather than raw prompts or outputs. Those records are organization-scoped and short-lived. Treat that as a provider-specific capability, not permission to log sensitive application content yourself.

For multi-tenant systems:

  • keep cache accounting separated where required by customer or privacy policy
  • do not use cache-hit observations to probe whether another tenant submitted matching content
  • avoid placing secrets or unnecessary personal data in shared stable prefixes
  • document the processing region and retention settings that govern the workload
  • restrict access to response IDs, usage records, and diagnostic logs

If the application cannot explain who may share a cache-accounting boundary, fix that ownership problem before tuning reuse.

Turn the Investigation Into a Regression Check

Once the team fixes an accidental miss, preserve the evidence as a small repeatable test:

  • use non-sensitive representative content above the minimum cacheable length
  • run the baseline and comparison with a pinned request manifest
  • assert that diagnostics are supported before interpreting the result
  • record cached and cache-write tokens without assuming an exact physical hit
  • fail only on a meaningful regression threshold, not one noisy request
  • pair the cache check with output and tool-behavior evaluations

Production dashboards should show trends by workload and request-manifest version. Alert on a sustained change in cached-token share, cost per accepted output, or time to first token. Use individual diagnostics to explain the trend, not as a high-cardinality log for every request.

Release Checklist

  • The reusable prefix and dynamic suffix are explicitly defined.
  • Baseline and comparison requests use the same intended organization, model, tier, tools, and settings.
  • Diagnostics and actual usage fields are both recorded.
  • The first cache-sensitive difference is identified with evidence.
  • The fix preserves behavior, safety, and data boundaries.
  • Cost and latency improve across several representative requests.
  • Output quality and tool behavior still pass their evaluations.
  • Monitoring can distinguish cache writes, reads, misses, and unsupported diagnostics.
  • Current provider documentation is linked for model support, retention, and pricing.

Prompt Cache Diagnostics turns “the cache seems worse” into a testable comparison. The durable practice is broader: version the request structure, measure reuse and outcomes together, and treat every optimization as a production change that needs evidence and a rollback path.

Ryan VerWey

Written by

Ryan VerWey

Ryan VerWey is a full-stack developer building tools and writing practical guides for working developers.