---
title: "Same context structure, different wiring per model: the three cache protocols"
date: "2026-07-30"
tag: "Cache"
lang: "en"
reading_minutes: 9
source: "https://neox-dev.com/blog/prompt-cache-three-protocols"
alternate: "https://neox-dev.com/md/blog/prompt-cache-three-protocols.md?lang=zh"
---

# Same context structure, different wiring per model: the three cache protocols

> The structure and discipline from the previous posts land differently on each model: Claude wants you to pin breakpoints, the OpenAI line wants you to stop touching bytes, Gemini has a separate API. We caught it live on DeepSeek — the first 35 messages byte-identical, broken at message 36, because of one extra Anthropic-style pin.

The previous two posts were about the context itself: [the four segments and how to nail down system/tools](/blog/prompt-cache-what-it-actually-saves), and [how to dedupe tool results](/blog/readfile-ledger-vs-claude-code). Those rules are model-agnostic — get the ordering right and it holds everywhere.

This post is the last step of landing them: the same well-ordered context, sent to Claude, DeepSeek, or Gemini, requires completely different work from the client. A user switches models in a second; if the client still "injects cache_control for everyone," the hit rate tells you immediately that the wiring is wrong.

## Three protocols, three answers to "who owns the hit"

![Three protocols: who pins breakpoints, who auto-hashes prefixes, who uses a side API](/site/blog/three-protocols-en.svg)

| Type | Typical targets | Who owns the hit | What the client should do |
|------|-----------------|------------------|---------------------------|
| `ephemeral` | Anthropic Claude (incl. compat paths) | Your explicit breakpoints | Attach `cache_control: { type: 'ephemeral' }` at agreed spots; breakpoint count is capped |
| `openai-prompt-cache` | OpenAI / DeepSeek / most GPT-compat | Server-side automatic prefix hash | **Do not** sprinkle `cache_control`; the useful knob is a stable `prompt_cache_key` (session affinity) |
| `gemini-context-cache` | Gemini | Separate Context Cache API | Don’t run it through the Chat Completions breakpoint injector |

One line: **Claude wants you to hammer nails; the OpenAI line wants you to stop touching bytes; Gemini is a different door.**

## ephemeral: pin well or pay twice

Anthropic’s path is explicit client pins. Spots we commonly use:

- last item in the tools array
- stable blocks near the front of system
- last user / tool message of the turn

Pins are capped (a handful, not infinite). Stable content under the pin; volatile content away from it — same physics as “static first,” two wordings.

Concrete failure: pinning per-turn reminders, todos, or MCP lists into the system head means you pay the write premium, miss next turn, and waste the write.

## openai-prompt-cache: meddling breaks the prefix

This line’s server does longest-prefix match. Porting Anthropic `cache_control` looks “more professional” and often **self-sabotages**.

On DeepSeek we caught a live break: msg[0..35] identical, break at msg[36] — because a `last_user_or_tool` pin walks backward each turn. Last turn’s “last message” isn’t last this turn, so the same message **had cache_control last turn and not this turn**, and the prefix snaps at the shape change.

Fix isn’t romantic: if schema says `type !== 'ephemeral'`, **do not inject** breakpoints. The knob you still own is a stable `prompt_cache_key` (we affinity on sessionId) so same-session traffic lands where prefixes can reuse — not a breakpoint; don’t mix them.

After the fix, hit rate moved from “obviously fake” back into our expected **>90%** steady band. The model didn’t get smarter; we stopped reshaping messages every turn.

## gemini-context-cache: don’t force the same injector

Gemini’s context cache is a separate lifecycle, not “extra fields on messages.” On the Chat Completions adapter we **leave breakpoints alone** — wrong handling is worse than none.

Wire create / reference / invalidate through its Context Cache API. Hammering Claude nails into Gemini is the wrong wrench.

## Three things it comes down to in code

Identify the type before doing anything: the model config declares whether it's `ephemeral`, `openai-prompt-cache`, or `gemini-context-cache`, and the runtime branches on it. We put the type in the schema precisely so nothing has to guess at runtime — guess wrong once and the bill and the feel both notice.

On the ephemeral line, pins should be few and stable, placed only on boundaries that genuinely survive across turns. Keep dynamic state away from them.

On the auto-prefix line, don't pin anything. Just keep the message bytes monotonic: fixed tool surface, append-only history, no decorative fields migrating each turn.

## The endpoint can still veto it

Even with the protocol wired correctly, the link can waste it. Between official direct connections and aggregator proxies we've measured a hit-rate cliff of 96% vs 14–20% — no amount of client-side care fills a hole dug at the endpoint.

Align your measurement formula first, too: some usage fields compute `cache_read / (cache_read + input)`, others don't. Before we aligned ours we printed a 187% hit rate. Confirm the formula before the number reaches a conclusion.

Three protocols aren't three "cache switches." They're three contracts about **who owns the hit**. Claude: you pin, it honors. The OpenAI line: it honors prefixes automatically, so don't deform them. Gemini: a separate door — don't impersonate it with someone else's syringe.
