Same context structure, different wiring per model: the three cache protocols
The structure and discipline from the previous posts land differently on each model: Claude wants you to pin breakpoints, the OpenAI line wants you to stop touching bytes, Gemini has a separate API. We caught it live on DeepSeek — the first 35 messages byte-identical, broken at message 36, because of one extra Anthropic-style pin.
The previous two posts were about the context itself: the four segments and how to nail down system/tools, and how to dedupe tool results. Those rules are model-agnostic — get the ordering right and it holds everywhere.
This post is the last step of landing them: the same well-ordered context, sent to Claude, DeepSeek, or Gemini, requires completely different work from the client. A user switches models in a second; if the client still "injects cache_control for everyone," the hit rate tells you immediately that the wiring is wrong.
Three protocols, three answers to "who owns the hit"
| Type | Typical targets | Who owns the hit | What the client should do |
|---|---|---|---|
ephemeral | Anthropic Claude (incl. compat paths) | Your explicit breakpoints | Attach cache_control: { type: 'ephemeral' } at agreed spots; breakpoint count is capped |
openai-prompt-cache | OpenAI / DeepSeek / most GPT-compat | Server-side automatic prefix hash | Do not sprinkle cache_control; the useful knob is a stable prompt_cache_key (session affinity) |
gemini-context-cache | Gemini | Separate Context Cache API | Don’t run it through the Chat Completions breakpoint injector |
One line: Claude wants you to hammer nails; the OpenAI line wants you to stop touching bytes; Gemini is a different door.
ephemeral: pin well or pay twice
Anthropic’s path is explicit client pins. Spots we commonly use:
- last item in the tools array
- stable blocks near the front of system
- last user / tool message of the turn
Pins are capped (a handful, not infinite). Stable content under the pin; volatile content away from it — same physics as “static first,” two wordings.
Concrete failure: pinning per-turn reminders, todos, or MCP lists into the system head means you pay the write premium, miss next turn, and waste the write.
openai-prompt-cache: meddling breaks the prefix
This line’s server does longest-prefix match. Porting Anthropic cache_control looks “more professional” and often self-sabotages.
On DeepSeek we caught a live break: msg[0..35] identical, break at msg[36] — because a last_user_or_tool pin walks backward each turn. Last turn’s “last message” isn’t last this turn, so the same message had cache_control last turn and not this turn, and the prefix snaps at the shape change.
Fix isn’t romantic: if schema says type !== 'ephemeral', do not inject breakpoints. The knob you still own is a stable prompt_cache_key (we affinity on sessionId) so same-session traffic lands where prefixes can reuse — not a breakpoint; don’t mix them.
After the fix, hit rate moved from “obviously fake” back into our expected >90% steady band. The model didn’t get smarter; we stopped reshaping messages every turn.
gemini-context-cache: don’t force the same injector
Gemini’s context cache is a separate lifecycle, not “extra fields on messages.” On the Chat Completions adapter we leave breakpoints alone — wrong handling is worse than none.
Wire create / reference / invalidate through its Context Cache API. Hammering Claude nails into Gemini is the wrong wrench.
Three things it comes down to in code
Identify the type before doing anything: the model config declares whether it's ephemeral, openai-prompt-cache, or gemini-context-cache, and the runtime branches on it. We put the type in the schema precisely so nothing has to guess at runtime — guess wrong once and the bill and the feel both notice.
On the ephemeral line, pins should be few and stable, placed only on boundaries that genuinely survive across turns. Keep dynamic state away from them.
On the auto-prefix line, don't pin anything. Just keep the message bytes monotonic: fixed tool surface, append-only history, no decorative fields migrating each turn.
The endpoint can still veto it
Even with the protocol wired correctly, the link can waste it. Between official direct connections and aggregator proxies we've measured a hit-rate cliff of 96% vs 14–20% — no amount of client-side care fills a hole dug at the endpoint.
Align your measurement formula first, too: some usage fields compute cache_read / (cache_read + input), others don't. Before we aligned ours we printed a 187% hit rate. Confirm the formula before the number reaches a conclusion.
Three protocols aren't three "cache switches." They're three contracts about who owns the hit. Claude: you pin, it honors. The OpenAI line: it honors prefixes automatically, so don't deform them. Gemini: a separate door — don't impersonate it with someone else's syringe.


