← All writing

Engineering7 min read

What Makes a Semantic Cache Safe Enough to Ship

A semantic cache that is too eager does not return an error. It returns a confident, well-formed, wrong answer. Here is how I tested for that while building ContextGate.

LLM applications repeat themselves constantly. The same system prompt, the same handful of user intents, the same question phrased six slightly different ways. Every repetition pays full latency and full token cost for work that was already done.

The obvious fix is a cache. The non-obvious part is that a cache for prompts is a fundamentally more dangerous object than a cache for URLs.

The failure mode nobody sees

A normal cache miss is visible. You ask for a key, it is not there, you do the work.

A semantic cache does not work in exact keys. It works in "close enough." You embed the incoming prompt, you search for a stored prompt whose vector is near it, and if the similarity clears a threshold you return the stored response.

So the failure mode is not a miss. It is a confident wrong hit.

Ask "what is our refund policy for annual plans" and get back the cached answer to "what is our refund policy for monthly plans." Nothing errors. Nothing logs. The response is well-formed, fluent, plausible, and wrong — and it looks exactly like a correct one.

That is the thing the system has to be engineered against.

Where the threshold lives

In ContextGate the similarity threshold is CACHE_SIMILARITY_THRESHOLD, defaulting to 0.88, and it is configuration rather than a constant inside the cache service.

That sounds like a small distinction. It is not.

The correct threshold is a property of the application, not of the cache. A support bot answering general questions can afford to be loose — a near-match is a good answer. A system answering questions about pricing, eligibility or policy cannot, because the near-matches are precisely the cases where a small difference in the question means a completely different correct answer.

If that number lives inside a service, tuning it means editing code and redeploying, which means it does not get tuned. If it lives in configuration, an operator can set it deliberately for their risk tolerance. The system stops making that decision on the application's behalf.

Test the property, not the number

The instinct is to write this test:

assert cosine_similarity(embed("..."), embed("...")) > 0.91

Do not. That test encodes a magic number produced by whatever embedder you happened to be using the day you wrote it. Change the embedding dimensions, change the tokenizer, swap the model, and the test fails without anything actually being wrong. So someone updates the number until it passes, and now the test asserts nothing.

The property that actually matters is an ordering:

def test_different_prompts_have_lower_similarity():
    base = embed_text("Explain what semantic caching is")
    similar = embed_text("Explain what semantic caching means")
    different = embed_text("What is the tallest mountain on Earth")

    assert cosine_similarity(base, different) < cosine_similarity(base, similar)

A near-paraphrase must rank above an unrelated question. That has to hold for any embedder worth using. It survives swapping the implementation, and when it fails, it has failed for a real reason: your embedding function has stopped distinguishing meaning, which is the one thing it exists to do.

The test is about the invariant. The invariant is what you are actually shipping.

Determinism is a testing decision

ContextGate V1 embeds with the hashing trick rather than a hosted embedding API: tokenize the text, hash each token into one of 384 buckets with a sign, sum, L2-normalize.

It is not as semantically sharp as a real embedding model. A true paraphrase using entirely different vocabulary scores lower than it should. That is a genuine limitation and the provider interface exists so it can be swapped.

But it buys three things that mattered more for V1:

  1. The same text always produces the same vector. Cache behaviour is reproducible, so a failing test means something.
  2. No key is required to run it. The whole stack — API, worker, queue, cache — comes up with docker compose up and answers on port 8000 with no account and no bill.
  3. No network call in the hot path. Embedding is arithmetic, so the cache lookup cannot be slower than the model call it replaces.

That third point is easy to miss. A cache that calls a remote embedding API before it can check itself has added a network round trip to every request just to sometimes save one.

The part that is actually fast

On a hit, the request never touches the queue, the worker or the provider. It is one vector search in Postgres and a return.

On a miss, it goes onto a Redis Stream, a worker consumes it, calls the provider, and writes the response back with its vector. Streams with a consumer group rather than a plain list, so a worker that dies mid-request leaves a pending entry that a restarted worker can reclaim instead of silently dropping the work.

That is the whole design. The interesting engineering is not the pipeline — it is the branch at the top, and whether you can trust it.

What I would tell someone building one

  • The dangerous case is the confident wrong hit, and it is invisible in logs. Design for it first.
  • Test relationships, not magic numbers. A test that gets "updated until it passes" is not a test.
  • Put the risk threshold where the person carrying the risk can set it.
  • Make the no-credentials path the default path. It is the difference between a system someone can evaluate and one they have to take your word for.

The code is public: github.com/Ejsav/V1.