← All work
Open sourceAI Systems

ContextGate

An LLM router that answers semantically similar prompts from a pgvector cache instead of paying the provider for the same work twice.

My role
Designed and built the system. Sole author.
Stack
Python 3.12FastAPIPostgreSQLpgvectorRedis StreamsSQLAlchemyPydanticDocker ComposepytestGitHub Actions

The problem

LLM applications send the same request over and over: repeated system prompts, FAQ-shaped questions, retries, a handful of common user intents. Every one of them pays full provider latency and full token cost, even when a near-identical prompt was answered seconds earlier. Hosted providers keep their own KV caches, but those are internal and not addressable from outside, so the repeated work has to be eliminated before the call is made.

What I built

A FastAPI service that sits between the application and the model provider. It embeds each incoming prompt with a deterministic local embedder, searches a PostgreSQL + pgvector table for a semantically close prompt, and returns the stored response on a hit. A miss goes onto a Redis Stream. A worker picks it up, sends it through a swappable provider adapter, and writes the answer back into the cache. It ships as a four-service Docker Compose stack: pytest suite, a GitHub Actions workflow on every push and pull request to main, and a benchmark script that measures the cached path against the uncached one.

The hard part

A semantic cache that is too eager returns the wrong answer confidently. That failure is invisible, because the response still looks like a reasonable answer. So the similarity threshold is configuration (CACHE_SIMILARITY_THRESHOLD, default 0.88), not a constant buried in a service. The tests assert the property that matters rather than a fixed score: a near-paraphrase has to rank higher against a prompt than an unrelated question does. The embedder is the hashing trick. Tokenize, hash each token into one of 384 signed buckets, L2-normalize. That makes cache behaviour deterministic and reproducible, and it means the whole stack runs and tests without a paid embedding API or any provider key.

Architecture

How a request moves through it.

  1. Client applicationin

    POST /v1/chat with a prompt, model, user and session.

  2. FastAPI routersvc

    Validates the request against a Pydantic schema and assigns a request id.

  3. Embedding servicesvc

    Hashing-trick embedder: 384 signed buckets, L2-normalized, deterministic.

  4. Postgres + pgvectordb

    Cosine search over stored prompt vectors against the configured threshold.

    ↳ Cache hit

    Above the similarity threshold the stored response returns immediately. One vector lookup, no queue, no provider, no tokens.

  5. Redis Streamqueue

    Miss is enqueued on a consumer group so a restarted worker can reclaim it.

  6. LLM workerworker

    Consumes the stream and calls the selected provider adapter.

  7. Provider adapterext

    Mock by default; Gemini activates only when a key is configured.

  8. Cache write-backdb

    Response and its prompt vector are stored for the next similar request.

  9. Responseout

    Returns the answer plus cache_hit, similarity and the provider that served it.

One request through ContextGate. The cache hit is the path that matters. It never reaches the queue or the provider.

Technical decisions

What was chosen, and what it cost.

A deterministic local embedder instead of a hosted embedding API.

WhyThe cache's behaviour has to be reproducible in a test and on a laptop with no API key. Hashing the tokens into signed buckets gives the same vector for the same text, every time, for free.

CostIt captures lexical overlap, not deep semantics. A true paraphrase with different words scores lower than a real embedding model would give it. The provider interface is swappable for exactly this reason.

The similarity threshold is configuration, not a constant.

WhyThe correct threshold depends on the application. A support bot can be looser than a system answering questions about pricing. Burying it in code guarantees it gets tuned by editing a service.

CostIt is one more value an operator has to understand and set deliberately.

Redis Streams with a consumer group rather than a simple list queue.

WhyA worker that dies mid-request should not silently lose the request. A consumer group tracks pending entries so a restarted worker can reclaim them.

CostMore moving parts than a list pop, and the pending-entry recovery is deliberately basic in V1.

A mock provider that is the default, not an afterthought.

WhyAnyone can clone the repository and run the full request path (queue, worker, cache write-back) without a key or a bill. That is what makes the system inspectable rather than describable.

CostThe default experience does not produce real model output.

Constraints

  • No provider key may be required to run, test, or benchmark the system.
  • Hosted model KV caches are internal to the provider and cannot be controlled from outside, and the design does not pretend otherwise.
  • V1 deliberately ships without auth, billing, streaming, or a dashboard; those are named as out of scope rather than implied.

How it is verified

  • pytest runs in GitHub Actions on every push and pull request into main.
  • Similarity tests assert a relationship (a near-paraphrase scores above an unrelated prompt) instead of pinning a brittle magic number.
  • The provider adapter contract is tested against the mock implementation.
  • The benchmark script measures the same prompt cold and cached and prints the delta, so the claim is reproduced rather than asserted.

Outcome

  • The cached path is a single vector lookup in Postgres: it skips the queue, the worker and the provider entirely.
  • The whole stack comes up with one command and answers on :8000 with no external credentials.
  • Published under MIT with a documented roadmap separating what V1 does from what it explicitly does not.

What I would tell someone building this

  • The dangerous failure in a semantic cache is not a miss, it is a confident wrong hit. Testing the ordering property caught more than testing an absolute score would have.
  • Making the no-key path the default path is what turned the project from a description into something a stranger can run.
  • Writing down what V1 does not do removed more ambiguity than another feature would have added.

Questions about how this was built, or want something like it?