Stop paying twice.
When the same request has been made before, Cache returns the saved answer instantly - without generating or charging for another one.
Hello, my name is
The answer was already generated.
Why pay again?
When the same AI request comes in again, Cache returns the existing answer instead of generating it again. Only another matching request can retrieve it.
For developers
Keep the API format you already use. Replace the upstream host with Cache and a matching cached answer is returned free, without requiring an account.
Anonymous mode can only return an existing cache hit. If your exact request is new, the gateway will ask for authorization instead of running it.
A better default
When the same request has been made before, Cache returns the saved answer instantly - without generating or charging for another one.
Every successful miss can save another answer for the next identical request. The gateway becomes more useful with every match.
A matching cached answer can be returned without an account. Answers are retrieved by sending the matching request, not by browsing a catalog.
How it works
Cache fingerprints the request, checks for the same request, and only calls the model when it needs something new.
Point an OpenAI-compatible completions request at the Cache gateway.
The model, prompt, and relevant settings must match to return the same cached answer.
A hit comes back free. A miss runs normally and can become the next person’s hit.
Designed for exact matches
Cache doesn't publish or index requests. It returns a saved answer only when another request matches the model, prompt, and relevant settings. Apply the same data-handling judgment you already use with your AI provider.