Docs you can actually read.

Cache developer docs · v1

Generate once.
Reuse on a match.

Send a familiar completion request. If the same request already exists, reuse it without an account for free. If it does not, bring a key, generate it once, and save the result so the same request doesn't need to be generated again.

Base URL

https://api.meetcache.ai
Cache API v1
Cache the otter wearing a dark hoodie and typing on a laptop

01 / Quickstart

Try the cache.

Start without an account. This request costs nothing when Cache already has an answer for the same request.

Anonymous · cURL
curl https://api.meetcache.ai/v1/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-sol",
    "prompt": "Why do otters hold hands?",
    "max_tokens": 120
  }'

Cache miss? Add your upstream API key. Cache detects the provider from the key, forwards the request, then stores the successful result for future reuse when another request matches.

BYOK · cURL
curl https://api.meetcache.ai/v1/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -d '{"model":"gpt-5.6-sol","prompt":"Why do otters hold hands?","max_tokens":120}'

02 / Core concept

One request. Three stops.

Cache looks nearby first, then durable answer storage, then the model. A result moves back through the same path.

  1. 01 · FAST

    Edge cache

    The closest reusable copy. Edge entries are kept for 24 hours.

  2. 02 · SAVED

    Reusable answer store

    The durable saved generation. A hit here also warms the edge.

  3. 03 · ON A MISS

    Upstream API

    Your key selects the provider. Successful results are saved for future matching requests.

Upstream errors and successful non-JSON responses pass through, but are not cached. Background storage or analytics failures never replace a successful model response.

03 / Cache identity

What counts as “the same”?

“Same request” means the method, URL, request body, model, prompt, and meaningful settings produce the same SHA-256 fingerprint. JSON object key order does not matter, but meaningful request differences do. Cache does not offer an API for listing or searching fingerprints.

InputIn the key?
Method + full URLYes
Request bodyYes
Provider metadataYes
stream / stream_optionsNo
AuthorizationNo

If the body is not valid JSON, Cache still forwards it and fingerprints the exact raw bytes.

04 / Authentication

No key on a hit.
A key on a miss.

NO ACCOUNT

Reuse what exists

Send no authorization header. Matching cached answers return free; a miss reaches upstream without credentials and will normally be rejected.

BYOK

Create what is missing

Send Authorization: Bearer …. You pay the upstream provider only when a new generation is needed.

Need identity-based limits or loaded credits? Compare access tiers →

05 / Encrypted cache

PRO

Cache ciphertext.
Keep the private key.

Pro customers can ask Cache to encrypt a successful generation before it is stored. Send an X25519 public key; Cache returns and stores only a JWE that the matching private key can decrypt.

Encrypted cache · cURL
curl https://api.meetcache.ai/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $CACHE_API_KEY" \
  -H "Cache-Encryption-Key: $PUBLIC_JWK_BASE64URL" \
  -d '{"model":"gpt-5.6-sol","input":"Summarize the attached report","stream":false}'

KEY FORMAT

Public X25519 JWK

Cache-Encryption-Key is the unpadded base64url encoding of a UTF-8 JWK: {"kty":"OKP","crv":"X25519","x":"…","kid":"…"}. The optional kid labels the key. A private d value is rejected.

ENVELOPE

Compact JWE

The response is application/jose using alg: ECDH-ES with X25519 and enc: A256GCM. When a cache entry is created, Cache generates one ephemeral key and IV, stores the resulting JWE, and discards the ephemeral private key. Cache hits replay that stored JWE unchanged. Decrypt it with the matching private JWK and a JOSE library.

Cache isolation
Encrypted requests never read from or write to the standard reusable cache. The recipient key's RFC 7638 thumbprint is part of the encrypted cache identity, so rotating keys creates a separate entry.
What is protected
Successful generated response bodies are encrypted before persistent or edge storage. Cache does not persist a plaintext copy. Uncached errors may still be returned as ordinary JSON.
Streaming
Encrypted cache requests require "stream": false because the authenticated JWE is returned as one complete value.

Upstream access. Your model provider still receives and processes the request and generated response unencrypted. This feature protects the response stored in Cache; it does not change how the upstream provider handles it.

Standards: JWE (RFC 7516), X25519 for JOSE (RFC 8037), and JWK thumbprints (RFC 7638).

06 / Streaming

Stream without splitting the cache.

Set "stream": true as usual. Stream settings are excluded from the fingerprint, so streaming and non-streaming callers share the same result.

ON A MISS

Live SSE

Chunks pass through immediately while Cache assembles a canonical JSON completion in the background.

ON A HIT

Cached SSE envelope

The canonical result is returned as an SSE data event followed by [DONE].

The answer is equivalent, but cached replay does not reproduce the original token-by-token chunk boundaries.

07 / API reference

The small API.

POST/v1/completions

Accepts the native prompt-style OpenAI completions body. Replace https://api.openai.com with https://api.meetcache.ai; the rest of the request stays the same.

Also supported: /v1/responses · /v1/messages

Content-Type
application/json
Response
JSON or SSE
GET/health

A lightweight liveness check. Returns {"status":"ok"}.

08 / Batch API

BETA

Reuse first. Generate the rest.

Use the standard OpenAI Files and Batches APIs through Cache. Every JSONL request is checked for a matching cached request first, and only misses are submitted to OpenAI.

OpenAI JSONL input
{"custom_id":"answer-001","method":"POST","url":"/v1/responses","body":{"model":"gpt-5.6-sol","input":"Why do otters hold hands?"}}

POST/v1/files → /v1/batches

Upload with purpose batch, then create the batch using the returned file ID.

GET/v1/batches/:id

Authenticated polls advance submission and finalization. Keep polling until the batch reaches a terminal status.

Supported batch endpoints: /v1/responses, /v1/chat/completions, and /v1/completions. A cache failure stops processing rather than silently generating every request.

09 / Errors

Failures stay recognizable.

Upstream status codes and response bodies pass through unchanged. Gateway routing errors use a compact JSON response.

400
The encryption key is malformed, contains private key material, or is combined with streaming.
401/403
The cache missed and the upstream key is missing, invalid, or not permitted.
404
The gateway route does not exist.
405
The route exists, but not for that HTTP method.
429
An upstream or account limit was reached.
500
The gateway could not complete a cache read or request operation.

10 / The important bit

Matching is not publishing.

Cache doesn't list, index, or provide search over standard cached requests. A saved answer can be retrieved without an account only by reproducing the same canonical request. Use the same data-handling practices you already apply when sending requests to the upstream AI provider.

Your authorization header is used for the upstream miss and is not part of the standard cache key.

That’s the whole idea.

Ask once. Answer many.

Make your first request ↑