Course ES
← back to chapter

Cache saver

A large static prefix sits in front of every question. Ask, then ask again — the second call reports most of its input as CACHED, billed cheaply. This is how a big system prompt or long context stays affordable.

Showcase — Cache saver

A large static prefix sits in front of every question. Ask, then ask again — the second call reports most of its input as CACHED, billed cheaply. This is how a big system prompt or long context stays affordable.

Run

bash bootstrap-secrets.sh              # reads ../../../../.env, writes secrets/
docker compose up --build              # default: PROVIDER=openai

Open http://localhost:3000. Gemini: PROVIDER=gemini docker compose up --build.

What's where

  • backend/ai_openai.py / backend/ai_gemini.py — big fixed prefix, reports cached token count.
  • frontend/app/page.tsx — ask twice and watch cached tokens jump.

Stop

docker compose down

Run locally

Download the project as a ZIP and run it with Docker. Brings up a FastAPI backend + Next.js frontend on localhost:3000.

Download cache-saver.zip

unzip cache-saver.zip
cd cache-saver
bash bootstrap-secrets.sh   # one-time: pulls API keys into ./secrets
docker compose up --build   # default provider: openai
# or:  PROVIDER=gemini docker compose up --build

Type some input, pick a provider, and run the same code shown in Source against the live API. Sign-in required.


  

The same modules the Run button hits. The whole project (frontend, Dockerfile, compose) is in the ZIP under README.

backend/ai_openai.py

"""Showcase 2 (OpenAI): watch prompt caching pay off.

A large static reference prefix sits in front of every question. Ask something,
then ask again: the second call reports most of its input tokens as CACHED, and
cached input is billed at a fraction of the price. This is how you make a
long-context or big-system-prompt app affordable — the fixed part is nearly free
after the first hit.
"""
from openai import OpenAI

_client = OpenAI()

_MODEL = "gpt-5.4-nano"

# The big fixed prefix — repeated to be large enough for caching to engage.
_CONTEXT = "Company handbook.\n" + (
    "Employees accrue 15 PTO days; sick leave is separate at 8 days. Remote work "
    "is allowed up to 3 days a week with approval. Expenses are reimbursed within "
    "30 days with a receipt. Parental leave is 12 weeks paid. " * 200
)


def run(question: str) -> str:
    q = question.strip() or "Summarize the handbook."
    response = _client.responses.create(
        model=_MODEL, instructions=_CONTEXT,
        input=[{"role": "user", "content": q}],
    )
    u = response.usage
    cached = getattr(getattr(u, "input_tokens_details", None), "cached_tokens", 0) or 0
    pct = (100 * cached / u.input_tokens) if u.input_tokens else 0
    return (f"{response.output_text}\n\n"
            f"--- metrics ---\n"
            f"input tokens:  {u.input_tokens}\n"
            f"cached tokens: {cached}  ({pct:.0f}% of input)\n"
            f"output tokens: {u.output_tokens}\n"
            f"Ask again — cached should jump and the effective cost drops.")

backend/ai_gemini.py

"""Showcase 2 (Gemini): watch prompt caching pay off.

Same large fixed prefix; Gemini reports cached_content_token_count on repeat
calls for its caching-capable models.
"""
import os

from google import genai
from google.genai import types

_client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

_MODEL = "gemini-3.1-flash-lite"

_CONTEXT = "Company handbook.\n" + (
    "Employees accrue 15 PTO days; sick leave is separate at 8 days. Remote work "
    "is allowed up to 3 days a week with approval. Expenses are reimbursed within "
    "30 days with a receipt. Parental leave is 12 weeks paid. " * 200
)


def run(question: str) -> str:
    q = question.strip() or "Summarize the handbook."
    response = _client.models.generate_content(
        model=_MODEL,
        contents=[types.Content(role="user", parts=[types.Part(text=q)])],
        config=types.GenerateContentConfig(system_instruction=_CONTEXT),
    )
    u = response.usage_metadata
    cached = getattr(u, "cached_content_token_count", 0) or 0
    pct = (100 * cached / u.prompt_token_count) if u.prompt_token_count else 0
    return (f"{response.text or ''}\n\n"
            f"--- metrics ---\n"
            f"input tokens:  {u.prompt_token_count}\n"
            f"cached tokens: {cached}  ({pct:.0f}% of input)\n"
            f"output tokens: {u.candidates_token_count or 0}\n"
            f"Ask again — cached should jump and the effective cost drops.")

Project files

  • .gitignore
  • README.es.md
  • README.md
  • backend/Dockerfile
  • backend/ai_gemini.py
  • backend/ai_openai.py
  • backend/main.py
  • backend/requirements.txt
  • bootstrap-secrets.sh
  • docker-compose.yml
  • frontend/Dockerfile
  • frontend/app/layout.tsx
  • frontend/app/page.tsx
  • frontend/next.config.ts
  • frontend/package.json
  • frontend/tsconfig.json