Course ES
← back to chapter

Image caption

Paste a public image URL and get an accessibility-style alt-text caption plus a

Showcase — Image caption

Paste a public image URL and get an accessibility-style alt-text caption plus a short description. It's the simplest multimodal call: one message whose content is an instruction text part and an image part, reasoned over together.

Run

bash bootstrap-secrets.sh              # reads ../../../../.env, writes secrets/
docker compose up --build              # default: PROVIDER=openai

Open http://localhost:3000. Gemini: PROVIDER=gemini docker compose up --build.

What's where

  • backend/ai_openai.py — passes the image URL directly to the model.
  • backend/ai_gemini.py — downloads the image and passes it as inline bytes.
  • frontend/app/page.tsx — URL box + caption.

Stop

docker compose down

Run locally

Download the project as a ZIP and run it with Docker. Brings up a FastAPI backend + Next.js frontend on localhost:3000.

Download image-caption.zip

unzip image-caption.zip
cd image-caption
bash bootstrap-secrets.sh   # one-time: pulls API keys into ./secrets
docker compose up --build   # default provider: openai
# or:  PROVIDER=gemini docker compose up --build

Type some input, pick a provider, and run the same code shown in Source against the live API. Sign-in required.


  

The same modules the Run button hits. The whole project (frontend, Dockerfile, compose) is in the ZIP under README.

backend/ai_openai.py

"""Showcase 1 (OpenAI): image captioning from a URL.

Paste a public image URL and get an accessibility-style caption plus a short
description. One multimodal message: an instruction text part and an image part.
"""
from openai import OpenAI

_client = OpenAI()

_MODEL = "gpt-5.4-nano"

_PROMPT = (
    "Write a concise one-sentence alt-text caption for this image, then a "
    "2-3 sentence description of what's in it. Label them 'Alt text:' and "
    "'Description:'."
)


def run(image_url: str) -> str:
    url = image_url.strip()
    if not url.startswith("http"):
        return "Paste a public image URL (starting with http:// or https://)."
    response = _client.responses.create(
        model=_MODEL,
        input=[{
            "role": "user",
            "content": [
                {"type": "input_text", "text": _PROMPT},
                {"type": "input_image", "image_url": url},
            ],
        }],
    )
    return response.output_text

backend/ai_gemini.py

"""Showcase 1 (Gemini): image captioning from a URL.

Same captioning prompt; Gemini takes the image as downloaded bytes.
"""
import os
import urllib.request

from google import genai
from google.genai import types

_client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

_MODEL = "gemini-3.1-flash-lite"

_PROMPT = (
    "Write a concise one-sentence alt-text caption for this image, then a "
    "2-3 sentence description of what's in it. Label them 'Alt text:' and "
    "'Description:'."
)


def _mime(url: str) -> str:
    u = url.lower()
    if u.endswith(".png"):
        return "image/png"
    if u.endswith(".webp"):
        return "image/webp"
    return "image/jpeg"


def run(image_url: str) -> str:
    url = image_url.strip()
    if not url.startswith("http"):
        return "Paste a public image URL (starting with http:// or https://)."
    image_bytes = urllib.request.urlopen(urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0"})).read()
    response = _client.models.generate_content(
        model=_MODEL,
        contents=[types.Part.from_bytes(data=image_bytes, mime_type=_mime(url)), _PROMPT],
    )
    return response.text or ""

Project files

  • .gitignore
  • README.es.md
  • README.md
  • backend/Dockerfile
  • backend/ai_gemini.py
  • backend/ai_openai.py
  • backend/main.py
  • backend/requirements.txt
  • bootstrap-secrets.sh
  • docker-compose.yml
  • frontend/Dockerfile
  • frontend/app/layout.tsx
  • frontend/app/page.tsx
  • frontend/next.config.ts
  • frontend/package.json
  • frontend/tsconfig.json