Course ES
← back to chapter

Visual Q&A

First line is a public image URL; the rest is your question about it. The model

Showcase — Visual Q&A

First line is a public image URL; the rest is your question about it. The model gets both as one multimodal message and answers grounded in what the image shows — counting objects, reading a sign, describing the scene.

Run

bash bootstrap-secrets.sh              # reads ../../../../.env, writes secrets/
docker compose up --build              # default: PROVIDER=openai

Open http://localhost:3000. Gemini: PROVIDER=gemini docker compose up --build.

What's where

  • backend/ai_openai.py / backend/ai_gemini.py — parse the first line as the image URL and the rest as the question, then a multimodal call.
  • frontend/app/page.tsx — URL-then-question box + answer.

Stop

docker compose down

Run locally

Download the project as a ZIP and run it with Docker. Brings up a FastAPI backend + Next.js frontend on localhost:3000.

Download visual-qa.zip

unzip visual-qa.zip
cd visual-qa
bash bootstrap-secrets.sh   # one-time: pulls API keys into ./secrets
docker compose up --build   # default provider: openai
# or:  PROVIDER=gemini docker compose up --build

Type some input, pick a provider, and run the same code shown in Source against the live API. Sign-in required.


  

The same modules the Run button hits. The whole project (frontend, Dockerfile, compose) is in the ZIP under README.

backend/ai_openai.py

"""Showcase 2 (OpenAI): visual question answering.

First line is a public image URL; the rest is your question about it. The model
gets both as one multimodal message and answers grounded in what the image
actually shows — counting objects, reading signs, judging the scene.
"""
from openai import OpenAI

_client = OpenAI()

_MODEL = "gpt-5.4-nano"


def run(text: str) -> str:
    lines = text.strip().splitlines()
    url = lines[0].strip() if lines else ""
    question = "\n".join(lines[1:]).strip() or "What is happening in this image?"
    if not url.startswith("http"):
        return "First line must be a public image URL (http/https); put your question on the lines below it."
    response = _client.responses.create(
        model=_MODEL,
        input=[{
            "role": "user",
            "content": [
                {"type": "input_text", "text": question},
                {"type": "input_image", "image_url": url},
            ],
        }],
    )
    return response.output_text

backend/ai_gemini.py

"""Showcase 2 (Gemini): visual question answering.

First line is the image URL, the rest is the question. Gemini takes the image
as downloaded bytes.
"""
import os
import urllib.request

from google import genai
from google.genai import types

_client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

_MODEL = "gemini-3.1-flash-lite"


def _mime(url: str) -> str:
    u = url.lower()
    if u.endswith(".png"):
        return "image/png"
    if u.endswith(".webp"):
        return "image/webp"
    return "image/jpeg"


def run(text: str) -> str:
    lines = text.strip().splitlines()
    url = lines[0].strip() if lines else ""
    question = "\n".join(lines[1:]).strip() or "What is happening in this image?"
    if not url.startswith("http"):
        return "First line must be a public image URL (http/https); put your question on the lines below it."
    image_bytes = urllib.request.urlopen(urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0"})).read()
    response = _client.models.generate_content(
        model=_MODEL,
        contents=[types.Part.from_bytes(data=image_bytes, mime_type=_mime(url)), question],
    )
    return response.text or ""

Project files

  • .gitignore
  • README.es.md
  • README.md
  • backend/Dockerfile
  • backend/ai_gemini.py
  • backend/ai_openai.py
  • backend/main.py
  • backend/requirements.txt
  • bootstrap-secrets.sh
  • docker-compose.yml
  • frontend/Dockerfile
  • frontend/app/layout.tsx
  • frontend/app/page.tsx
  • frontend/next.config.ts
  • frontend/package.json
  • frontend/tsconfig.json