Visual Q&A
First line is a public image URL; the rest is your question about it. The model
Showcase — Visual Q&A
First line is a public image URL; the rest is your question about it. The model gets both as one multimodal message and answers grounded in what the image shows — counting objects, reading a sign, describing the scene.
Run
bash bootstrap-secrets.sh # reads ../../../../.env, writes secrets/
docker compose up --build # default: PROVIDER=openai
Open http://localhost:3000. Gemini: PROVIDER=gemini docker compose up --build.
What's where
backend/ai_openai.py/backend/ai_gemini.py— parse the first line as the image URL and the rest as the question, then a multimodal call.frontend/app/page.tsx— URL-then-question box + answer.
Stop
docker compose down
Run locally
Download the project as a ZIP and run it with Docker. Brings up a FastAPI backend + Next.js frontend on localhost:3000.
unzip visual-qa.zip
cd visual-qa
bash bootstrap-secrets.sh # one-time: pulls API keys into ./secrets
docker compose up --build # default provider: openai
# or: PROVIDER=gemini docker compose up --build
Type some input, pick a provider, and run the same code shown in Source against the live API. Sign-in required.
The same modules the Run button hits. The whole project (frontend, Dockerfile, compose) is in the ZIP under README.
backend/ai_openai.py
"""Showcase 2 (OpenAI): visual question answering.
First line is a public image URL; the rest is your question about it. The model
gets both as one multimodal message and answers grounded in what the image
actually shows — counting objects, reading signs, judging the scene.
"""
from openai import OpenAI
_client = OpenAI()
_MODEL = "gpt-5.4-nano"
def run(text: str) -> str:
lines = text.strip().splitlines()
url = lines[0].strip() if lines else ""
question = "\n".join(lines[1:]).strip() or "What is happening in this image?"
if not url.startswith("http"):
return "First line must be a public image URL (http/https); put your question on the lines below it."
response = _client.responses.create(
model=_MODEL,
input=[{
"role": "user",
"content": [
{"type": "input_text", "text": question},
{"type": "input_image", "image_url": url},
],
}],
)
return response.output_text
backend/ai_gemini.py
"""Showcase 2 (Gemini): visual question answering.
First line is the image URL, the rest is the question. Gemini takes the image
as downloaded bytes.
"""
import os
import urllib.request
from google import genai
from google.genai import types
_client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
_MODEL = "gemini-3.1-flash-lite"
def _mime(url: str) -> str:
u = url.lower()
if u.endswith(".png"):
return "image/png"
if u.endswith(".webp"):
return "image/webp"
return "image/jpeg"
def run(text: str) -> str:
lines = text.strip().splitlines()
url = lines[0].strip() if lines else ""
question = "\n".join(lines[1:]).strip() or "What is happening in this image?"
if not url.startswith("http"):
return "First line must be a public image URL (http/https); put your question on the lines below it."
image_bytes = urllib.request.urlopen(urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0"})).read()
response = _client.models.generate_content(
model=_MODEL,
contents=[types.Part.from_bytes(data=image_bytes, mime_type=_mime(url)), question],
)
return response.text or ""
Project files
.gitignoreREADME.es.mdREADME.mdbackend/Dockerfilebackend/ai_gemini.pybackend/ai_openai.pybackend/main.pybackend/requirements.txtbootstrap-secrets.shdocker-compose.ymlfrontend/Dockerfilefrontend/app/layout.tsxfrontend/app/page.tsxfrontend/next.config.tsfrontend/package.jsonfrontend/tsconfig.json