Preguntas y respuestas visuales
La primera línea es una image URL pública; el resto es tu pregunta sobre ella. El
Showcase — Preguntas y respuestas visuales
La primera línea es una image URL pública; el resto es tu pregunta sobre ella. El modelo recibe ambas como un solo mensaje multimodal y responde con grounding en lo que muestra la imagen — contar objetos, leer un letrero, describir la escena.
Córrelo
bash bootstrap-secrets.sh # reads ../../../../.env, writes secrets/
docker compose up --build # default: PROVIDER=openai
Abre http://localhost:3000. Gemini: PROVIDER=gemini docker compose up --build.
Qué hay aquí
backend/ai_openai.py/backend/ai_gemini.py— parsean la primera línea como la image URL y el resto como la pregunta, luego una llamada multimodal.frontend/app/page.tsx— caja de URL-luego-pregunta + respuesta.
Detenlo
docker compose down
Ejecútalo en tu máquina
Descarga el proyecto como ZIP y córrelo con Docker. Levanta un backend FastAPI y un frontend Next.js en localhost:3000.
unzip visual-qa.zip
cd visual-qa
bash bootstrap-secrets.sh # one-time: pulls API keys into ./secrets
docker compose up --build # default provider: openai
# or: PROVIDER=gemini docker compose up --build
Escribe algo, elige un proveedor y ejecuta el mismo código de Código contra la API real. Requiere iniciar sesión.
Los mismos módulos que ejecuta el botón Run. El proyecto completo (frontend, Dockerfile, compose) está en el ZIP, pestaña README.
backend/ai_openai.py
"""Showcase 2 (OpenAI): visual question answering.
First line is a public image URL; the rest is your question about it. The model
gets both as one multimodal message and answers grounded in what the image
actually shows — counting objects, reading signs, judging the scene.
"""
from openai import OpenAI
_client = OpenAI()
_MODEL = "gpt-5.4-nano"
def run(text: str) -> str:
lines = text.strip().splitlines()
url = lines[0].strip() if lines else ""
question = "\n".join(lines[1:]).strip() or "What is happening in this image?"
if not url.startswith("http"):
return "First line must be a public image URL (http/https); put your question on the lines below it."
response = _client.responses.create(
model=_MODEL,
input=[{
"role": "user",
"content": [
{"type": "input_text", "text": question},
{"type": "input_image", "image_url": url},
],
}],
)
return response.output_text
backend/ai_gemini.py
"""Showcase 2 (Gemini): visual question answering.
First line is the image URL, the rest is the question. Gemini takes the image
as downloaded bytes.
"""
import os
import urllib.request
from google import genai
from google.genai import types
_client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
_MODEL = "gemini-3.1-flash-lite"
def _mime(url: str) -> str:
u = url.lower()
if u.endswith(".png"):
return "image/png"
if u.endswith(".webp"):
return "image/webp"
return "image/jpeg"
def run(text: str) -> str:
lines = text.strip().splitlines()
url = lines[0].strip() if lines else ""
question = "\n".join(lines[1:]).strip() or "What is happening in this image?"
if not url.startswith("http"):
return "First line must be a public image URL (http/https); put your question on the lines below it."
image_bytes = urllib.request.urlopen(urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0"})).read()
response = _client.models.generate_content(
model=_MODEL,
contents=[types.Part.from_bytes(data=image_bytes, mime_type=_mime(url)), question],
)
return response.text or ""
Archivos del proyecto
.gitignoreREADME.es.mdREADME.mdbackend/Dockerfilebackend/ai_gemini.pybackend/ai_openai.pybackend/main.pybackend/requirements.txtbootstrap-secrets.shdocker-compose.ymlfrontend/Dockerfilefrontend/app/layout.tsxfrontend/app/page.tsxfrontend/next.config.tsfrontend/package.jsonfrontend/tsconfig.json