Course EN
← back to chapter

Preguntas y respuestas visuales

La primera línea es una image URL pública; el resto es tu pregunta sobre ella. El

Showcase — Preguntas y respuestas visuales

La primera línea es una image URL pública; el resto es tu pregunta sobre ella. El modelo recibe ambas como un solo mensaje multimodal y responde con grounding en lo que muestra la imagen — contar objetos, leer un letrero, describir la escena.

Córrelo

bash bootstrap-secrets.sh              # reads ../../../../.env, writes secrets/
docker compose up --build              # default: PROVIDER=openai

Abre http://localhost:3000. Gemini: PROVIDER=gemini docker compose up --build.

Qué hay aquí

  • backend/ai_openai.py / backend/ai_gemini.py — parsean la primera línea como la image URL y el resto como la pregunta, luego una llamada multimodal.
  • frontend/app/page.tsx — caja de URL-luego-pregunta + respuesta.

Detenlo

docker compose down

Ejecútalo en tu máquina

Descarga el proyecto como ZIP y córrelo con Docker. Levanta un backend FastAPI y un frontend Next.js en localhost:3000.

Descargar visual-qa.zip

unzip visual-qa.zip
cd visual-qa
bash bootstrap-secrets.sh   # one-time: pulls API keys into ./secrets
docker compose up --build   # default provider: openai
# or:  PROVIDER=gemini docker compose up --build

Escribe algo, elige un proveedor y ejecuta el mismo código de Código contra la API real. Requiere iniciar sesión.


  

Los mismos módulos que ejecuta el botón Run. El proyecto completo (frontend, Dockerfile, compose) está en el ZIP, pestaña README.

backend/ai_openai.py

"""Showcase 2 (OpenAI): visual question answering.

First line is a public image URL; the rest is your question about it. The model
gets both as one multimodal message and answers grounded in what the image
actually shows — counting objects, reading signs, judging the scene.
"""
from openai import OpenAI

_client = OpenAI()

_MODEL = "gpt-5.4-nano"


def run(text: str) -> str:
    lines = text.strip().splitlines()
    url = lines[0].strip() if lines else ""
    question = "\n".join(lines[1:]).strip() or "What is happening in this image?"
    if not url.startswith("http"):
        return "First line must be a public image URL (http/https); put your question on the lines below it."
    response = _client.responses.create(
        model=_MODEL,
        input=[{
            "role": "user",
            "content": [
                {"type": "input_text", "text": question},
                {"type": "input_image", "image_url": url},
            ],
        }],
    )
    return response.output_text

backend/ai_gemini.py

"""Showcase 2 (Gemini): visual question answering.

First line is the image URL, the rest is the question. Gemini takes the image
as downloaded bytes.
"""
import os
import urllib.request

from google import genai
from google.genai import types

_client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

_MODEL = "gemini-3.1-flash-lite"


def _mime(url: str) -> str:
    u = url.lower()
    if u.endswith(".png"):
        return "image/png"
    if u.endswith(".webp"):
        return "image/webp"
    return "image/jpeg"


def run(text: str) -> str:
    lines = text.strip().splitlines()
    url = lines[0].strip() if lines else ""
    question = "\n".join(lines[1:]).strip() or "What is happening in this image?"
    if not url.startswith("http"):
        return "First line must be a public image URL (http/https); put your question on the lines below it."
    image_bytes = urllib.request.urlopen(urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0"})).read()
    response = _client.models.generate_content(
        model=_MODEL,
        contents=[types.Part.from_bytes(data=image_bytes, mime_type=_mime(url)), question],
    )
    return response.text or ""

Archivos del proyecto

  • .gitignore
  • README.es.md
  • README.md
  • backend/Dockerfile
  • backend/ai_gemini.py
  • backend/ai_openai.py
  • backend/main.py
  • backend/requirements.txt
  • bootstrap-secrets.sh
  • docker-compose.yml
  • frontend/Dockerfile
  • frontend/app/layout.tsx
  • frontend/app/page.tsx
  • frontend/next.config.ts
  • frontend/package.json
  • frontend/tsconfig.json