This repository has been archived on 2026-08-16. You can view files and clone it, but cannot push or open issues or pull requests.
ollie-9p/prompts/tools-image.md

1.9 KiB

Image Tools

Read and interpret image files using vision capabilities. Use call_tool.

Dependencies: Python 3, Pillow (for PNG resizing when over budget).


image_read

Read an image file and return it as a visual content block for interpretation. Automatically recompresses/resizes PNGs exceeding 128 KB.

Args: [path] or [path, --flags...]

Flags: --region=x,y,w,h (crop region), --scale=N (scale factor), --quality=N (compression quality)

call_tool: calls=[{tool: "image_read", args: ["/abs/path/to/image.png"]}]
call_tool: calls=[{tool: "image_read", args: ["/abs/path/to/screenshot.png", "--region=0,0,800,600"]}]
call_tool: calls=[{tool: "image_read", args: ["/abs/path/to/diagram.png", "--scale=0.5"]}]

Returns: Image content block (base64-encoded) that the model can see and interpret, plus a text summary with filename, size, and MIME type.

Constraints:

  • Absolute paths only.
  • File must exist.
  • Budget: 128 KB max after recompression. Large PNGs are progressively scaled down (75% → 50% → 35% → 25%) until they fit.
  • Supported formats: any format with a recognized MIME type (PNG, JPEG, GIF, WebP, SVG, BMP, etc.).

When to use

  • Interpret visual content: diagrams, charts, UI mockups, error screenshots, handwritten notes.
  • Verify rendered output: check that generated HTML/CSS/SVG looks correct.
  • Read text from images: OCR-like extraction from screenshots or scanned documents.
  • Debug UI issues: capture with gui_screenshot, then read with image_read.
# Capture then interpret workflow
call_tool: calls=[{tool: "gui_screenshot", args: ["active", "--output=/tmp/debug.png"]}]
call_tool: calls=[{tool: "image_read", args: ["/tmp/debug.png"]}]

Note: image_read makes the image visible to the model. Without it, image files are just opaque byte blobs. Use it any time you need to see what an image contains.