1.9 KiB
1.9 KiB
Image Tools
Read and interpret image files using vision capabilities. Use call_tool.
Dependencies: Python 3, Pillow (for PNG resizing when over budget).
image_read
Read an image file and return it as a visual content block for interpretation. Automatically recompresses/resizes PNGs exceeding 128 KB.
Args: [path] or [path, --flags...]
Flags: --region=x,y,w,h (crop region), --scale=N (scale factor), --quality=N (compression quality)
call_tool: calls=[{tool: "image_read", args: ["/abs/path/to/image.png"]}]
call_tool: calls=[{tool: "image_read", args: ["/abs/path/to/screenshot.png", "--region=0,0,800,600"]}]
call_tool: calls=[{tool: "image_read", args: ["/abs/path/to/diagram.png", "--scale=0.5"]}]
Returns: Image content block (base64-encoded) that the model can see and interpret, plus a text summary with filename, size, and MIME type.
Constraints:
- Absolute paths only.
- File must exist.
- Budget: 128 KB max after recompression. Large PNGs are progressively scaled down (75% → 50% → 35% → 25%) until they fit.
- Supported formats: any format with a recognized MIME type (PNG, JPEG, GIF, WebP, SVG, BMP, etc.).
When to use
- Interpret visual content: diagrams, charts, UI mockups, error screenshots, handwritten notes.
- Verify rendered output: check that generated HTML/CSS/SVG looks correct.
- Read text from images: OCR-like extraction from screenshots or scanned documents.
- Debug UI issues: capture with
gui_screenshot, then read withimage_read.
# Capture then interpret workflow
call_tool: calls=[{tool: "gui_screenshot", args: ["active", "--output=/tmp/debug.png"]}]
call_tool: calls=[{tool: "image_read", args: ["/tmp/debug.png"]}]
Note: image_read makes the image visible to the model. Without it, image files are just opaque byte blobs. Use it any time you need to see what an image contains.