prompts: add tools-image.md for image_read tool

This commit is contained in:
Levi Neely 2026-07-16 20:13:21 +02:00
parent 6c07c308e1
commit b82e4d857a
1 changed files with 46 additions and 0 deletions

46
prompts/tools-image.md Normal file
View File

@ -0,0 +1,46 @@
# Image Tools
Read and interpret image files using vision capabilities. Use `call_tool`.
**Dependencies**: Python 3, Pillow (for PNG resizing when over budget).
---
## image_read
Read an image file and return it as a visual content block for interpretation. Automatically recompresses/resizes PNGs exceeding 128 KB.
**Args**: `[path]` or `[path, --flags...]`
Flags: `--region=x,y,w,h` (crop region), `--scale=N` (scale factor), `--quality=N` (compression quality)
```
call_tool: calls=[{tool: "image_read", args: ["/abs/path/to/image.png"]}]
call_tool: calls=[{tool: "image_read", args: ["/abs/path/to/screenshot.png", "--region=0,0,800,600"]}]
call_tool: calls=[{tool: "image_read", args: ["/abs/path/to/diagram.png", "--scale=0.5"]}]
```
**Returns**: Image content block (base64-encoded) that the model can see and interpret, plus a text summary with filename, size, and MIME type.
**Constraints**:
- Absolute paths only.
- File must exist.
- Budget: 128 KB max after recompression. Large PNGs are progressively scaled down (75% → 50% → 35% → 25%) until they fit.
- Supported formats: any format with a recognized MIME type (PNG, JPEG, GIF, WebP, SVG, BMP, etc.).
---
## When to use
- **Interpret visual content**: diagrams, charts, UI mockups, error screenshots, handwritten notes.
- **Verify rendered output**: check that generated HTML/CSS/SVG looks correct.
- **Read text from images**: OCR-like extraction from screenshots or scanned documents.
- **Debug UI issues**: capture with `gui_screenshot`, then read with `image_read`.
```
# Capture then interpret workflow
call_tool: calls=[{tool: "gui_screenshot", args: ["active", "--output=/tmp/debug.png"]}]
call_tool: calls=[{tool: "image_read", args: ["/tmp/debug.png"]}]
```
**Note**: `image_read` makes the image visible to the model. Without it, image files are just opaque byte blobs. Use it any time you need to *see* what an image contains.