Can the model see images? How to confirm?
Distinguishing image attachments, previews, and the model's image processing capabilities.
First confirm the image source
Images can be provided as task materials or may be generated or read by the Agent from the environment. Being able to preview an image in the interface is different from the current model successfully completing image analysis; tools may report unsupported formats, sizes, or capabilities.
Do not assume the model has seen the details in the image just because it repeats the file name.
Make the question specific enough
Provide clear images and specify the area or phenomenon you want to check. For example, compare the button states on two interfaces or read a specific error segment in a screenshot.
When checking interface screenshots, you can first ask the model to describe each button's text and current state, then inquire about the reasons. For security alerts or important amounts, it is best to also provide the original text in a copyable format; image-based guesses and exact text should be clearly distinguished in the response.
- Confirm the attachment upload is complete or the image path truly exists.
- Request a description of key visible facts first, then an explanation.
- Supplement unclear text with the original text to avoid relying on guesswork for recognition.
How to proceed if the image cannot be read
If prompted that the current model does not support the corresponding image result, use available support options in the interface or provide textual materials to continue. Oversized images can be cropped to relevant areas; PDFs and SVGs should not be treated as ordinary bitmaps. Small text, occlusion, and layout in images may affect judgment; key numbers, paths, and error codes should be verified against the original image.