One of the things that most people don’t know is that artificial intelligence is not able to seenor does he know what is in an image. When you describe a photograph in detail, it is the result of interpreting a set of pixels, based on their training. He interprets, he does not see. That is why Google has introduced the Agentic Vision or Agentic Vision, to Gemini 3 Flash.
AI models like Gemini process the world they “see” with a single static glance. If they miss a minute detail, like the serial number on a microchip or a distant road sign, they are forced to guess. That’s why they make mistakes.
Agentic Vision in Gemini 3 Flash radically changes the way you analyze an image. Instead of just taking a quick look and trying to guess, Google’s AI now investigate the image with a reasoning modeland is even able to write code in real time to zoom in and refine what you see.
This is how Google’s Agentic Vision works
As Google explains on its blog, Agentic Vision carries out three steps in image recognition: thought, action and observation.
During the cycle of Thoughtthe AI analyzes the user’s query and the initial image, and formulates a multi-step plan.
In the cycle of Actiongenerates and runs Python code to manipulate the images (for example, crop, rotate, annotate) or analyze them (perform calculations, count bounding boxes, etc.), to better understand it.
In the cycle of Observationthe transformed image is added to the context window. This allows the AI to inspect the new data with better context before generating a final response.
As seen in the previous graph, with this new technique Google ensures that improves image recognition by 5 to 10%in different benchmarks. It may not seem like much, but when we talk about reducing errors in a vital task such as recognizing images, which can be part of a police report or professional work, it is a significant improvement.
Agentic Vision is now available via the Gemini API in Google AI Studio and Vertex AI. It’s also in the Gemini app within the Reasoning drop-down menu. Developers can try the demo in Google AI Studio or experiment with the feature in AI Studio Playground by enabling Code execution en Tools.