Someone built a tool that lets AI agents draw annotations directly on your screen:…
By AI Update World · 2026-10-09

AI agents have become increasingly sophisticated at understanding and interacting with digital environments. But there's a persistent gap between what these systems perceive and what humans can verify they're actually seeing. Annotation tools that allow agents to mark up screens in real time represent a straightforward approach to that transparency problem, and they raise interesting questions about how we communicate with and monitor automated systems.
The challenge of agent verification
When an AI system operates autonomously on a computer, a human observer faces a genuine puzzle: what is the agent actually looking at? Traditional logs and transcripts can describe what happened, but they don't show which elements on screen the system attended to, which it ignored, or which it misunderstood. If an agent makes a mistake, you might see the error outcome but not the visual reasoning that led there. This creates what researchers sometimes call the "black box screen interaction" problem. Unlike a human colleague who can explain "I clicked that button because I thought it said submit," an autonomous agent leaves no inherent visual trail.
Visual communication and automation
Humans have always relied on visual markup to coordinate work. Managers draw red circles on printouts. Architects annotate blueprints. Teachers mark up student work with arrows and corrections. These marks serve a dual purpose: they're instructions to others, and they're evidence of attention and intent. When humans watch other humans work, they naturally observe what they're pointing at and focusing on. Extending that same visual language to AI agents creates a bridge between human and machine workflows. Instead of asking "did the agent see X," you can literally see where the agent drew an arrow pointing at X.
How annotation layers work technically
Screen annotation typically operates as an overlay, a transparent layer drawn on top of the existing display. Rather than modifying the original screen content, the system renders shapes, text, and indicators on top of it. From a technical standpoint, this can happen at different levels: some approaches hook into the graphics pipeline and inject drawing commands, while others capture the screen, add annotations programmatically, and display the result. For AI agents, the annotation capability would need to be integrated into whatever computer vision or screen-reading system the agent uses. When the agent identifies a clickable element, a form field, or a region of interest, it can simultaneously render a visual marker at that location.
Agency and interpretability in automation
There's a subtle but important distinction between observation and agency. When a human reads a logfile saying "clicked button at coordinates 400, 300," they still don't know why that location mattered. But when an agent draws a box around an element and labels it "Submit Button," it's making an interpretive