Gemini API adds computer-use tier for desktop automation
Google exposes screen-level control through the Gemini API with per-application permission scoping, targeting RPA vendors and internal tooling teams. Demo content.
Google has added a computer-use tier to the Gemini API, exposing screen-level desktop control with per-application permission scoping.
The capability
- ▸Screenshot-grounded clicking, typing and scrolling
- ▸Permission manifests that restrict agents to named applications
- ▸Audit logs of every action for compliance teams
Target market
The tier aims squarely at RPA (robotic process automation) vendors and enterprise internal-tools teams still maintaining brittle screen-scraping bots built a decade ago.
"Legacy RPA breaks when a button moves 10 pixels. Vision-grounded agents just find the button." — automation consultant
Pricing is usage-based with a substantial enterprise floor, signaling Google wants committed deployments rather than hobbyist experiments.
Source: Google DeepMind


