Live

Gemini API adds computer-use tier for desktop automation

Google exposes screen-level control through the Gemini API with per-application permission scoping, targeting RPA vendors and internal tooling teams. Demo content.

MVMara Vidović
Published 28 Jun 2026, 12:20Updated 25 Jul 2026, 17:272 min read
Gemini API adds computer-use tier for desktop automation

Google has added a computer-use tier to the Gemini API, exposing screen-level desktop control with per-application permission scoping.

The capability

  • Screenshot-grounded clicking, typing and scrolling
  • Permission manifests that restrict agents to named applications
  • Audit logs of every action for compliance teams

Target market

The tier aims squarely at RPA (robotic process automation) vendors and enterprise internal-tools teams still maintaining brittle screen-scraping bots built a decade ago.

"Legacy RPA breaks when a button moves 10 pixels. Vision-grounded agents just find the button." — automation consultant

Pricing is usage-based with a substantial enterprise floor, signaling Google wants committed deployments rather than hobbyist experiments.

Source: Google DeepMind

Related stories

All →
Gemini API adds computer-use tier for desktop automation — DailyPrompt