Sharing is caring!

In 2024, the AI landscape witnessed a pivotal moment with the introduction of Rabbit R1. Although it faced significant challenges and was deemed a failure, Rabbit R1 introduced the concept of Large Action Models (LAMs), paving the way for the development of autonomous AI agents capable of interacting with digital environments.

As of early 2025, three major competitors have emerged in the realm of computer-using agents:

Claude’s Compute

Developed by Anthropic, Claude’s Compute-first appeared in late 2024-enables AI to interact with desktop environments by simulating human actions such as clicking, typing, and navigating. Currently, this feature requires some technical knowledge to set up and use effectively via the API, limiting its accessibility for non-technical users.

Google Mariner

An experimental agent from Google’s DeepMind, Project Mariner is designed to navigate and interact with web pages autonomously. While still in its research phase, Mariner is being tested with a small group of users, and its integration within Google’s ecosystem suggests it could excel in workflows involving Gmail, Google Docs, and other Google services.

OpenAI Operator

OpenAI’s entry into the field, Operator, is powered by the Computer-Using Agent (CUA) model. Operator interacts with web pages by viewing screenshots and performing mouse and keyboard actions, allowing it to autonomously perform tasks such as filling out forms, ordering groceries, and even creating memes. It has set new performance benchmarks, achieving a 38.1% success rate on the OSWorld benchmark for full operating system tasks, surpassing previous models.

How The Battle Will Unfold?

OpenAI has publicly declared that Operator outperforms Anthropic’s “Computer Use” (a simplified version of Claude 3.5 Sonnet that can carry out basic computer tasks) and Google’s Mariner (built on Gemini 2.0 for web browsing).

OpenAI’s confidence is buoyed by benchmark results:

  • On OSWorld, which tests tasks like merging PDF files or editing images, Operator (CUA) scores 38.1%, trouncing Claude’s Compute at 22.0% (Mariner doesn’t participate here because it is web-only).
  • On WebVoyager, which tests browser-based tasks, Operator scores 87%, narrowly edging out Mariner at 83.5%, and leaving Anthropic’s offering behind at 56%.

However, benchmarks only paint part of the picture. Real-world computer usage includes CAPTCHA challenges, unexpected website layouts, sudden security prompts, and other chaotic elements that can trip up even the most advanced AI. While CUA’s “ask-before-acting” safety mechanism is designed to address ethical concerns and avoid misuse (e.g., building a bioweapon or executing malicious hacks), hidden instructions or prompt injections on a webpage could still derail an agent’s actions. OpenAI acknowledges these challenges and has employed “red teams” to probe possible vulnerabilities.

This sentiment underscores the reality: all three companies — Anthropic, Google, and OpenAI — have a shared end goal of transforming how humans use computers, yet their paths differ in focus, ethics, and ecosystem integration.

  • Anthropic’s emphasis on AI alignment and safety may resonate with enterprise customers worried about security.
  • Google’s massive, browser-based ecosystem could give Mariner a unique edge, especially in seamlessly integrating across Gmail, Docs, and future services.
  • OpenAI’s Operator aims to be the broadest solution, with near-term browser dominance and long-term potential to control entire operating systems via planned APIs.

Conclusion

From Rabbit R1’s short-lived run and introduction of Large Action Models to today’s showdown among Claude’s Compute, Google Mariner, and OpenAI Operator, the AI world has shifted from simple chatbots to autonomous agents capable of extensive digital actions. Operator’s early lead in benchmarks illustrates how rapidly AI capabilities are evolving, but real-world challenges and ethical considerations will determine which agent truly shapes our future interaction with computers.

For now, the AI community watch keenly as OpenAI, Anthropic, and Google gear up for what promises to be a prolonged — and transformative — competition. The next few years will reveal whether these LAM-powered autonomous agents can reliably conquer the messy, unpredictable realm of real-world computer tasks while maintaining user trust and safety. Regardless of who comes out on top, the race itself is already reshaping the way we think about and use technology every day.

Sharing is caring!