Anthropic Adds Computer Use to Claude 3.5 Sonnet: Driving Software Through the Screen

Anthropic introduced computer use in public beta for Claude 3.5 Sonnet, letting the model read screenshots and drive mouse and keyboard. SWE-bench Verified went from 33.4% to 49.0%.

Anthropic Adds Computer Use to Claude 3.5 Sonnet: Driving Software Through the Screen

In October 2024, Anthropic introduced “computer use” in public beta alongside an upgrade to Claude 3.5 Sonnet1. The feature lets the model perceive a computer screen the way a person does and operate software through mouse and keyboard input.

Computer use: looking at a screen and acting on it

What computer use amounts to is looking at a screen, moving a cursor, clicking buttons, and typing text1. The difference from earlier approaches is that software without a dedicated API can be driven through the same interface a human would use.

Anthropic was explicit at launch, however, that the feature was “still experimental—at times cumbersome and error-prone”1. It also noted that actions people perform effortlessly — scrolling, dragging, zooming — presented challenges for Claude at the time1. This was a release that signalled a direction rather than a finished, production-grade capability.

This shape of work — an AI operating its own environment to get a task done — went on to become the central theme of what is now called AI agents.

Coding performance

The Claude 3.5 Sonnet upgrade announced at the same time raised the score on the coding benchmark SWE-bench Verified from 33.4% to 49.0%1.

Anthropic also published comments from companies using it. GitLab reported stronger reasoning — up to 10% across use cases — with no added latency, and Cognition cited substantial improvements in coding, planning, and problem-solving1. Both comments assess the new Claude 3.5 Sonnet model in general; neither was offered as an account of deploying computer use.

How safety was handled

Anthropic stated plainly that computer use “may provide a new vector for more familiar threats such as spam, misinformation, or fraud”1. In response, it said it had developed new classifiers that can identify when computer use is being used and whether harm is occurring1.

Being able to operate a screen means the scope of granted permissions becomes the scope of potential damage. That concern later became an industry-wide design question about what authority to give agents and how to stop them.

The angle on legacy business systems

The ability to drive existing systems that expose no API, through the screen, can matter in environments where older software is still in place. Production-management systems and similar in-house tools without integration points could become candidates for automation without being rewritten.

That said, the maturity at launch was what Anthropic itself called experimental — not something to place under core operations as-is. Whether the approach would become usable in practice depended on the model and tooling improvements that followed.

Sources

  1. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku - Anthropic (October 22, 2024)

We publish the latest AI news every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →