Wain AI/Tech Blog

AI news and trends worldwide, updated nearly every day

How Claude Code's macOS Sandbox Could Be Escaped Is Now Public - Researchers Say It Was Fixed in 2.1.247

How Claude Code's macOS Sandbox Could Be Escaped Is Now Public - Researchers Say It Was Fixed in 2.1.247

On September 11, 2026, the startup Accomplish published how opening an untrusted repository in Claude Code could let commands run outside the macOS sandbox without a permission prompt. It says it reported the issue on July 13 and that it was fixed in 2.1.247.

Accomplish, a startup with an office in Tel Aviv, published a technical write-up on September 11, 2026 detailing “Beltdown,” a technique for getting out of Claude Code’s sandbox on macOS1. By the company’s account, if you open an untrusted repository in Claude Code, a command planted in that repository can end up running outside the sandbox with your own user privileges. No permission prompt appears.

According to Accomplish, Anthropic fixed the problem in Claude Code 2.1.247 after receiving the report1. If you use Claude Code on macOS, the first thing to check is whether your installed version is 2.1.247 or later.

No Prompt Appeared Because the Sandbox Was On

Accomplish set up its test on the strict side. It turned the sandbox on, chose “don’t ask,” the most restrictive permission mode, opened a repository, and sent one short message. That alone, it says, was enough for a command originating from the repository to run outside the sandbox with no confirmation1. The company adds that the same thing could be triggered through indirect prompt injection, where instructions hidden in external content steer the AI.

As for why no prompt showed up, Accomplish points to the sandbox itself, calling this counterintuitive. With the sandbox enabled, Claude Code confines the commands the agent runs using macOS Seatbelt, and in exchange stops asking before running them1. The design reduces confirmations on the assumption that the agent is contained, so a hole in the containment leads straight to execution without confirmation.

Four Defenses, and a Way Around Each

The hole was not in the agent’s commands but in Claude Code itself, the harness. To build an index of the repository, the harness runs its own git commands in the background, outside the sandbox. Git has a setting called core.fsmonitor; when it is present in .git/config, git executes it as a shell command every time it inspects the working tree1. In other words, getting the harness’s git to read a .git/config carrying this setting means a command runs outside the sandbox.

Anthropic had defenses in place for this route. Accomplish sorts them into four and shows how each was sidestepped1.

The first defense added flags that disable core.fsmonitor and similar options to the git commands run outside the sandbox. But one call, git ls-files, which refreshes the file index, had been left out. The second blocked writes to the .git folder, yet it only protected the .git at the project root; there was no Seatbelt rule preventing a folder inside a subdirectory from being renamed to .git. The third was the assumption that the harness runs git at the project root, so an untouched root .git would be enough. In practice, git ran in whichever directory the Bash tool had used most recently. The fourth was that the harness does not rerun git on its own. That was undone by automatic skill loading: when Claude reads a file, it looks for skills in that folder and loads them, and loading a skill triggers a refresh of the file index.

Chained together, a setup script creates git configuration in a folder with a different name and writes core.fsmonitor into it, renames that folder to .git inside a subdirectory, moves the working directory there, and ends by prompting Claude to read a build report in the same folder. According to the company, once Claude reads the report, the rest proceeds with no action from the user1.

From Report to Fix

In the timeline Accomplish published, it reported the issue to Anthropic on July 13, 2026, and it was triaged that same day. A first round of hardening shipped in 2.1.223 on August 6, but some git calls were missed and the escape moved to a different call, so Accomplish sent over the remaining ones. The complete fix came in 2.1.247 on August 261. The company writes that core.fsmonitor is now emptied for all git commands issued by the harness, leaving a repository’s settings unable to execute anything, and it also describes Anthropic’s triage as fast.

Claude Code’s CHANGELOG does have entries for 2.1.223 and 2.1.247, but the word fsmonitor does not appear anywhere in the CHANGELOG2. This fix is not named there. Anyone who tracks security fixes through the CHANGELOG alone would miss this kind of change. Accomplish’s post does not say whether the flaw was actually exploited.

According to reporting by Upstarts Media, Accomplish also reported issues to Cursor and OpenAI this summer, in addition to Anthropic. The one issue reported to Cursor in July and the two reported to OpenAI were reportedly fixed in about a week3. After the article was published, an OpenAI spokesperson issued a statement saying both issues had been addressed in August. Anthropic and Cursor did not comment on the record.

What to Check on Your Own Setup

Claude Code has been going through a series of changes to how approval itself works. Starting August 14, new sessions on the Pro, Max, and Team plans default to auto mode, replacing per-call confirmation with blocking by a classifier. Version 2.1.251, added on August 28, also fixed problems including one where files outside the approved scope could be read or written through symbolic links. In designs that cut down on confirmations, whatever runs outside those confirmations, in this case the harness’s own git, becomes the premise that safety rests on. Beltdown can be read as a case where that premise broke.

In practice, the response comes down to three things. The first is checking your version: confirm that your Claude Code is 2.1.247 or later. Environments that have turned off auto-update, or that distribute a pinned version internally, need particular attention. The second is how you handle unfamiliar repositories. When opening code you cannot vouch for, such as when reviewing open-source projects or checking code received from outside, do not assume the sandbox makes it safe; consider opening it in a disposable VM or container instead. Accomplish itself says it runs the entire agent inside a VM and keeps real credentials out of that VM1. The third is how you follow information: watching what the people who find the flaws publish, in addition to the CHANGELOG, reduces what slips through.

Sources

  1. Beltdown: Escaping the Claude Code sandbox - Accomplish official blog (September 11, 2026; technical explanation and timeline from the researchers who found it)
  2. Claude Code CHANGELOG - Anthropic official repository (entries for 2.1.223 and 2.1.247)
  3. Claude Code, Codex, And Cursor Have Leaky Sandbox Problems You Don’t Hear About - Upstarts Media (September 10, 2026)

We publish the latest AI news nearly every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →