Wain AI/Tech Blog

AI news and trends worldwide, updated nearly every day

Meta Launches Personal AI Agent 'Muse' in the US — A Design That Puts Permission Decisions Outside the Model

Meta Launches Personal AI Agent 'Muse' in the US — A Design That Puts Permission Decisions Outside the Model

Meta launched Muse in the US on September 8, 2026, a personal AI agent that operates email, payments, and a browser on a user's behalf. A technical post published the same day discloses the design: a separate agent called Sentinel holds sole permission authority on a dedicated VM, and the agent never sees real credentials. The bug bounty pays up to $300,000.

Meta launched its personal AI agent Muse in the US on September 8, 20261. It is said to take on sending emails, booking travel, opening a browser and filling out forms, and negotiating on a user’s behalf, and the company describes it as a “secure, private personal AI agent.”

This is a consumer product announcement, but for anyone building an agent themselves, the substance is in the technical post published the same day2. What was given up, and what was moved outside the model, in order to make an agent that touches email, payments, and a browser work at all — that is disclosed at the implementation level.

What It Does, and What It Costs

Muse runs on a dedicated virtual machine called Muse Secure VM. The company explains this is “because personal agents need a new kind of secure computer,” and both the agent and the person’s data live on that VM1. Besides the Muse app, it can be used by talking to it directly in WhatsApp. The model behind it is Muse Spark, which the company positions as its most capable model to date.

For tasks that take more time, it keeps working after the app is closed and comes back when something changes or when it needs approval. The examples given for approval are sending an email and making a purchase1.

Checkout can be done with Link, built by Stripe. The company says Muse is “the first AI agent covered by Link’s purchase protections,” which include free coverage for damaged or lost items, price drops, no-fee returns, and a return guarantee on eligible purchases1. Link’s wallet for agents generates a one-time-use card, so real card details never reach the merchant. Shop Pay and 1Password support are described as coming soon.

Availability is the US, rolling out on iOS, Android, and muse.ai. AI glasses support is said to be coming soon1. On pricing, the official announcement says only that it is “free for most of what people need, with subscription plans for people who want to do more.” According to TechCrunch’s reporting, two paid plans are available at launch — Power at $20/month and Maximum at $100/month — and Meta believes most people will remain on the free tier3. The publication also reports that a payment card is required to get started, and that the app includes a usage meter showing what percentage of usage remains.

Separating the Permission Authority from the Model

The technical post is written by Tarek Sheasha of Meta Superintelligence Labs, and it states its design premise up front: design the system on the assumption that the agent may be under attack, and limit the potential damage2. The harness runs in its own isolated cell, it does not see real credentials, and every interaction with the outside world runs through a Sentinel that the agent cannot override.

That Sentinel is a separate, host-side agent, distinct from Muse. It is described as the sole permission authority for connector actions and for all network egress. In the company’s own words, Muse proposes actions, but only Sentinel can grant permission to perform them2.

The isolation is spelled out concretely. The agent harness and the working filesystem run inside a systemd-nspawn runtime container, and root inside that container is mapped to an unprivileged host user — so root in the cell is not host root2. Security-sensitive services sit outside this runtime cell: an independent set of models and classifiers that detect prompt injection and other threats (hatch-safety), workers that execute connector code with tightly scoped privileges (privsep), credential storage (hatch-authd), and Sentinel. The stated reason for that placement is that an attacker cannot disable the protections themselves.

The handling of credentials is the core of this design. Code inside the runtime cell only ever sees a “surrogate” token minted by authd, and the real credential is substituted by Sentinel at the network boundary. The company therefore argues that any attempt to coerce the agent into revealing actual secrets — through prompt injection or otherwise — is futile2. The official announcement puts the same point in plainer language: Muse has no visibility into people’s passwords or payment methods1.

Treating Approvals as Capabilities, Not Conversation

How human approval is handled also departs from a normal chat UI. When Sentinel resolves to ask the user, the approval dialog is presented directly within the client UI, not via the conversation with Muse, and the answer is routed straight back to Sentinel2. Presumably because taking approval inside the conversation exposes the approval itself to manipulation.

The granularity of approval is defined as well. Grants are strict capabilities rather than conversational suggestions, bound to a particular connector or destination and use case. The user chooses among one-time, session-scoped, task-scoped, time-bounded, and perpetual permission2.

Asking for approval too often, however, makes a system unusable. Hence “tainted egress,” a kernel-level data flow tracking mechanism. Each tool execution process starts clean and becomes tainted the moment it reads user data. Tainted or unverifiable processes lose auto-allow and fall back to the normal approval flow2. The implementation uses eBPF cgroup programs and Linux Security Module hooks the team added.

The email connector reflects a different kind of trade-off. It filters out one-time tokens, password reset links, and login magic links using both deterministic filters and a classifier model2. The point is to keep “letting it read your mail” from becoming a path to taking over your other accounts.

Putting human approval in front of outbound actions is not unique to this product. In September, Google let Gemini Spark operate Google Photos, asking permission each time a shared album is created or an email is sent, and editing by making a copy rather than overwriting the original. What differs with Muse is that the decision is physically pulled out of the same process as the model and enforced by OS-level isolation.

Layered Defense Against Prompt Injection, and a Bug Bounty

Prompt injection defenses are described in four layers: training the model itself, harness-level labeling of externally sourced data as untrusted, an ensemble of detection classifiers, and human approval for actions that move data out of the VM2. The company rates Muse Spark 1.3 as “close to SOTA” on this capability. That model is the same one that appeared as the latest version in Meta’s cheaper “data-for-access” tier in early September.

Browser-side limits are explicit too. The sub-agent driving the browser sees an accessibility tree snapshot rather than the raw DOM, cannot run JavaScript in the page context, and Chrome devtools are disabled2. For payments, on sites where payment credentials are already on file, checkout pages are detected and an approval showing the exact purchase details is requested every time. On new sites, a single-use card number is issued, tied to that merchant, that amount, and a limited period.

On top of that, the company opened its bug bounty to the public the same day. It pays up to $300,000 for valid reports, including up to $130,000 for successful prompt injection attempts that affect one user2. Publishing defense claims and bounty amounts together is a posture that assumes external verification.

What Is Guaranteed, and What Is Not

Worth reading carefully is that the company states the limits of the current architecture itself. The technical post says that while access by Meta personnel is restricted through operational policies, this does not prevent Meta from accessing data when necessary to support, secure, or operate the service2. In other words, “private” today is an operational restriction, not a cryptographic impossibility.

What is meant to change that is Muse Confidential VM, planned for later this year. The whole VM would be encrypted with a key only the user holds, so that not even Meta can access it1, and the design and source code have begun to be made available to external auditors, with a continuous audit inspectable by anyone once it launches2. Whether it actually ships this year is what will decide how this product is judged on privacy.

Phrases like “the world’s first personal AI agent built for everyone” and “our most capable model to date” are Meta’s own claims. TechCrunch adds the caveat that the security claims will require deeper investigation by security experts, and notes that the announcement came less than two weeks after Meta agreed to an $18 billion multistate settlement in a lawsuit over social media’s consumer harms3.

For anyone assembling an agent in-house, this technical post is useful less as competitive analysis than as a design reference. Move the permission decision outside the model, keep credentials invisible to the agent, treat approvals as capabilities and detach them from the conversation — those three are design choices that can be adopted independently of which model is used. Conversely, anyone weighing Muse for work use should copy two conditions straight into their evaluation table: US-only availability, and the fact that until Confidential VM ships, the protection remains operational rather than cryptographic.

Sources

  1. Introducing Muse: The World’s First Personal AI Agent Built for Everyone - Meta Newsroom official announcement (September 8, 2026)
  2. How We Built Safety Into Muse - Meta AI Research, technical post by Tarek Sheasha (September 8, 2026)
  3. Meta debuts its Muse AI agent. Will consumers trust it? - TechCrunch (September 8, 2026)

We publish the latest AI news nearly every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →