nanoMuse is an open-source (GPL-3.0) personal agent that puts a full agent on each of a person’s devices, with on-device screen control, a policy gate called Sentinel on every action, memory as plain files, and a bring-your-own model, as an open counterpart to Meta’s closed Muse.
A personal agent is a program that acts for one person across their accounts and devices for weeks, remembers them, speaks first when it matters, and keeps a log of what it did. In September 2026 Meta launched Muse, the first product to assemble these pieces. Within three weeks OpenAI, Manus, and a startup called Today shipped similar things. All of them run one Linux VM per person in a vendor’s cloud, with the model trained by the vendor and the code closed. Only a small gadget SDK was opened.
The authors argue the shape is wrong: an agent that holds a person’s life should be inspectable by them, run on hardware they own when possible, and use a model they chose. The cloud-VM agents can also only reach things with an API or a web page, which leaves out most of the apps on a phone, from banking to government services. That gap, plus the lack of an open counterpart to Muse, is what nanoMuse exists to fill.
nanoMuse is one agent replicated on each device the person owns (Android, iOS, macOS, Windows, Linux, web), with a thin shared relay so the devices can see one conversation. Each device runs its own agent loop and its own policy gate, so one device can work with no server at all.
Four design choices carry the system:
•
Sentinel: the agent never calls a tool directly. Every call passes through a gate that checks, in a fixed order, denied tools, the person’s rules, allow/ask lists, a taint flag (once a conversation has read private data, any outbound call becomes an ask), the tool’s risk level, and a few hard warnings (destructive shell commands, a committing word like “confirm payment” in a button label). Approvals are scoped grants: once, this conversation, or always for a named recipient/site/folder. Payments and passwords are never remembered.
•
Hands: to operate an app, the agent climbs a ladder and takes the lowest rung that works: a skill, a CLI tool, an MCP (Model Context Protocol) server, a web fetch with the user’s login, a browser, and last the device’s own screen. The screen rung uses a one-action-per-screenshot loop ported from UI-TARS-desktop on computers and from OpenMinis on Android. On iOS there are no hands, because iOS forbids one app from operating another.
•
Memory as files: Markdown files the person can open and edit (SOUL.md for the agent, USER.md for the person, GLOBAL.md for durable facts, a dated diary, HEARTBEAT.md for routines), plus a line-level store with undo on the desktop.
•
Model is the person’s choice: any OpenAI-compatible endpoint, a ChatGPT plan via OpenAI’s Codex auth flow, or a local model server. The relay never sees tool outputs or files, only the text of synced conversations if the person opts in.
The screen-control loop looks roughly like this:
while not done:
screenshot = device.capture()
a11y = device.accessibility_elements() # if available
thought, sentence, gesture = model(goal, screenshot, a11y, history[-1])
decision = sentinel.check(gesture, sentence, a11y_at_target)
if decision.needs_approval:
if not user.approve(sentence, scope=decision.scope): break
device.perform(gesture)
history.append(sentence)
The sentence the model writes before acting (“I’ll tap Confirm payment”) is what makes the step governable: a deterministic policy cannot judge a picture, but it can scan words for commit verbs and force an approval.
This is a system report, not a benchmarked contribution. The paper explicitly says it has no success rate and no count of human take-overs to report for the Hands, and would rather say so than estimate. The claims it does make are about shape and cost.
Size and cost (from release files and provider price lists, not from usage measurement):
•
Android app: 38 MB download; desktop app: 256-498 MB download, about 928 MB installed on Linux, roughly 0.5 GB idle memory.
•
Relay: one Python process plus a SQLite file, about 85 MB idle, sized for a 1 vCPU / 1 GB server.
•
Self-hosting a relay: on the order of ¥30-¥60/month in mainland China or US$4-6/month elsewhere for the smallest tier, plus model usage. A day of chat is estimated at “a few fen” on Bailian’s October 2026 prices.
Against Muse, nanoMuse adds hands on the Android screen, on-device execution (no per-person VM), full openness including the relay, and model choice. It drops Muse’s per-person VM with kernel-level taint tracking and surrogate credentials, Muse’s wallet, and the polish of a large team. Against OpenMuse (CopilotKit’s open server-with-windows design), nanoMuse’s agent lives on the device, which is what gives it the phone’s screen.
The reading of Muse itself is sourced carefully: each claim is tagged documented, prompt (read from a copy of Muse’s October 2026 production system prompt), observed, or inferred. The authors do not reproduce the prompt and quote at most a few words at a time.
•
If you want to understand the 2026 personal-agent category, the paper’s framing (five questions: what it knows, where it acts, whom it answers to, how it learns, how it is paid for) and its table of Muse / Today / Manus Cue / dots / OpenMuse / nanoMuse is a compact map. Worth reading before building anything in this space.
•
If you want an agent that operates native phone apps (not just web and APIs), the on-device screen loop is one of the few open options on Android. On iOS the platform rules out screen control for anyone, so plan around that.
•
If you are designing an approval system for a tool-calling agent, the Sentinel’s decision order (deny / user rules / allow-ask / taint / risk / hard warnings) and the “model narrates its action in one sentence, policy judges the sentence” pattern are directly borrowable. The authors note this is a policy boundary, not a privilege boundary like Muse’s: a compromised device compromises the agent, so do not treat it as a sandbox against a hostile model.
•
Worth testing: running nanoMuse against public suites like AndroidWorld, OSWorld, MemGUI-Bench, and OS-Harm. The authors list this as roadmap work and have not yet published numbers, so any evaluation would be new information.
•
Artifacts are released: GitHub, project site at nanomuse.cn, with Android, iOS (TestFlight), desktop, web, and the relay all under GPL-3.0.
•
No measured performance of the Hands. Success rates, take-over counts, and benchmark scores do not exist yet in the paper.
•
The Sentinel runs in the same trust domain as the agent on each device. There is no outside host holding real credentials, so device compromise means agent compromise.
•
Memory files do not yet track which model wrote a line, when, or with what confidence, and there is no rule for re-checking stale lines. The authors call it “notes of a colleague one does not yet fully trust.”
•
Dense professional screens where buttons have no labels or accessibility text defeat the policy check; the person watching is the last line. On Linux, Hands work only in X11 sessions.
•
The community relay stores the text of synced conversations on a server the user does not run, and the “help improve nanoMuse’s AI models” switch is on by default on that relay. The authors say this is the choice they are least sure of.
•
The account of Muse rests on four public documents, the shipped clients, and a leaked production prompt whose provenance the authors cannot independently verify; prompts also change between deployments.