Studio

Studio Code

A coding agent you can leave alone, because of what it cannot reach.

Each project gets its own container that survives between sessions. Clone a repository or upload a zip, then plan, build, review and preview — with the agent’s reach decided by a switch rather than by a paragraph in a prompt.

The modes are tool sets, not prompts

In Plan mode the agent is handed nothing that can write a file. Not told not to — not given the tool. A model that has been argued, jailbroken or prompt-injected into trying has nothing to call. That is the enforcement; the sentence in the system prompt exists only so it does not waste three steps discovering it.

Ask is the same idea taken one step further, and it is the mode for the middle of a task — where a setting lives, what a value should be, which file does the thing. It reads your project and answers. It is also the only mode with no way to write the plan, which is what makes it safe to interrupt yourself with: stopping to ask a question cannot cost you the plan you were working through.

The agent canAskPlanBuild
Read, search, git diffYesYesYes
Run read-only commandsYesYesYes
Change the planNo tool for itYesYes
Write and edit filesNo tool for itNo tool for itYes
Install, build, run testsNo tool for itNo tool for itYes
CommitNo tool for itNo tool for itYes
Delete, push, anything unrecognisedNo tool for itNo tool for itAsks first

Then decide how much runs unattended

Plan and Build decide which tools exist on a turn. The approval mode decides which of the calls they make may run without you. Every command is classified before it runs, and the mode is a threshold over that.

Auto

Ordinary work runs. Deleting, pushing and anything unrecognised still stops and asks.

Best when you are watching the transcript and want throughput.

Manual

Every change asks first — each file write, each install, each commit.

Best on a repository you care about, or the first time on a new one.

Read-only

Nothing that changes anything is permitted at all.

Best for understanding a codebase you have just been handed.

Anything genuinely dangerous — sudo, curl … | sh, touching ~/.sshor git hooks, writing outside the workspace — is refused in every mode including Manual. A prompt you can click through is a prompt you will click through.

Say this plainly: the container is the boundary

Your machine is never in reach. The work happens in a container that runs unprivileged, drops every capability, cannot gain new ones, and has memory, process and file-size ceilings. It is firewalled off the internal network, so one project cannot see another’s, the database, or the host.

What the approval classifier protects is something narrower and still worth protecting: your code. It is the difference between an agent that runs rm -rf and one that asks first. Both are contained; only one costs you an afternoon.

Everything the agent runs is written to an audit log with its exit code and duration — including the calls that were refused, which is exactly the sort of thing worth being able to look up.

Four views of the same workspace

Files

The real tree in the container, expanded a directory at a time. Open anything to read it, even while the container is asleep.

Plan

Written by the agent itself as it works, not summarised afterwards by a second model. On a coding session the plan is the long-term memory.

Changes

git diff HEAD, the same read the agent makes before it commits. Reviewed by a separate model that sees the diff and nothing else — a reviewer told why the code is right tends to agree that it is.

Terminal and preview

Commands replay live, because a four-minute install is unwatchable inside a collapsed block. And the dev server renders in the page, sandboxed, so you can look at what was built.

A different model for each job, and a bill that separates them

Set one model to plan, another to write, a third to review. Planning rewards a model that thinks; reviewing rewards a sceptical one; most of the tokens go to the coder re-reading files, which is where a cheaper model pays for itself. The usage page splits spend by role, by project and by model, so the trade is one you can actually see.

Studio Code needs a topped-up account: it uses server CPU rather than model tokens, so your own API key does not cover it.