# Where One Tenant Ends: Isolation on AgentCell

> We assume the code in a cell is buggy and the agent that wrote it may be hostile. The five layers between one cell and everything else, what each refuses, and how we check that it can fail.

Published: 2026-09-24  
Canonical: https://agentcell.dev/blog/tenant-isolation  
Markdown: https://agentcell.dev/blog/tenant-isolation.md

Most of the code AgentCell runs was written by an agent in an afternoon, and nobody reviewed it. Some of it will have bugs a stranger could exploit. Some of it may be deliberately hostile: anyone can sign up, and an agent with a shell is a realistic adversary in its own right.

So we assume whatever is inside a cell is compromised, and ask what a compromised cell can reach. This post goes through the layers between one cell and everything else, roughly in the order a request or an escape attempt would meet them. For each one we say what it refuses and how we check that it really does.

We try to keep what runs today apart from what is only designed. The things that aren't built yet are listed at the end, and if you're evaluating us you may want to skip there first.

## Layer 1: a separate kernel

Every cell that runs code runs as a Kata Containers microVM. (A static site runs none: the edge serves its files, and there's no process of yours on our machines to escape from.) It's still a container image, but it boots inside a lightweight virtual machine with its own guest kernel instead of sharing the worker's kernel through namespaces.

We measure this rather than take it on faith. An isolation check runs inside a live cell and compares what it sees with the worker around it. The kernel version inside the cell is the guest's, and differs from the worker's. The cell sees three or four processes, where the worker has about a hundred. Files that exist on the worker aren't visible from inside. All of it still holds after the cell is moved to a different worker.

The control run is more interesting. Run the same app as an ordinary container, with the `runc` runtime, and the check fails, though only just: the ordinary container passed five of the six checks and failed only on the kernel. A shared-kernel container can hide processes and files from itself quite convincingly. It can't hide the kernel, and the kernel is what an escape goes through.

The scheduler configuration enforces the runtime. Kata is the only runtime permitted and also the default, so leaving the runtime unset doesn't fall back to something weaker, and the ordinary runtime has been removed from workers entirely. Privileged containers and host volumes aren't allowed. Building an image needs more privilege than running one, so builds get it only on separate build machines that never run customer apps, and the build itself runs inside a microVM as well. That includes the `npm install` and `npm run build` behind a static site.

## Layer 2: one door into the scheduler

A microVM only helps if every workload actually gets one. Nomad, our scheduler, will run whatever job it's handed, so nothing talks to it directly. Its API is firewalled off, and the only way in is an admission proxy that inspects each job before Nomad sees it.

The proxy refuses any task that doesn't run under Kata, and any task that doesn't join the cell's own private network (host networking is refused by name). Scheduled (cron) jobs have to be exactly the shape we allow: overlapping runs prohibited, a fixed time zone, a run-time ceiling, no services, no ports. We built this check test-first: thirteen deliberately wrong scheduled jobs, each seen getting through before the check existed and refused once it did. They stay in the test. Requests to force a scheduled job to run early are refused, whoever sends them.

It also blocks two scheduler endpoints that answer without a credential, one of which lists every running workload. Paths are normalised before matching, so `//v1/metrics` doesn't slip past a check for `/v1/metrics`.

We treat the proxy as the boundary and assume our own client can be bypassed. A customer who pulls a deploy credential out of our tooling and hand-crafts a job still meets the same refusals. The same checks also run on every rendered job before we submit it, so a bad template fails at our end first. We've watched each of those checks fail.

## Layer 3: a network of one

Each cell gets its own network bridge, with nothing else attached. Separate bridges don't route to each other, and on every worker a firewall rule drops traffic from any cell bridge to other cells, to anything on the platform's private overlay network, to RFC 1918 private address space, and to the worker's own services, including its DNS.

The rule matches cell bridges by name prefix, so there's no list to update after each deploy. A new cell is covered as soon as its bridge exists, with no gap while a rule catches up. To check it, we read the drop counters from inside a live cell instead of reading the ruleset: attempts to reach the scheduler, service discovery, the registry and storage are all dropped, and the counters go up.

Two lint rules guard this layer. Firewall chains may not use a drop policy, because a drop policy on the wrong hook cuts off the whole machine rather than the cell. And every job must declare its network, since a job that forgets gets a default network the firewall doesn't cover. Both lint rules have tests that make them fail.

The next step for this layer is outbound traffic. A cell can reach the public internet, and there's no per-cell allowlist and no default-deny for outbound connections. We've built and measured a default-deny design with an allowlist in a prototype, but it isn't live. Until it is, think of a cell as a machine that can reach anything on the internet and nothing of ours or anyone else's.

## Layer 4: the front door

A cell isn't reachable until something decides who may reach it. Before the gate existed, the wildcard hostname for cells answered every request with a 404, and we didn't serve a single app publicly until authorization was in place. From the outside, a cell that answers a stranger with a 200 looks exactly like one that answers its owner.

Every cell has its own Cloudflare Access application, listing the people it's shared with. We could have used one wildcard application for all cells, but a wildcard authenticates people, not tenants: it would let any signed-in customer open any cell. Sharing is just that list. Adding or removing a person changes it without a redeploy, and removing the last person puts an explicit deny-all policy in place.

The edge router checks the signed identity assertion itself: the signature against published keys, a pinned algorithm, issuer and expiry. It then strips cookies, `Authorization` and the assertion before forwarding, so the cell never holds a credential. A cell name that doesn't exist gets the same answer as one that does.

The router's first version, early in development, only checked that the identity header was present. Our own adversarial test caught that the day it was written: with a test cell's login application removed, a header set to the literal string `forged` was accepted. Full signature verification replaced it that same day, long before any app was served publicly.

Every configuration run now sends a forged assertion and one claiming the `none` algorithm, and both must be refused.

Static sites get one more check. A valid signature proves the assertion came from our Access account, not which cell it was issued for, so for a static site the router also requires the assertion's audience to name that cell's own Access application. We tested it live: another organisation's genuine sign-in, replayed at a static site, gets a 403 and none of the site's bytes. Web cells don't have this check yet. For them, the cell's own Access application is the only per-cell binding until the router's check is extended to them. The episode also gave us the rule described further down: a test suite that has never sent a bad header can't tell you the check works.

## Layer 5: the API

The control plane's API is where a token turns into an action. Tokens are stored as salted scrypt hashes, and we checked the whole table: none can be recovered from it. Comparison is constant-time, and a lint check catches any code that goes back to comparing with `==`.

Every operation is bound to the token's org. Ask for another org's cell and you get the same "not found" as for a cell that doesn't exist, byte for byte. Scopes are enforced separately from identity, so a deploy token asking for an admin operation gets `forbidden`, a different answer from `unauthenticated`. Revoking a token takes effect on the next request, with no cache to wait out.

Rate limits are per credential and apply only after authentication. In the other order, someone with no credential could make us allocate a rate-limit bucket for every token they care to invent. Forty unauthenticated requests get forty `unauthenticated` refusals and allocate nothing, and one token being throttled doesn't affect another in the same org. Authentication itself runs a bounded number at a time and refuses beyond that rather than queueing. A per-address limit at Cloudflare's edge sits in front of all of it.

Every refusal has a fixed, typed error code, and a contract test pins those codes against the public client.

Our own operational secrets are encrypted in the repository with SOPS and age keys, and reach machines only as files readable by root. The control plane stores no customer credential, just token hashes, which it can't replay.

## How we know a check works

The forged-header test gave us one rule, and we apply it to every security check we add:

> **A check counts only after it has been seen failing** against the code it protects: the unfixed file, a deliberately broken copy, or a live gate we removed on purpose.

The cross-tenant test is the clearest case. It removes a real gate in front of a live cell, confirms the test now fails, then puts the gate back. On its last run it passed six of six, reading the marker each cell served, with the negative control live. It now includes a static site, where the other org's credential must get a 403 with no marker and no bytes.

Authentication code gets mutants, for example a verifier that only checks a header is present, and each mutant has to fail the suite. Test doubles are held to the same standard. A review found a test double enforcing database rules on its own, which would have hidden a broken query, so we made it stop. Three mutants that break the real SQL now fail.

Every security-relevant change also gets a separate adversarial review before it merges, from a reviewer whose only job is to break tenant isolation, authentication or scoping. On our device-login flow, the first round found that a poll with a caller-chosen value could collect someone else's token. That would have been a cross-tenant credential mint. We redesigned it before the flow was ever switched on.

A test that has only ever passed hasn't shown it can catch anything.

## What's next

A few pieces are designed and on their way:

- Cells can reach the public internet. Allowlists and default-deny for outbound traffic are designed and prototyped, but not enforced.
- There's no way to inject secrets into a cell yet. Keep keys out of your source anyway.
- Spend caps and per-cell access logs are designed. Neither is built.
- For abuse, per-account ceilings and takedown tooling are designed. The only live caps today are on scheduled jobs per org and on static sites, at most 10 per org.
- The router's per-cell token check covers static sites only. Web cells don't get it yet.
- We don't add a Content-Security-Policy to responses from your app yet. Our own login pages carry a strict one.

When one of these ships, it will go through the same testing as the layers above, and we'll write it up.

---

*Part of a series. Start with [what happens when you run agentcell deploy](/blog/how-a-deploy-runs). Next: [breaking it on purpose](/blog/breaking-it-on-purpose). Or [deploy now](/docs/deploy/).*
