How distributed Claude Code is built with nfltr

Claude Code already knows how to split up work. It starts subagents, lets them run in the background, and picks up their results when they finish. That works well as long as everything the work needs is on the machine where Claude Code runs.

Often it is not:

  • The dataset lives on a VM and should not leave it.
  • The failing service is only reachable from inside one network.
  • Three independent fixes would go faster on three machines at once.
  • The machine doing the work gets restarted halfway through.

nfltr keeps Claude Code's model and removes the one-machine limit. Here is how it is built, what went wrong along the way, and how we test it.

The idea: Claude Code's hub, made distributed

In nfltr, one Claude session is the hub. It gets a small set of tools that mirror Claude Code's own subagent model:

  • spawn_agent starts an agent from a brief and returns its id at once.
  • send_message reaches an agent. A running agent gets it in its live session; a finished one continues in the same session and workspace.
  • stop_agent cancels an agent. Its partial work comes back as evidence.
  • list_agents shows each agent's status, machine, progress and last activity.
  • wait_for_agents blocks until results arrive and returns them.
  • list_nodes shows the joined machines, their descriptions, and how many more agents each may start.
  • start_monitor runs a command on a machine without a model. Each line it prints becomes an event for the hub.

The difference from subagents is where the agents run: on any machine you have joined, and they keep running when the hub goes away.

nfltr orch "<goal>" starts Claude Code as the hub in your terminal. Or add the hub to your own session with claude mcp add nfltr -- nfltr mcp --toolset hub; with nfltr orch hub install-hook, a finished agent wakes an idle session by itself.

Machines join once, with nfltr node join --max-agents 2. A node stays idle until a spawn needs it, then launches a Claude Code agent for that task; the agent exits when the task ends. One join per machine, no configuration per task.

How the pieces connect Your terminal or Claude Code session is the hub. It and every joined machine (a laptop, three QA VMs, an app VM and a private database VM) open outbound connections to the nfltr.xyz relay, which carries messages and runs no model. Each machine shows its agent slots. Your terminal or Claude Code session the hub: decides what runs where outbound nfltr.xyz relay carries messages, runs no model outbound laptop the checkout, git push agents 1/2 QA VMs ×3 nightly suites and logs agents 1/2 each app VM the service, its logs agents 0/2 database VM private data stays here agents 1/2
Every arrow is a connection its source opens toward the relay; no machine accepts inbound connections. Filled dots are running agents, hollow dots free slots under the machine's --max-agents. The machines are the ones from the demos.
  1. The hub calls spawn_agent with a brief and, say, a machine, and gets an agent id at once.
  2. No agent is free there, so the hub asks that machine's node, through the relay, to launch one.
  3. The node starts a Claude Code agent for this task.
  4. The agent connects out to the relay and registers.
  5. The hub sends the brief; the agent decides how to do the work.
  6. Progress comes back; list_agents shows it.
  7. The result, with any branch and commit pushed, comes back.
  8. The hub stores it and marks it delivered as wait_for_agents returns it: exactly once, even across a hub restart.
  9. The agent exits and the slot is free again.

One agent on demand, from spawn_agent to its result. Solid arrows are messages; dots on the relay's line mark messages it carries.

Design principles

Provide, don't decide

nfltr never picks a machine, a model or a plan for you. A machine or labels named in a spawn are hard filters; among the matches, the runtime only finds a free slot. If several machines could run a monitor, start_monitor asks for one to be named rather than choosing. Model, effort and budgets pass through when set; unset, Claude Code on that machine decides.

The same goes for what we tell the model. In one campaign the hub spawned agents in only 2 of 7 passing runs. Instead of an instruction to delegate, we added facts to the tool results: free capacity, and that agents run in parallel, each with its own workspace and tokens. The choice stays with the hub.

Capabilities are opt-in per machine

What an agent may do on a machine is that machine's decision:

  • Without --allow-all-tools, a node's agents can edit files, but commands that need approval (shell, git push, reading a path outside the checkout) are refused, and the turn fails with tools_not_allowed, naming the flag.
  • Monitors run only with --allow-monitors: they use that machine's credentials.
  • A node clones only repositories it lists with --allow-repo.
  • The API key comes from the environment or the saved config, never from the command line.

Every result arrives exactly once

Agents keep running when the hub disconnects. The hub stores each completion and marks it delivered before wait_for_agents returns it, so a hub that reattaches with the same id gets everything it missed, none of it twice. Monitor events work the same way. In your own Claude Code session the hub id comes from the project directory, so reopening the project reattaches.

  1. The hub calls start_monitor on a node joined with --allow-monitors.
  2. The node runs the command; no model runs while it watches.
  3. The session ends its turn and sits idle.
  4. The command prints a line; the node numbers it and keeps it until acknowledged.
  5. The hub's long-poll pull through the relay carries its cursor.
  6. The node answers with the new event.
  7. The hub stores it before its next pull acknowledges it, so nothing is lost or repeated.
  8. The session wakes, and wait_for_agents returns the line as untrusted data.

A monitor line waking an idle session. Watching costs no model calls; each event that arrives starts one turn.

Each hub is isolated

A hub sees and controls only its own agents and monitors. A second session in the same project cannot share its hub; it fails and names the holder.

Agents get a clean Claude config

In real-Claude runs, agents picked up the host user's hooks and memory, and one used Claude Code's cross-session messaging to ask another local session for data. Now agents load no user settings, hooks, memory, plugins or MCP servers from the host (the workspace's project context still applies), and cross-session messaging is off. The node verifies both before any task; each is an opt-in per node.

What the relay does, and doesn't

Every node, agent and hub opens a long-lived connection out to the relay, over TLS, authenticated with your account key. No machine opens an inbound port, joins a VPN or needs a firewall change, so it works behind NAT. The relay routes each message to the agent it is addressed to, stamping the sender from the authenticated connection so nobody can speak as another agent. It runs no model, runs none of your code, and makes no decisions: it never picks a machine, a worker or a retry.

How a result survives a relay restart

Exactly-once delivery is a chain, not one component. The hub keeps its agents' state in its own durable store; each agent keeps its finished result until the hub acknowledges it; every task's events carry a gapless sequence number; and a resumed stream replays what was missed and then says so explicitly.

  1. The hub dispatches a task to an agent through the relay.
  2. The agent accepts it and streams progress, each event numbered.
  3. The relay restarts (a deploy, a crash).
  4. The agent keeps working; its machine has everything it needs.
  5. Hub and agent redial at once; a spawn issued meanwhile waits for them instead of failing.
  6. The hub resumes the task from the last event it has.
  7. The agent replays the events the hub missed, then marks the replay complete.
  8. The result arrives.
  9. The hub stores it and acknowledges; only then does the agent drop its copy and exit.

A task across a relay restart. In the soaks, a relay killed or restarted at a random point had its nodes back in about a second (p50), and every agent still completed exactly once.
More of how it works: ownership leases, the capacity feed, idle behaviour, deploys
  • One owner per task. When processes share a task store (relay replicas, a restarted hub), each task records an owner and an epoch, and each process holds a renewed lease. A task whose owner's lease lapsed can be taken over with the next epoch, and the store rejects any write from a stale owner, so a restart or a second replica never runs a task twice or loses it. A hub restarted after kill -9 takes the dead process's tasks over at once.
  • No polling for capacity. The relay publishes a small feed of changes (a slot freed, a worker or node connected); a spawn waiting for a machine wakes on the event. With nothing happening, the feed costs nothing.
  • Holding for a reconnecting agent. A dispatch or resume for an agent that is reconnecting is held briefly rather than failed.
  • Artifacts. Where artifacts pass through the relay's store, they are addressed by SHA-256, checked chunk by chunk and as a whole, and resumed from the last acknowledged offset; a mismatch fails closed.
  • Dormant when idle. Periodic work first checks a cheap revision number and reads nothing if it has not changed. Idle after the final hub-on-nodes soak, the relay used 0.23 % CPU (2026-09-28); the production relay idled at about 100 MB and 1 % CPU (2026-09-27).
  • Deploys. nfltr.xyz runs on one VM. A deploy drains, restarts and waits for health before taking traffic; clients ride out the gap for up to 2 minutes, and nodes redial within a second, jittered so a fleet does not reconnect in lockstep. The site is served directly rather than through a CDN proxy, which cut long polls after 100 seconds and answered every request with an error during deploys.

What the relay can see

Connections are encrypted to the relay with TLS. On top of that, the frames between the hub and its agents and nodes (briefs, progress, results, artifacts, monitor lines) are encrypted end to end with X25519 and AES-256-GCM by default, so the relay forwards ciphertext. That is not the whole picture, and we would rather you hear it from us:

  • The dashboard summary is readable. The hub also sends the relay a summary of each task for the dashboard: the brief, progress, the result text (up to 12 KiB), steers, usage, branch and commit, and artifact names and hashes. The relay stores it. Treat the relay as able to read your prompts and results.
  • The end-to-end keys are not authenticated. They protect against a relay that only watches; a relay could in principle sit in the middle of the key exchange, though no code in it does. And an end with encryption on still accepts unencrypted messages rather than refusing them.
  • Metadata is always visible: agent and machine ids, labels and machine descriptions, timing and message sizes.
  • The dashboard can act on your work. The relay carries your own dashboard actions on tasks you started (answer, approve, reject, abort, steer, pause) to the hub. It cannot start new work on your machines.

What does not go through the relay: your repositories and data, which stay on the machines that work with them, apart from what agents put in their answers.

How the relay is tested and how reliable it is: relay restarts at random points passed 40 / 40 and 10 / 10 goals (one relay) and 6 / 6 SIGTERM, 3 / 3 SIGKILL and 3 / 3 deploy windows (hub on nodes); ownership leases linearizable over about 630,000 operations; canary 25 / 25

The relay is exercised on every rung of the testing ladder: SIGTERM and SIGKILL at random points, deploy windows of 30–90 s, the history audit, the linearizability checks of the ownership leases, and the production canary through the real relay. Relay-side defects those runs found and fixed:

  • The relay kept a context for every finished request on long-lived connections: heap growth of 4.6 MB in 33 minutes, none after the fix.
  • An idle hub re-sent its whole dashboard summary every 5 seconds (58 % of the idle hub's CPU), and the relay's idle CPU went to decoding it. It now sends only on change.
  • Status requests decoded each task's full event history: the relay reached 143.8 % CPU under polling (2026-07-24). A state-only read took 25 polls to 0 full decodes.
  • A relay sharing Redis re-read every task every 5 seconds while idle, for lack of a revision counter.
  • Dispatches and resumes sent while an agent reconnected failed as "target agent not connected"; they are now held briefly. With faster redialing, an in-flight task resumed 0.5 s after a restart instead of 7.5 s (p50).
  • A running agent whose hub had exited, or that had been quiet for over 5 minutes, disappeared from the dashboard; such rows now stay, marked when their publisher is gone.
  • Artifact integrity headers were dropped on the way out of the relay; artifact downloads now bypass the layer that dropped them.

Still open for the relay: multi-replica relays have no soak run yet and no production multi-host proof; the Kubernetes chaos runs on hub sessions are pending; rollback has not been run on the real VM.

nfltr vs. Claude Code over SSH or a VPN

You can already point one Claude Code session at other machines: give it SSH, or put everything on a VPN. That works, and for some jobs it is the right call. The difference is where the agent runs.

One remote-controlling session vs. an agent on each machine Left: one Claude Code session on the laptop runs commands on two machines over SSH, and their output flows back into its one context. Right: the hub on the laptop spawns an agent on each machine through the relay; each agent works in place and only its result comes back. Claude Code over SSH one session gets all output VMno agent DB VMno agent inbound SSH nfltr hub gets results only VMhas agent DB VMhas agent outbound
Left, commands and their output cross the network into one session. Right, each machine runs its own agent and only its answer returns.
One Claude Code session with SSH or a VPN, and nfltr.
Claude Code over SSH or a VPNnfltr
Where work happensOne session drives every machine remotelyAn agent on each machine, with that machine's tools and project context
DataCommand output flows back into one contextProcessed in place; only results return (the demos return counts, not rows)
Parallel workBackground commands, all feeding one sessionAgents on several machines at once, up to each machine's limit
A dropped connectionThe running command dies with itThe agent keeps going; its result arrives exactly once
Network exposureInbound SSH or a VPN, keys that reach the whole machineOutbound only; tools, monitors and repositories opt-in per machine
Watching for somethingThe session polls, and every poll costs tokensA monitor costs nothing until a line arrives
Setup and costNone if you already reach the box; one model sessionA join per machine and the relay; a model session per agent

SSH is the better choice for a quick one-off command on one machine you already reach: nothing to set up, no relay to depend on, and one model session instead of one per agent. A VPN solves reachability, not the agent model: the session still runs in one place and pulls everything back to it.

What we learned building it

The hub didn't know other machines existed. Given a plain goal about a dataset on a VM, early takes had Claude try gcloud in its own shell. The fix was the instructions the MCP server sends when a session connects: facts about the setup, no policy.

"The QA environment" wasn't covered. The instructions named machines, datasets and services as things that may be elsewhere, but not environments, test suites or logs, and the hub gave up without listing the nodes:

Before: 0 of 3

Goal: "Run the full regression suite now in the QA environment that tests our release branch"

After about 8 seconds, without listing the nodes:

I don't have access to a QA environment

After: 3 of 3

The fact we added to the server's instructions, paraphrased: this machine is one of several; an environment (QA, staging, production), test suite or logs missing here may be on another machine; what is missing here says nothing about the others.

Each run found the release environment and reported the flaky cache test: 34 of 35 passed, and reruns flip-flopped.

"Run it again there once it's back" needs a spawn that can wait. In the resilience demo the VM's node is killed mid-job; the hub sees node_lost and issues the job again. At first a spawn for an absent machine failed after 30 seconds, and the hub announced a "scheduled check" it could not make. Now the spawn's timeout_ms bounds the wait, and it runs when the machine rejoins.

A noisy watcher wakes the session on every line. Each monitor event starts a turn, which is a model call. A command that prints a timestamped "ok" every few seconds wakes it every time, and deduplication cannot match lines that differ; one that prints only changes or failures keeps it quiet. The tool description says so; nfltr does not rewrite lines.

Agent reports were cut to their first line. The result of any agent that changed a repository carried only the first line of its answer:

Before

What the hub received from an agent that pushed a fix branch:

Pushed successfully. Summary:

It asked twice more. That run cost $2.00; the other two runs of the goal cost $0.89 and $1.03.

After

The completions of agents that changed a repository carried their whole answer, 21 and 26 lines in the next runs, and the rerun of that goal needed no follow-up question.

Hubs sharing a store took each other's agents. In a test with no model spend, a second hub on the same machine acknowledged the first hub's parked agents, so their processes exited: every process sharing the one task store acted as a replica of the same hub. Each hub now has its own store and ignores agents not in it.

How we test it

Most of what goes wrong in a distributed system has nothing to do with the model: a relay restarts mid-dispatch, a machine dies, a laptop sleeps. We test all of that without a model, and spend on real Claude only to see whether it chooses well with what nfltr gives it.

The testing ladder From bottom to top: unit and contract tests, with a guard that fails any test that runs a real agent CLI; end-to-end hub tests with a scripted fake Claude; fault injection and soaks; oracles (a history audit and linearizability checks); Kubernetes chaos experiments; a production canary every 15 minutes; and at the top, capped real-Claude campaigns judged by oracles. Only the top rung spends on models. Real-Claude campaigns capped spend, judged by oracles Production canary every 15 minutes, hosted relay, $0 Kubernetes chaos relay, Redis and worker faults, $0 Oracles over what happened history audit, linearizability checks Faults and soaks restarts, kills, sleep, deploy windows, hours of load End-to-end hub tests, scripted fake Claude real relay, nodes and CLI; zero model calls Unit and contract tests a guard fails any test that runs a real agent CLI
Each rung runs what the one below cannot: the fake tests the machinery, faults test recovery, oracles check the recorded history, the canary tests production, and real Claude tests the model's choices. Only the top rung spends on models.

The rules we test by

  • Judge by outcome, not prose. A run passes when the world shows it, not when the model says so.
  • Fix the root cause once, where it lives, with a regression test that fails before the fix.
  • Never widen a timeout to hide a race. A spawn right after a relay restart failed because the node had not re-registered; the fix waits for its registration event within the existing bound.
  • Omission stays omission. A model, effort or budget nobody set is never filled in; a pre-commit check rejects concrete defaults.
  • Tests never run a real agent CLI. A guard puts failing stand-ins for claude and the other agent CLIs first on each test's path and fails the run if one is called.

A scripted fake Claude

nfltr reaches Claude only through the claude CLI, so a deterministic stand-in for that CLI replaces every model call. It follows a scenario file, answers as an agent or as the hub session, and interprets no prose; anything it does not recognise exits with an error, so drift from the real CLI fails loudly. It proves machinery, not model quality.

End-to-end hub tests with the fake: spawn, steer, stop, continue, reattach, isolation, monitors, upgrade

These start a real relay, real nodes and the real CLI, and drive the hub over MCP the way Claude Code does: spawns constrained by machine and by repository, a steer into a running agent, stop, a continuation that resumes the same Claude session, a disconnect and a reattach that receives the missed completion exactly once and nothing else, and a kill -9 of the hub whose successor still gets it. Two hubs must not see each other's agents. Monitor lines printed while no hub runs arrive exactly once after reattach, and a runaway yes is held to its declared rate (5 lines delivered, 4.5 million counted as dropped in two seconds). An upgrade test checks that an agent parked in the old shared store is still delivered once. A host-local ladder (readiness, spawn, edit and verify, steer, continue, stop, reattach, on-demand nodes) passes all 8 rungs in 144–172 s on a laptop.

Tests that run a freshly written fake executable used to time out on macOS, which checks a new executable on its first run (0.2–0.9 s idle, seconds under load). The test helper now runs each stub once when it writes it; the product's own bounds stay as they are.

Faults and soaks

With the fake, a soak driver injects one fault per goal at a random point and runs goals back to back for half an hour or more. A goal passes when every agent ends exactly once within 2 minutes of the fault ending (completed, or node_lost where the fault took its node), monitor events arrive exactly once with no gap, neither hub sees the other's agents, the history audit is clean, and nothing is left running. Rates, not gates, from one shared laptop.

The hub on nodes, nine fault classes (2026-09-28)

One relay, three joined nodes launching agents on demand (one runs a monitor printing a tick every second), and two hubs on the same relay. Before is the first runs of the harness; after is the final 30-minute soak.

Fault matrix before and after the fixes, fake Claude, no model spend.
FaultBeforeAfter (goals passed)
Relay SIGTERM or SIGKILL, down 1–5 sfail at spawn the agent exited and the turn hungpass SIGTERM 6 / 6, SIGKILL 3 / 3
Relay deploy window, 30–90 sfail at spawn the turn failed or gave up too earlypass 3 / 3
Node kill -9 mid-agent, then rejoinfail the loss arrived after 5 minutes, or neverpass 4 / 4, node_lost 5–7 s after the kill
Node restart under a running monitor, hub awayfail monitor lines lost; the agent's outcome never arrivedpass 2 / 2, every line delivered
Monitor's node killed for goodfail the monitor stayed running until the node rejoinedpass 3 / 3, lost 30 s after the kill
Hub kill -9, then reattachfail an agent never exited and kept taking workpass 7 / 7
Node frozen 60–200 s (a closed laptop lid)passpass 6 / 6
Hub frozen 60–200 s0 / 6 a stray agent; an empty success in place of a real resultpass 2 / 2
Two hubs on one relaypass no cross-hub visibilitypass 36 / 36 goals
The six 30-minute soaks, each on the fixes so far. Bar: share of goals passed.
SoakShare passedGoals passed
Run 117 / 24
Run 228 / 30
Run 328 / 29
Run 422 / 25
Run 522 / 24
Run 6, final36 / 36

The final soak ran 36 of 36 goals and 144 agents (134 completed, 10 node_lost on the node a fault took) with no duplicate completion or monitor event and nothing left behind. Idle afterwards, the relay used 0.23 % CPU and the nodes 0.02–0.03 %. The runs found 13 defects; each was fixed at its owner with a regression test.

Recovery in the final soak (p50): relay restart to all nodes back about 1 s, node restart to listed 0.35–0.46 s, hub reattach 114 ms, after a freeze 15 ms (hub) and 356 ms (node)
FaultRecovery p50 / p95Measured as
Relay SIGTERM957 / 1,254 msrelay start to all nodes listed
Relay SIGKILL1,244 / 1,246 mssame
Deploy window 30–90 s958 / 2,218 mssame
Node kill -9 and rejoin354 / 364 msrestart to listed
Node restart under a monitor, hub away462 / 483 mssame, including the hub's reattach
Hub kill -9 and reattach114 / 151 msreattach to first answer
Node frozen 60–200 s356 / 369 msresume to listed
Hub frozen 60–200 s15 / 19 msresume to answer

Spawn to completion across all goals was 14.1 s p50 and 173.6 s p95; the fake agent's turn is 8 s, and the time includes each fault's own duration.

The 13 defects, in short
  • A node restart lost the monitor lines it held; they are now kept on disk and delivered after the restart.
  • A node_lost failure could be dropped in transit; it is now kept until the hub acknowledges it.
  • An agent a restarted node had reaped never reported; the hub now notices its agent is gone and ends the turn node_lost.
  • Two cases where a result recorded while the hub was away was never acknowledged, so the agent kept its slot and took new work.
  • A killed node's worktrees stayed behind.
  • Three cases where a status check made up an empty success and dropped the real result (branch, commit, text) the agent still held.
  • A node agent that dropped a not-yet-accepted task during a relay restart exited as idle.
  • A turn refused before it ran ended failed instead of being placed again.
  • A turn waiting for its node after a relay restart gave up just before the node came back: the wait equalled the node's reconnect bound.
  • A monitor on a node that is gone stayed running; it now ends lost, once.

One relay, restarts and long soaks (2026-09-27)

An earlier round, with fleet workers, a node and one hub, restarted the relay at random points and ran a two-hour soak. After its fixes, an in-flight task resumed 0.5 s (p50) after the relay was ready, down from 7.5 s.

Single-relay results and recovery chart: relay restarts 40 / 40 goals (SIGTERM) and 10 / 10 (SIGKILL), node kills 20 / 20, a two-hour soak 303 / 304
After a relay restart (SIGTERM runs), seconds from the relay being ready. Top bar: before the follow-up fixes; bottom bar: after.
MeasureBefore and afterBefore → after
In-flight task resumed, p507.5 s → 0.5 s
In-flight task resumed, p9520.4 s → 1.1 s
All workers reconnected, p508.7 s → 0.7 s
All workers reconnected, p9530.7 s → 1.3 s
One relay, fake Claude, no model spend, 2026-09-27.
FaultWhat is checkedResult
Each hub operation 50 times: spawn (by machine, by repository, with a clone), steer, continue, stop, budget stop, disconnect and reattachpasses, no orphan processes or worktreespass 50 / 50 each
Relay SIGTERM at random points, 57 restartsevery agent exactly once, commits on origin, audit cleanpass 40 / 40 goals, 120 agents, 0 duplicate or lost
Relay SIGKILL at random pointssamepass 10 / 10 goals, 17 restarts (first round 9 / 10)
Spawn while the relay is down (deploy window)the spawn waits and runspass 5 / 5 goals (failed before the fix)
Node kill -9 mid-agent, 20 timesthe hub gets node_lost; a new spawn on the rejoined node completespass 20 / 20, p50 206 ms to node_lost
Hub kill -9, then reattachmissed completion delivered oncepass in 30 s (was 111–120 s)
Two-hour soak, 304 goals, 912 agentsevery goal within its timeout303 / 304 the miss spanned an 11.6-minute laptop sleep; its agents still completed exactly once
Idle relay and hub for 15 minutesnear dormantfixed hub 0 GC/min, ~0.3 % CPU (it had re-read its store every 5 s)

The first restart round found a dispatch abandoned when its stream dropped during a relay restart, a resume refused while the worker reconnected, a node-launched worker that refused its task stranding the agent for about 15 minutes, a dead Claude session's socket mistaken for a live one, and an idle hub that re-read its store every 5 seconds. A second round fixed recovery speed, leftover worktrees, a lost acknowledgement that kept a worker alive, and memory the relay kept for every finished request.

Oracles over what happened

A history audit. In the spirit of a Jepsen checker, a tool reads a run's event history and reports violations, never changing anything: a finished task going live again, a result accepted twice, a gap in a task's event cursor, a completion from the wrong worker, rejected output on the main branch, a running task silent too long.

Linearizability checks. Each task has one owner at a time, fenced by a lease. With porcupine, concurrent replicas run against the in-memory, SQLite and Redis stores under lease expiry, crashes and dropped connections, checked against a model of the contract. After the fixes, 500 seeds (about 630,000 operations) were linearizable for all four store targets (2026-09-25).

What the linearizability checks found: every store flagged at first, four root causes

A clock read twice in one decision, a clock read before taking the write lock, lease reads not watched to commit, and a retried renewal that extended a lease twice. Each was fixed with a guard test. The 500-seed run included about 1,200 operations with unknown outcomes. A second model over event appends and per-task sequences found no violation.

Kubernetes chaos

A three-node kind cluster runs three relays sharing Redis, workers and a hub. Each experiment is one fault (relay, Redis and worker kills, a rollout restart, partitions, a hub kill; with Chaos Mesh, latency, partitions and clock skew) plus a steady state that must hold before and after, judged from the journal rather than the chaos tool's own status.

pending The experiments now run hub sessions; results will be added here after the first run.

A production canary every 15 minutes

A hub on a dedicated VM talks to the hosted relay with real TLS, auth and deploys. Each run starts a fresh hub, spawns a fake agent on each of two canary nodes (one pushes a branch) and a monitor, and requires everything exactly once, the branch on the origin, no other hub's agents in view, and nothing left running. The first 25 runs, over 70 minutes with a relay deploy between two of them, all passed.

The first 25 canary runs (50 agents), laptop hub to the hosted relay, 2026-09-28. Bar: p50; tick: p95; line: max, on a 0–12 s scale. The fake agent's own work is 3 s of spawn to completion.
StepLatencyp50 / p95 / max
spawn_agent call1.5 / 2.2 / 2.4 s
Spawn to first progress3.3 / 5.7 / 6.6 s
Spawn to completion8.5 / 10.9 / 11.6 s
Monitor start to first line1.1 / 1.6 / 1.9 s

Real Claude, judged by outcome

For the model's choices we run capped campaigns on real machines. Goals are plain English and name no tool; each run is judged by an oracle: statistics recomputed on the VM, branches cloned fresh and tested, a service answering its health check. The latest campaign (2026-09-28): 33 runs on the hosted relay, $14.01 in total, 29 passed. Three failures were one goal before the instruction fix above; the fourth was a wrong count the hub passed on unchecked. No run stalled, exceeded its cap, left anything running, or used a machine the goal ruled out.

The latest real-Claude campaign: 33 runs on the hosted relay, $14.01 in total, each judged by an oracle.
GoalPass ratePassed
Statistic on a file that exists only on a VM4 / 5
Two backlog items in parallel on two machines3 / 3
Same, laptop only, after the report fix1 / 1
Broken service on a VM: investigate and fix3 / 3
Machine killed mid-job, re-run when it is back3 / 3
Same goal, no fault injected (the test driver missed its trigger)1 / 1
QA triage, staging: a stopped database3 / 3
QA triage, main: regression, bisect, fix branch with a test3 / 3
No machine named: find it from the node descriptions5 / 5
“Run the release suite”, before the instruction fix0 / 3
Same goal, after the fix3 / 3
Per-run notes: the wrong count, and two passing runs that still found problems

The wrong count: the agent compared numbers as strings, and the hub relayed it (model behaviour; the means in the same answer were right). The passing runs: the truncated report (a bug, fixed) and a spawn refused on the laptop for a missing ssh alias, where the error named the cause and the hub moved the work to a VM. We also read every hub transcript and state file afterwards.

We also compare the hub with a single agent on the same tasks, with the same snapshot, goal, oracle and caps. In an earlier comparison on nine scenarios built for tasks where coordinating agents should help, the hub passed 7 and one agent alone passed 5.

What is still open

  • A monitor whose machine is away past the reconnect bound (a laptop asleep for a few minutes) ends lost and is not revived when the machine wakes; the hub can start it again. Lines a killed node never read from the command are gone.
  • After a kill -9 of the hub, the missed completion arrives within 30 s rather than at once.
  • The idle hub is not fully dormant yet (0.36 % CPU, about 5 GC/min in the single-relay soak, possibly partly the sampler).
  • The history audit does not see monitor events; the soak harness checks those itself.
  • A result carrying artifacts that is recorded while the hub was away is acknowledged, but artifact content only the live transfer carries is not delivered on that path.

Demos

Three of the homepage demos, recorded on real machines:

  • Overnight QA triage. A plain Claude Code session starts a monitor on each of three QA environments and goes idle. When the nightly runs finish, the monitors wake it with nobody typing, and it triages each failure where it happened: a rounding regression pinned to its commit and fixed on a branch with a test, a flaky test rerun, a stopped Postgres started. No logs leave the environments. 677 seconds; session $1.63, seven agents $1.19.
  • An incident across three machines. A monitor wakes the session when a bad deploy starts failing. It investigates on the app VM, runs an aggregate-only query on the private database VM, fixes the code with regression tests, deploys and shows it healthy.

Loaded only when you press play. Pauses over two seconds are shortened; nothing else is edited. Models and budgets on screen were picked for the recording.

Try it

On each machine:

curl -fsSL https://nfltr.xyz/install.sh | sh
nfltr config add-api-key <key>
nfltr node join --max-agents 2 --allow-all-tools

Then give Claude a goal:

nfltr orch "<goal>"

Or use your own Claude Code session:

claude mcp add nfltr -- nfltr mcp --toolset hub
nfltr orch hub install-hook
claude

More: Use nfltr from Claude Code, Join machines as nodes, and the hub tools reference.

What's next

Planned, not shipped:

  • An always-on hub. Run the hub on a machine that stays up, so overnight work keeps being handled while your laptop sleeps, and reattach to it from your laptop in the morning.
  • Triggers from external systems. Alertmanager, Jenkins or GitHub webhooks feeding the same inbox as completions and monitor events, with the same limits, their text marked as untrusted data.
  • Schedules for overnight and recurring work. Today a monitor whose command is a loop does this; a scheduler on the node that emits an event on each tick is the next step.
  • A node-scoped key, before nodes run on hosts you do not fully trust, such as PR runners or shared CI.
  • Less for the relay to see. An option to keep prompts and results off the relay (a status-only dashboard), end-to-end keys authenticated and tied to your account, refusing unencrypted messages when encryption is on, and a switch to turn the dashboard's command channel off.

← All posts