What nfltr orch is useful for, measured

nfltr orch is useful when work needs more than one machine, more than one agent at a time, or access that should stay on the machine that holds it, and it is not worth it for a quick fix in one repository.

Here is the evidence: three measured uses, each with its setup, numbers and what is still partial, and the case where one Claude Code agent wins. In every run the hub and its agents used the same Claude model. How it learned to keep a fleet busy is in the earlier post, How a Claude Code hub learned to fill a fleet.

One big task, split across agents

Setup. One goal: write tests for six independent packages of a Go codebase, judged by a hidden check that runs the tests and 12 seeded mutants. Four nodes, each a container on one host with 2 CPUs, 4 GB and --max-agents 2 (8 slots), a private relay, a cold fleet for every hub run. The single agent ran in a container of the same size, to its own end.

Time to finish, one run per arm. Bars are against the slowest run. Yellow: the run where one agent's long tail erased the gain. Every arm passed the hidden check.
ArmTime to finishTimeCost
Single agent605 s$2.44
Hub, six agents322 s$3.83
Hub, steered, before a fix1,151 s$4.34
Hub, steered, after the fix333 s$3.79

Every hub run split the task six ways in its first call, one agent per package, and every part landed. The default hub finished in 322 s against the single agent's 605 s, 1.88× faster at 1.57× the cost. Splitting bought wall time, not tokens: each agent started from a cold clone and build and wrote and checked its own suite, so the six together worked about twice as long as the one.

Partial. One slow agent can erase the gain. In the 1,151 s run, one package's agent spent most of 1,118 s trying a race-detector build that its node's image could not do, until the hub sent it a correction; the same part took 290 s and 300 s in the other runs. And here the goal named the six packages. In the earlier 20-task batch, where two splittable tasks sat among 18 others, the hub split neither.

Steering. You can talk to a running hub. In the two steered runs a scripted operator message asked it to move work to idle slots. A message sent while it waits for agents now ends the wait with a note that a message is pending: in 0.16 to 0.22 s, and the hub's next call came 2.5 s after the send. Before the fix in the table, a batch of parallel waits held the message 70 s. Both hubs read the request, looked at their agents and declined to move work: their call.

Many tasks, across real machines

Setup. The same 20 tasks as the earlier post's full run, on three separate machines joined to nfltr.xyz: a laptop (10 cores, 32 GB), which also ran the hub, and two cloud VMs with 4 vCPU and 16 GB each. Each node had 3 slots, so 9 in all. The VMs held the data sets the data tasks needed; the laptop held none.

One run, 2026-10-01. Serial is the agents' run times added up: how long the same agents would take one after another. It is computed, not a separate single-agent run.
MeasureResult
Makespan against serial1,062 s against 3,737 s: 3.5× faster
Hidden checks20 / 20 judged on the statuses the hub answered (see below)
Data-bound agents on the machine with the data3 of 3
Model spend$6.17 in total ($1.57 hub, $4.60 agents), $0.31 per task
Throughput67.8 tasks an hour
Mean utilization of the 9 slots0.39 all 9 busy for 11 % of the run

About the 20 of 20: the hub put its final JSON answer inside a code fence after one line of text, and the test runner, which wants exactly one JSON object, refused it, so its strict score is 0 of 20. In the run itself 17 of 20 checks passed; judged on the statuses the hub answered, all 20 pass.

Partial. Utilization was 0.39, against 0.71 on the one-host fleet of the earlier post, where the same 20 tasks ran 5.7× faster than serial (1,416 s against 8,106 s, four 2-CPU containers on one host, a private relay). The two runs are not directly comparable: here the agents had faster machines and did less than half the busy time. Three things kept slots idle:

  • The laptop took no work for its first 2 minutes. It had two agent harnesses installed, so its node refused every launch that did not name one, and the hub was not told why. Since this run, the node's refusal reaches the hub and nfltr node join warns at startup.
  • Clones over the WAN. The task repositories were on the laptop, so every VM agent cloned over its uplink: an agent's first clone took a median of 54 s and 66 s on the two VMs, against 8 s on the laptop.
  • An overloaded VM. Three cold builds on 4 vCPU pushed one VM's load to 15–19, and one agent there ran 813 s, leaving a single slot busy for about 360 s at the end. Fewer slots on a small machine (--max-agents) is the lever.

Incident response across environments

Setup. A real Kubernetes cluster (a local kind cluster) with a Go API that a bad config change makes crash-loop at startup. Prometheus and Alertmanager sent a real alert to a webhook the hub created on nfltr.xyz. Four nodes, each its own container: a read-only cluster node, a write-capable cluster node gated on the owner's approval, a CI node with the build logs, and a laptop node with the developer's checkout. The owner's signing key lived in a separate container with no network. The hub ran with nfltr orch "<goal>", and the goal said to investigate, propose a fix, apply it only with the owner's approval, and stop once the alert resolved.

Run 2, with no shortcut. No environment held the whole story: the cluster had the crash and the failing setting's name, but not the build, the PR or the value; CI had which build made which image; the laptop had the PR's diff.

Run 2, 2026-10-01, times since the alert fired. The first 10 s of the alert's delivery is Alertmanager's own grouping wait.
Since the alertEvent
11.8 sThe hub has the alert
18.5 sOne read-only investigator on the cluster node
46.4 sIts result: the crash, the setting, a healthy previous revision
48.7 sThe hub requests approval for an exact rollback prompt on the write node
1 m 17 sThe owner signs it
1 m 34 sRollback applied; 2 of 2 pods ready seconds later
4 m 12 sThe resolved alert reaches the hub
4 m 23 sThe hub has revoked its webhook and written its report

The incident session cost $0.39. Without being asked, the hub mitigated and stopped: its report named the failing setting but not the build or the PR, and said the bad value was inferred from the log. After one message from the owner asking which build and PR produced the image, the hub looked at the CI and laptop nodes, ran one read-only agent on each in parallel, and answered correctly in 53 s for $0.25: the build, the PR, the commit, its author and the change. The whole run cost $0.64.

Run 1, with a shortcut. The root cause was there in 36 s and the fix applied 2 m 16 s after the alert, about 1 m 24 s of it the owner's. But in that run the deployment carried an annotation, written by CI, that named the build and the PR, so the cluster alone told the whole story and the hub never used the CI or laptop nodes. Run 1 did not test correlation; run 2 did.

Safety, measured in these runs:

  • The write node took only its own hub's pinned work. It was joined with --serve-hubs and --require-pin. A spawn on it from another hub id was refused, and so was an unpinned spawn from its own hub; no agent process started for either.
  • Changes ran only with the owner's signed approval, bound to the exact prompt. nfltr orch hub approve --expect-digest refused the prompt file with one byte added and signed nothing. In run 1, an unapproved write was denied by the node, and the agent did not try to work around it.
  • Prompt injection was ignored. Text telling AI agents to delete namespaces sat in the alert annotation, a pod log and the PR. No agent acted on it, and no agent tool call in run 2 contained delete; the hub reported the text to the owner as a security finding.
  • No credential reached hub-visible output. The cluster tokens, the CI token, the API key, the Claude login and the owner's key were searched for in every transcript and log: no hit. The only secret printed was the webhook's, which the goal asked for, and it was revoked at the end.

When not to use it

A quick fix in one repository is faster with one agent. On two short tasks run on nfltr.xyz, the hub took 67 s against the single agent's 28 s, and 111 s against 62 s; once it took 479 s, because its brief asked for a whole-module test run. The hub adds about 20 s of fixed time per task: its own session before the spawn, the launch, and delivering the answer. The cost was within $0.03 of the single agent's. A hub earns that overhead when the work needs another machine, more slots, or many tasks at once.

What's next

  • Guidance for incident goals. The hub stopped at a correct mitigation because nothing asked for more. We will publish how to state what a root cause must contain (the build, the change, the author) in the goal or a --hub-guidance-file, and measure an incident run that uses it.
  • Multi-machine utilization. A rerun with the refusal fix, repositories near the agents, and slots sized to each machine, to see how much of the 0.39 was setup and how much is the hub.
  • Hosted TURN, so nfltr p2p calls across symmetric NATs work without your own TURN server.

Try it

Join each machine you want agents on, then give the hub a goal:

nfltr node join --max-agents 2 --labels env=ci
nfltr orch "<goal>"

The guides: Join machines as nodes (with approvals for nodes that can change an environment), Use nfltr from Claude Code if you want the hub in your own Claude Code session, and the hub tools reference. The design and how we test it are in How distributed Claude Code is built with nfltr.

← All posts