How a Claude Code hub learned to fill a fleet
Give a Claude Code hub twenty independent tasks and eight places to run agents. Does it keep all eight busy, send the work that needs a machine's data to that machine, and stay off the machine that is already loaded? An operating system does this for processes without thinking about it. For a hub it took eight rounds of measurement and fixing, and the most important fix was not about scheduling at all.
This post covers what we measured, what we found, what we changed, and what is still partial. Every number comes from a recorded run, with its conditions next to it.
What we measured
Each round used the same kind of fleet: four nodes joined with nfltr node join --max-agents 2, so 8 agent slots. Each node was a container on one host, limited to 2 CPUs and 4 GB. In the many-task rounds (7 and 8), one node was kept busy by other load, and each data set lived on exactly one node. The hub and every agent used the same Claude model, the relay was a private one (not nfltr.xyz), and each run had a spending cap. The tasks were real coding, test, docs and data tasks, each judged by a hidden check: a test suite, a build, or a recomputed number, never keywords in the answer.
From the hub's state and a 5-second utilization sample we computed:
- Makespan: from the session's start to the hub's final answer.
- Busy time: the agents' run times added up, in slot-seconds. Busy time is also the serial time: how long the same agents would take one after another. It is computed, not a separate single-agent run.
- Ideal: busy time divided by 8 slots. A makespan of 1.0× ideal means no slot was ever idle.
- Mean utilization: busy slots over all slots, averaged over the run.
- Idle while ready: slot-seconds idle while some task was waiting, either for a slot or for the hub to start it.
- Refill: the time from a slot freeing up to the next agent starting there.
The hub saw a quarter of what we told it
In earlier rounds the hub ran a task that touched 40 call sites in 11 independent packages through exactly one agent, every time. Asked afterwards, in the same session, why it used one agent, it gave four reasons: it could not see the work's size, it did not know how many slots were free until after its first spawn, several agents pushing to one branch would need rebasing, and several agents would cost several times as much.
Most of those were answered in the instructions the hub's MCP server sends when a session connects. The hub had not seen them. Claude Code (version 2.1.285 in these tests) cuts an MCP server's instructions and each tool's description at 2,048 characters and marks the cut … [truncated]. A parameter's description is passed whole. We measured it two ways:
- A probe server with 6,000-character texts and a numbered marker every ~480 characters. The model listed markers up to offset 1,927 in the instructions and the tool description, and all 13 markers in the parameter description.
- A hub session asked to quote the last lines it had of its instructions returned the text up to exactly character 2,048.
The hub's instructions were 8,677 characters, so every hub saw about 24 % of them. The parts it never saw included that agents run in parallel, how several pushes land on one branch, that several agents' wall time is the longest one's rather than the sum, and the framing of agent output as untrusted data.
The fix was to fit the limit, not to add rules. The session instructions are now a core of at most 2,048 characters, most important first. The facts moved to where they arrive whole: each tool's own description, kept under the limit, and its parameter descriptions. The number of free agent slots is in the instructions from the session's start. A test fails the build if any of these texts grows past the limit, and a fact counts only if it sits where the hub can actually read it.
Before: 0 of 7
No hub run used more than one agent on the tasks built to be split. None looked at the free slots before spawning.
After: 2 of 6
Both runs of one task used four agents, one per package, and quoted the facts they weighed. The other two tasks stayed at one agent.
That was not yet a win. On those small tasks one agent was faster and cheaper: the four-agent runs took 95 s and 126 s and cost $1.01 and $0.97, against 80 s and $0.41 for a single agent. Each part was under a minute of model work, and a fresh agent's start (a new Claude session, its own clone, a cold build of its own checkout) cost about as much as the part itself. The hub's one-agent choices matched those economics. Splitting pays when the parts are large, or when there are more tasks than slots, so that is what we measured next.
What kept slots idle
With the facts reaching the hub, the remaining idle time came from mechanics and wording. Each item below was found in a measured run and fixed with a regression test.
A freed slot was invisible for 30 seconds. When a node's agent finished, its connection closed before its process exited, and nothing told the hub when the node reaped the process and freed the slot. A waiting spawn found out only from a 30-second backstop. In a probe, the next agent started 39.7 s after a slot freed. Now the agent marks its slot free before it stops, and in the same probe the next agent started 0.3 s after.
One call, many agents. spawn_agents starts up to 64 agents in one call. Each entry is validated and placed on its own, and an entry beyond the free slots waits for one instead of failing. The hub can queue the whole batch at once, and a freed slot goes to the next waiting agent without another model turn.
Wording made hubs place work by hand. In round 6 the hub pinned every code task to a node itself, two or three per node (that round's utilization, 0.32, was the lowest). Nothing required a pin. Three phrases suggested one: the node tool said a node starts an agent "for a spawn_agent it can host (constraints.machine = its machine)", which reads as "only for a spawn naming it"; an unpinned spawn "takes the first free slot", which reads as packing onto one machine; and the task brief said agents run "each in its own checkout on its own machine". The facts now say what placement does. Unpinned work goes to a node with a free slot, an idle node first, on any machine. In the next two runs the hub pinned 0 of 18 code tasks each time and pinned every data reader to its data node. That is one run per variant, and hubs vary from run to run.
Spare CPU before free slots. Placement used to rank nodes by running agents and free slots before spare CPU, so in round 7 the loaded node still got 3 and 5 agents. When nodes report their CPU, unpinned work now goes to the node with the most spare CPU first. Placement still never moves a hard pin and never picks a machine for a pinned agent.
A constraint in the wrong place. Hubs sometimes put a field such as mutates_workspace next to the prompt rather than inside constraints. A multi-spawn refused the entry without saying where the field belongs; a single spawn ignored the field without saying so. Both now refuse it with an error that says where the field belongs, and nothing is moved for the hub.
Your own scheduling policy. If you have one, state it in a file: nfltr orch --hub-guidance-file <path>. The hub gets it verbatim in its session prompt and its spawn tool, and without the flag nothing is added. In the one run with guidance, utilization was 0.91 against 0.88 without guidance, but 1 of 25 briefs reflected the guidance and the hub never cited it, so we cannot credit the difference to it.
Utilization, round by round
Rounds 2 to 6 ran 8 tasks on 8 slots. With no queue, utilization there is capped by the longest task, and the number to watch is idle-while-ready. Round 7 was the first with 20 tasks on 8 slots, so there was a queue. Both of its runs were stopped early, the first by the test runner's batch limit and the second by hand when the round's budget ran out, and their figures cover only the time up to the stop. Round 8 ran the same 20 tasks to the end.
| Run and makespan | Mean utilization | Util. | Vs ideal |
|---|---|---|---|
| Round 2 8 tasks, 600 s | 0.50 | 2.0× | |
| Round 3 8 tasks, 786 s | 0.44 | 2.3× | |
| Round 4 8 tasks, 803 s | 0.47 | 2.1× | |
| Round 5 8 tasks, 885 s | 0.46 | 2.2× | |
| Round 6 8 tasks, 1,339 s, hand-pinned | 0.32 | 3.1× | |
| Round 7, run 1 20 tasks, stopped at 651 s | 0.88 | 1.13× | |
| Round 7, run 2 20 tasks, stopped at 623 s | 0.91 | 1.07× | |
| Round 8 20 tasks, 1,416 s, to the end | 0.71 | 1.40× |
Round 8's mean is lower than round 7's because it includes the start and the tail, which round 7's stopped runs never reached. While work was queued, round 8 had all 8 slots busy from 120 s to 960 s.
The full run
Round 8 ran the 20 tasks to the hub's own final answer on the fleet above, with no guidance file. The tasks: 15 short ones (bug fixes, a small feature, a flaky test, a rename, a package move, docs, tests, a config migration, a verification, an investigation, and three that were already done, ambiguous or impossible), two longer test-writing tasks that could be split, one coupled change across a type and its users, and two data tasks whose data existed on one node each.
| Measure | Result |
|---|---|
| Makespan against serial | 1,416 s against 8,106 s: 5.7× faster than the same agents one after another |
| Makespan against ideal | 1.40× (ideal 1,013 s) |
| All 8 slots busy | 60 % of the run; mean utilization 0.71 |
| Refill, freed slot to next agent | median 1 s, p90 3 s, max 46 s (17 refills); up to 12 agents waited for a slot at once |
| Data-bound agents on the node with the data | 4 of 4 |
| The loaded node | 5 of 23 agents, each started while all 6 slots elsewhere were busy |
| Hidden checks | 20 / 20 judged on the status the hub answered (see below) |
| Answered status as expected | 19 / 20 one impossible task was answered not_completed, not one of the statuses the test accepts |
| Model spend | $14.67 in total ($1.66 hub, $13.01 agents), $0.73 per task |
| Throughput | 50.9 tasks an hour |
About the 20 of 20: the run's cap was $14.50, and the last turns reported their cost after the hub's final answer, which brought the total to $14.67. The test runner then marked every task over budget, so its strict score is 0 of 20; in the run itself, 17 of 20 checks passed and the other three failed only on that over-budget mark. Judged on the status the hub answered, all 20 pass. The cap check lagging turn cost is a known gap.
What is still partial
- No splitting yet. The hub did not split either of the two splittable tasks; 0 splits in the run. One big task still goes to one agent.
- The tail. From 1,140 s one slot was busy: a verification task ran 820 s on the loaded node. That tail and a 70 s start are why the makespan is 1.4× the ideal. Placement is work-conserving: when every other slot is busy, a free slot on a loaded node is used. You can lower that node's
--max-agents, or the hub can soft-pin work elsewhere. - The start. The first spawn came at +70 s: the hub wrote 20 briefs, and its first call used the tool's short name without the server prefix, which Claude Code answered with "No such tool available".
- Pinned work waited for its machine. The two data readers, pinned to their nodes, waited 550 s and 774 s, because agents that could run anywhere took the slots freed on those nodes first. Since the run, a slot freed on a machine goes to the pinned agent waiting for it. That change has tests but has not been measured live yet. In this run's timeline the readers would have started at about +84 s and +264 s; the makespan would not have changed, because the tail set it.
- One run. These are the numbers of one full run on one host's containers, not an average and not machines on separate networks.
Turns that end, and cost next to one agent
Alongside the fleet work, two rounds compared the hub with a single plain Claude Code agent on the same tasks, with the same snapshot, checks and caps, on nfltr.xyz.
- Cost reached parity. Round 6, seven tasks: hub $1.64 as reported (about $1.87 including one killed session's unreported usage) against $1.84 for the single agent, both 7 of 7 on the checks. Round 7, two tasks run twice: the hub's spend within $0.03 of the single agent's, tokens -22 % to +11 %.
- A turn now ends at its answer. In round 6 an agent answered while its own background commands still ran, the turn stayed open 584 s and was then cancelled at its time limit, after a correct push. Now the answer ends the turn: in round 7 the same situation completed 6 s after the answer, carried the pushed commit, and told the hub which background tasks had been left. All 4 hub turns in round 7 completed on their own.
- Polling mostly gone. Agents' polling loops took 1,712 s in round 5 and 2 s of tool time in round 6. In round 7 one agent still waited with a 300 s shell loop. How an agent waits is the model's choice, and the hub now sees which long commands an agent ran and how long they took.
- Still slower on short tasks. With tightly scoped briefs the hub took 39 to 49 s more than one agent on these short single-repository tasks: its own session, a launch and a clone. A hub is worth it when the work needs another machine, more slots, or many tasks at once, not for one small fix.
Also new for hubs
- Triggers. A hub can create a webhook URL on nfltr.xyz with its own secret. A CI job or an alerting system posts JSON to it, and each event reaches the hub exactly once, as untrusted data, under rate and size limits the hub sets.
- Schedules. A cron expression or a single time wakes the hub. A schedule fires while the hub is running; fires missed while it was away arrive as one event with a count.
inspect_node. A node's live CPU, memory, free disk and GPUs. On nodes joined with--allow-inspect-files, also file names, sizes and times, never contents.- Soft and hard pins.
constraints.pin:hardruns only on the named machine and waits for a free slot there;softprefers it and otherwise takes a free slot elsewhere. - Several agents in one call with
spawn_agents, as above. - Owner-approved changes for incidents. A node can refuse risky commands (for example
kubectl apply) unless the turn carries your signed approval for exactly that prompt on that node:nfltr orch hub approve. See Approvals. - Answers from a paired device (beta, off by default). Pair a phone or browser with
nfltr orch hub pair, run the hub with--remote-messages=answers, and answer its questions and approvals away from the terminal. They travel end to end encrypted; the relay forwards ciphertext only.
All of these are in the hub tools reference and the nodes guide.
nfltr.xyz itself
- Served directly. nfltr.xyz and every
*.nfltr.xyzname (shares, tunnels) are served by our own server, with no CDN proxy in the path, under a wildcard certificate issued with the DNS-01 challenge. The zone is signed with DNSSEC. - 1 GiB streaming uploads. Request bodies up to 1 GiB stream through to your
nfltr httpornfltr shareendpoint as they arrive; a body over the limit is refused with a 413 that names it. Before, a proxy in front refused uploads over about 100 MB and held the body until it had most of it. - Fair upload slots. A client already streaming a large upload does not take a second slot while another client waits.
- Alerting. Uptime checks from three regions page us when two of them fail for five minutes, and certificates alert well before expiry or on the first failed renewal. In a drill, a test alert opened an incident 3 min 23 s after it was created, and its email arrived.
- An SLO. A proposed 99.5 % monthly availability objective for nfltr.xyz, measured from those checks. It covers the health endpoint only, not tunnel traffic.
- Rollback and restore, drilled. One command rolls back a release and the same command rolls forward; in the drill the public site returned errors for 3.0 s rolling back and 2.4 s rolling forward. A backup was restored to a working relay in a drill. Backups are copied off the server daily and before every deploy, to storage the server can write to but cannot list or delete.
Calls with nfltr p2p
- Calls work end to end.
nfltr p2p callandnfltr p2p recvopen a browser call between two machines: audio and video both ways, chat and screen share, verified in headless browsers through a local relay and through nfltr.xyz. In those runs both peers were on one host, so the media took the direct path. - Across NATs. In a lab with real NAT routers, peers behind ordinary (cone) NATs connected directly. Where either side is behind a symmetric NAT, a call needs a TURN server for its media; with one configured, those calls connected in the lab. Hosted TURN on nfltr.xyz is not enabled yet. You can use your own TURN server with
nfltr p2p call --turn <url>, credentials taken from the environment. - Fixed: repeated frames were dropped. The relay discarded any frame identical to one it had already seen from the same client. Real data repeats: two file chunks of zeros, two message headers of the same size. Calls that fell back to the relay lost their signaling, and a file with repeated blocks stopped part way. The relay now delivers every frame; HTTP responses with repeated chunks were exposed to the same drop, and the one fix covers them.
What's next
- Large parts, measured. A batch where some tasks are big enough that splitting them pays, run on the hub and on a single agent, to see whether the hub splits when it should. The framework will keep providing facts, such as each agent's measured start cost, and the hub keeps the decision.
- The kept slot, live. A run that measures how long pinned work waits now that a freed slot goes to it.
- A cap that cannot overshoot. Spending caps that account for turns still reporting their cost.
- Hosted TURN, so calls across symmetric NATs work without your own server.
To try the hub: Use nfltr from Claude Code and Join machines as nodes. The earlier post, How distributed Claude Code is built with nfltr, covers the design and how we test it.