Claude Code agent teams and nfltr orch, measured

Claude Code has an experimental feature, agent teams: one interactive session, the lead, starts teammates that share a task list and message each other. nfltr orch does a related job another way: a hub session starts agents on machines joined as nodes and collects their results. We ran the same batch of eight coding tasks both ways, on the same machines with the same Claude model and limits, to see what each gives.

The short version: every arm landed all eight tasks, and on one machine neither was reliably faster. The differences are in what survives when the lead or hub process is killed, in limits and reported cost, and in where agents can run. nfltr took more setup, and spreading this batch over four machines made it slower. Every number below comes from the run log of 2026-10-02. Each main arm ran twice and the others once, so read the differences as what these runs showed, not as rates.

How each works

Agent teams, as the Claude Code documentation describes them: disabled by default and turned on with one environment variable, CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1. The lead spawns teammates, each a separate Claude Code instance with its own context window. Teammates share a task list, message each other directly by name, and you can talk to any of them from the lead's terminal. Teams form only in an interactive session, never with claude -p. The documentation lists, among the feature's limitations, that /resume does not restore in-process teammates.

nfltr orch: the hub is a Claude Code session with the hub tools (spawn_agent, send_message, wait_for_agents and others). Each agent is a non-interactive Claude Code process on a node, a machine you join with nfltr node join, which sets how many agents it runs at once. Agents report their result and cost to the hub. The hub's state is kept under its hub id, so a new session with the same id receives the results it missed.

The runs

Eight tasks from our task-diversity suite: two bug fixes, a feature, a rename, new tests, a docs change, an investigation and a refactor across many packages. A hidden test judges each task by its exit code. Every lead and hub got the same goal text; only the paragraph on how to run other agents differed. Every arm had 4 agent slots, and every lead, hub, teammate and agent used the same Claude model, in Claude Code 2.1.287.

  • Team: an interactive lead with agent teams on, teammates on a Mac (10 cores, 32 GB). Its brief said at most 4 teammates at once.
  • Hub, one machine: a hub session with only the hub tools; agents on the Mac, joined as one node running at most 4.
  • Hub, four machines: the same hub; agents on the Mac and three cloud VMs (4 vCPU, 16 GB each), one at a time per machine.
  • Lead with hub tools: an agent-teams lead with the hub tools added; teammates on the Mac (at most 2, per its brief) and agents on two of the VMs, one each.
One batch of 8 tasks per run, 2026-10-02. Every run passed 8 of 8 tasks with no operator intervention. Makespan is batch start to the lead's or hub's final answer. Team costs are priced from transcripts (see below); the others are as Claude Code reported them. Tokens are all kinds, cache reads included.
RunMakespanCostTokensMost at once
Team, Mac, run 1298 s$1.483.35 M6
Team, Mac, run 2494 s$1.843.84 M4
Hub, Mac node, run 1361 s$1.562.41 M4
Hub, Mac node, run 2385 s$1.451.99 M4
Hub, Mac + 3 VMs519 s$1.422.00 M4
Lead with hub tools1969 s$2.023.82 M2

In every run, one task first failed for a reason of our own making: the test runner's temporary directory had a path too long for a macOS unix socket. Rerun with a short path, it passed in every arm; on the untouched starting code it fails. The other seven passed in every run. No arm touched anything outside its tasks.

Where agent teams are as good or better

  • Same quality. All eight tasks landed and passed in every arm. Nothing here separates the two on outcome.
  • No speed winner on one machine. The fastest run was a team (298 s), whose lead ran up to 6 teammates at once. Its repeat, with 4 at once, took 494 s, slower than both hub runs (361 s and 385 s). The same bug fix took 34 s in one hub run and 225 s in the other. The Mac was shared with other builds: its load reached 45, and the background load moved by up to about 5 cores within one run, so the comparison is not clean. The variance between runs of the same arm is larger than the gap between arms.
  • Cost in the same range. $1.42 to $1.84 per 8-task batch for the team and hub runs, about $0.18 to $0.23 a task. Teams read more tokens (3.35 M and 3.84 M against 1.99 M to 2.41 M), mostly cached reads of each teammate's Claude Code context, which cost little at cache-read prices. The team costs were obtained differently: an interactive session reports no cost, so we priced every message in the lead's and teammates' transcripts at rates fitted to the same model's reported costs in non-interactive runs. The fit reproduced its two calibration reports exactly; no team run reports a cost to check it against.
  • Setup is one environment variable and an interactive session: no nodes, no git origin every machine can reach, no credentials on other machines. Adding the three VMs for the four-machine run took about 15 minutes.
  • Teammates talk to each other, and you can message any of them. Peer discussion is the documented strength of teams; our tasks were independent, so it was not measured here.
  • Concurrency on demand. A lead can run more at once than a fixed slot count, and the fastest run did.

Where nfltr orch differs

A killed lead or hub. In a separate test, each worker ran a command of about 180 s, and the lead or hub was killed with its whole process tree about 65 s in, then restarted the documented way after the workers would have finished.

  • Hub: the agents kept running on their node. A new session with the same hub id got both results from its first wait_for_agents and answered within 8 s.
  • Team: claude --continue restored the lead's conversation, not its teammates: 0 of 2 results recovered (“No agent named 'tm1' is reachable”). The teammates' shell commands kept running as orphans, their output going nowhere.
  • Lead with hub tools: the agent's result came back; the teammate's was lost.

Limits. A team runs as many teammates at once as the lead decides: told at most 4, the first team lead ran 6. A node's --max-agents held every hub run at its slots. Either can be what you want; only one is enforced. For spend, the hub counts each agent's reported cost and accepts hub-wide caps (--hub-max-tokens, --hub-max-usd, see Budgets) that stop running turns once passed. No run here came near its $6 limit, so a cap was not hit. An interactive team reports no cost at all.

Unattended runs. Teams never form with claude -p, so a team cannot run as a scripted or scheduled job, or as an agent on another machine. Hub agents always run non-interactively, and the hub in the hub runs was itself a claude -p session.

Other machines. Agents run wherever a node is joined, so work can go to the machine that holds the data or the tools.

A lead must be asked for a team. With a brief that stated only facts, our first team lead did every task itself, one after another, and was still on task 3 of 8 when we stopped it at 7.5 minutes. One sentence asking for the work to be done by teammates made it delegate all eight. The hub session in these runs has no edit or shell tools, so it has nothing to do but start agents.

What did not go well for nfltr

  • More machines did not help this batch. Four machines with one slot each took 519 s, against 361 s and 385 s on the Mac alone with 4 slots. The VMs were mostly idle (3 % and 2 % average busy; the one running a package's tests averaged 59 %): these tasks wait on the model, not on CPU. The run paid for distribution instead. The git origin was on the Mac, so agents on the VMs cloned through an SSH tunnel, and the first such clone took 121 s. Since these runs, a released change serves an agent's clone and fetch of an https origin from its node's repository mirror, which first updates itself from the origin, so only what is new crosses the network (Repositories). The first clone on each machine still crosses it.
  • The lead with hub tools took 1969 s. Offered teammates and agents, the lead started 8 hub agents and no teammate, so two VM slots carried the batch while 6 agents waited. Two agent turns also took 13 minutes each; one of them built the whole module (163 s) and then spent 610 s in a wait loop it wrote itself, though its brief asked only for its package's build, vet and test. That makespan measures one lead's choice and two agents' detours, not what the combination can do, and one run says nothing about how often a lead picks either. A node's operator can deny a command shape like that loop for every agent on the node with a tool policy (--tool-policy-file, see A shared machine); nfltr ships no pattern of its own.
  • Setup. Nodes, a git origin every node can reach and credentials on each machine; for the VMs, about 15 minutes of provisioning.
  • Any node of an account can take an unpinned spawn. In a probe, an agent from our hub was placed on a node joined to the same account for other work. --serve-hubs limits which hubs a node serves, not which nodes a hub uses; pin a spawn with the machine or labels constraint when it matters.

Using both

The two combine: a team lead can have the hub tools too, with teammates on this machine and agents on your nodes. Add the hub server, install its hooks and start the lead with agent teams on:

claude mcp add nfltr -- nfltr mcp --toolset hub
nfltr orch hub install-hook
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 claude

We found one gap doing this: a lead waiting in wait_for_agents saw a teammate's idle notification, which carries the teammate's final answer, only when Claude Code moved the call to the background, 118 s after the teammate went idle. Since these runs, nfltr orch hub install-hook also adds a TeammateIdle hook, so that wait ends the moment a teammate goes idle; a message a teammate sends without going idle still waits until the lead's turn ends. It also keeps the agent-teams switch out of the environment of agents on your nodes. Teammates still end with the lead; the agents do not. The setup is in Use nfltr from Claude Code.

Which to use

  • Agent teams for work on one machine, in an interactive session, where agents discussing and challenging each other helps, and where one environment variable is all the setup you want.
  • nfltr orch for work that must survive the session being closed or killed, run on the machine where the data or tools are, run unattended, or stay under a spending cap.
  • Both when part of the work is discussion on your machine and part must run elsewhere or outlive the session.

The whole comparison, including probes and discarded runs, used $12.13 of model spend.

Try it

Join each machine you want agents on, then give the hub a goal:

nfltr node join --max-agents 2
nfltr orch "<goal>"

The guides: Join machines as nodes, Use nfltr from Claude Code, and the hub tools reference. For agent teams, see the Claude Code documentation. What a hub is useful for, measured on other work, is in What nfltr orch is useful for, measured.

← All posts