nfltr orch refactored its own CLI

We gave an nfltr orch hub a goal twice in one day: split a large package of the nfltr CLI's own code into smaller ones. The hub and its agents did the refactor. In the first round they split the orchestration commands into seven packages, and the operator had to fix three things before the result passed the repository's checks. In the second round they split the CLI's root package into eight, and the operator fixed nothing.

This post covers how the runs were set up, what the second one measured against the first, what the first round's fixes changed, what got slower, and the gaps the second round found. Every number comes from the two run logs.

Setup

Both rounds used the same fleet: four nodes on one Mac, each joined with --max-agents 2, so 8 agent slots. The hub and every agent used the same Claude model. The spending cap was $30 in round 1 and $25 in round 2.

The plan was made without a model. A reference graph of the package's declarations shows what each command reaches; what two or more commands reach is a shared base, split out first, and with it out the parts no longer touch each other. The goal named the parts, their order and the rules: move code without changing behavior, keep every test, push only to the run's own branches. It pointed every agent at the plan.

What the hub did

In round 2 the hub followed the plan's shape, with no message from the operator:

  • At the start it spawned the base and one part that does not depend on it, together.
  • At 14.8 minutes, after the base reported, it spawned the six remaining parts in one call, each on its own branch off the base's commit.
  • At 36.5 minutes, after all six reported, it spawned one agent to merge the parts. Every conflict was in import, handler or setter lines of files that stay in the root package, none in moved code.
  • At 61.3 minutes it wrote its final answer.

Every part reported its build, vet and short tests passing, and test names identical before and after. All 8 parts and the merge were pushed. The operator rebased the result onto main and ran the repository's gates; all passed.

One run per round, 2026-10-02, both on four nodes with 8 slots on one Mac. Makespan is first spawn to last completion. Utilization is busy agent time over 8 slots.
Round 1: orchestration commandsRound 2: root package
Packages split out78
Agent turns811: 9 work turns, 1 stray placeholder, 1 follow-up message
Agent processes12, 4 of them gave up unclaimed11, none unclaimed
Most agents at once6 of 8 slots6 of 8 slots
Makespan43.6 min61.1 min
Utilization, overall / parallel phase0.22 / 0.350.30 / 0.60
Model spend$8.50about $5.95 ($6.81 reported, see below)
Operator steers00
Operator code fixes30

What the split bought

The root package went from 11,998 to 4,945 source lines, 59 % smaller, and from 225 to 83 top-level tests. No test was lost: 2,190 test names across the CLI before and after. The root package's own tests went from 28.9 s to 18.4 s, 37 % faster.

In round 1 the orchestration commands' package went from 22,769 to 9,778 source lines, 57 % smaller.

What got slower

The CLI's orchestration test gate got about 3 s slower: 36.7 s and 35.2 s after the split, against 32.5 s and 32.8 s before. The gate now covers the 8 new packages. Three packages in it, the root and two of the new ones, each build the full CLI binary for those tests, about 6 s of link each, and their relay tests now run beside one another. Inside the gate they take 26 to 31 s each, against 14 to 18 s alone. Two remedies are open: build the binary once per gate run, or keep the packages that are not orchestration commands out of this gate (they still run in the full suite). Neither is applied yet.

The run took longer than round 1, 61.1 minutes against 43.6, because each agent worked longer, not because agents waited. Parts ran 10 to 21 minutes, against 5 to 12 in round 1, and the merge 24.3 minutes against 9.2. In round 2 every agent ran the root package's tests and the repository's commit hook (vet, a Windows build, a dead-code check) while five others did the same on one Mac; one agent saw the root package's short tests take 233 s. Agents on separate machines, or checks scoped to what a part changed, would shorten it.

Running the hook was the brief's change after round 1. There, agents skipped it, and the operator's three fixes were two things the hook catches (an unused copied helper and three missing markers) and one bug in the hook itself, which read moved code as new. In round 2 every commit went through the fixed hook, and nothing needed fixing.

The round 1 fixes, measured live

  • A slow clone no longer stalls other agents. In round 1 the hub held a hub-wide lock while it waited for a worker to accept a turn, and a worker first clones the repository, which took 1.5 to 3.5 minutes, so every other spawn waited behind that clone. 4 of 12 agent processes gave up after their 2-minute claim timeout, and one part started 6.7 minutes after its spawn. Now no call to a node or worker runs under a hub-wide lock. In round 2 the six parts' agents launched in the same second and each started 3.9 to 18.3 s later. The wait in the parallel phase fell from 6.7 minutes to 0.3.
  • One repository mirror per machine. In round 1 each of the four nodes cloned its own mirror, about 330 MB each. In round 2 one 341 MB mirror served all four. Its one clone took 2.5 minutes; the other first-wave agent waited on the mirror's lock and started 3 s after the clone finished, and the second wave found the mirror ready.

The 3 s steps between the six starts come from that lock: each agent refreshes the mirror in turn. Letting agents that waited use the refresh that just finished is the next fix.

Gaps round 2 found

All three are being fixed.

  • A wake-up turn replaced the final answer. The merging agent gave its full report, then a timer it had started in the background finished, and Claude Code woke the session for one more turn. That turn's one line, that the notification was only a finished timer, became the agent's result. The hub noticed, sent one message asking for the full report and got it 9 s later. The fix being made keeps the first answer of a session as the result.
  • A resumed session's usage was counted twice. The follow-up turn reported the whole session's totals, and the hub added them to the first turn's. About $0.86 was counted twice, so the run reported $6.81 for about $5.95 of spend, and a spending cap would trip early.
  • The hub spawned a placeholder. Between its two spawn calls it started one agent whose whole objective was a template placeholder. The agent refused in 4 s for $0.05, and the hub called it a mistake and stopped it. This was the hub's error, and nfltr does not judge objectives; the fix is a sentence in the hub's instructions.

At the end, before they were stopped, the idle nodes had used 0.3 to 0.4 s of CPU each in 72 minutes, and the relay 5.8 s.

Try it

Join each machine you want agents on, then give the hub a goal:

nfltr node join --max-agents 2
nfltr orch "<goal>"

The guides: Join machines as nodes, Use nfltr from Claude Code if you want the hub in your own Claude Code session, and the hub tools reference. What a hub is useful for, measured on other work, is in What nfltr orch is useful for, measured.

← All posts