Build your own agent platform on nfltr
Hosted agent products give a team a box to type a task into, an agent that works on it somewhere else, and a result. We wanted to know how much of that a team can build itself on nfltr, with the agents on its own machines: build boxes, staging, production. So we wrote a small example service, published it, and recorded one working day on it.
Why not just Claude Code over SSH?
If you are one developer with a few machines you already reach, SSH or Tailscale plus Claude Code is probably the better setup: nothing to install, no relay to depend on, and one model session instead of one per agent. nfltr earns its place when some of these matter:
- Many machines. One request starts agents on several machines at once, each working where its data and tools are, up to each machine's limit. Twenty tasks on a Mac and two VMs finished 5.7× faster than one after another (Filling a fleet).
- Triggers. Work starts from a webhook, a failed CI run or a schedule, not only from someone at a terminal. Below, the investigation was on the pull request 33 s after CI failed.
- Policy enforced on the machine. Each machine's owner sets what agents there may do, and the machine checks it, not the session that asked. Below, the production node denied
kubectl rollout restartto a task without a grant. - Owner-signed approvals. A risky step runs only with the owner's signature over exactly that prompt, once: the node refused the same grant a second time.
- One record of who ran what and what it cost. Below, the platform showed cost per developer, and the relay showed each hub's agents under the key that started them.
- Results that outlive the requester. Agents keep running when the session or service that started them goes away. In a separate test a hub was killed mid-run; restarted with the same hub id, it got both results from its first wait and answered within 8 s (measured here).
More on the difference: nfltr vs. Claude Code over SSH or a VPN.
What it is
The service is an ordinary web app: about 800 lines of Python with FastAPI and SQLite, the GitHub integration included. It holds users, tokens, a task table and an audit log. For each team it starts one hub, nfltr mcp --toolset hub, as a child process and calls its tools over MCP, the same tools Claude calls when Claude is the hub: spawn_agent when someone submits a task, wait_for_agents to collect results, list_nodes for the machines, request_approval for production. There is no model in the service or the hub. The agents are Claude Code processes that nodes start on your machines, one per task, under each machine's own policy.
Tasks come in four ways: an HTTP call, a webhook, a /agent ... comment on a pull request, and a failed CI run. Developers can also skip the service and use the same machines from their own Claude Code, each on their own hub.
One recorded day
On 2026-10-04 we ran it on nfltr v1.0.673 against nfltr.xyz with three machines: two cloud VMs as a build box and a staging box, and a production machine that was simulated (its kubectl only writes a log). The repository and its CI on GitHub were real, and so was every Claude Code session. Each person had their own nfltr key, and the model was set per task, not by a default.
| Part | What happened | Model cost |
|---|---|---|
| A developer offloads two tasks from her own Claude Code | One spawn_agents call: the build box ran a flaky test 20 times (2 failures), the staging box checked the service read-only (about 8 % gateway timeouts). 35 s for the session | $0.47 |
| A pull request fails CI; the platform investigates | The comment was on the pull request 33 s after the failure. The agent reproduced the failure in 4 of 6 runs and named the cause: a retry test with a random gateway and no seed fails about 22 % of the time | $0.10 |
/agent fix ... and push a branch | A new branch with one test file changed and no build artefacts; the result comment came 31 s after the task started. CI passed on it, and the pull request's author had the build box run the test 30 times on the branch from her own Claude Code: 0 failures | $0.26 |
| A production restart | Without the owner's grant the node's policy denied kubectl rollout restart. The owner signed exactly the prompt the agent would get; with that grant the agent ran one restart. The same grant submitted again was refused by the node, naming the task that had used it | $0.29 |
| Who started what, and the cost | The platform's task list and cost per developer; on the relay, each hub's work under the key that started it | $0.05 |
Model spend for the day was $1.18, against a cap of $6. The two VMs ran for about 19 minutes, about $0.09.
Two things did not go as planned. The flaky test passed CI six times before it failed once, so the investigation waited for the seventh attempt. And the platform's GitHub token had no permission to read CI logs, so the investigation task said so and the agent worked from the repository alone.
The day also showed four small gaps, fixed since the recording: build machines had no git identity, so the agent committed under a name it chose; the team's house rules were added to every prompt, so the production owner signed them too, and now only tasks that name a repository get them; the platform's hub listed a machine that serves only another hub; and a dashboard key generated with an empty name got a random one, so the name is now required.
What is not included
- Organisations, single sign-on and quotas. The example has teams, users and bearer tokens; put your sign-on in front and map identities to teams.
- High availability. One service instance owns each team's hub at a time.
- A container per task. Agents get their own directory and a clean Claude config; for stronger isolation run nodes as separate users, in containers or in VMs.
- Spawn retries. A spawn has no idempotency key yet, so a service that crashes mid-spawn may leave an agent it does not know about; its result still arrives and is logged.
- Other agent runtimes. The nodes run Claude Code.
Try it
The tutorial walks through the parts: joining machines with their policies, driving the hub from code, receiving results without losing any, approvals and teams: Build your own agent platform. The service is at github.com/nfltr/agent-platform-example (Apache-2.0); copy it and change what does not fit. If you build on it, tell us what you needed that it does not do.