Build Your Own Agent Platform
Hosted agent products give you a box to type a task into, an agent that works on it in the cloud, and a result. This tutorial builds the same thing on your own machines: a small web service where your teams submit tasks over HTTP or from a webhook, Claude Code agents run them on machines you choose, and the answer, the machine and the cost come back to your service. nfltr supplies the machines, the transport, durable delivery and per-machine policy; your service makes every decision, with no model in it.
The complete service is at github.com/nfltr/agent-platform-example (Apache-2.0): copy it and adapt it.
/agent comment pushes a fix branch, a production restart runs only with the owner's signed grant and the same grant replayed is refused, and the platform shows who started what and what it cost. Recorded on 2026-10-04 with nfltr v1.0.673 on nfltr.xyz: two cloud VMs as build and staging machines, a simulated production machine, a real GitHub repository and CI; $1.18 of model spend.The shape
| Part | What it is | Runs a model? |
|---|---|---|
| Your service | An ordinary web app (the example is a few hundred lines of Python): users, tokens, a task table, an audit log, a page. | No |
| The hub | nfltr mcp --toolset hub, started by your service as a child process, one per team. An MCP server whose tools start, message, stop and list agents and return their results. | No |
| The relay | nfltr.xyz, or your own relay (beta). Carries messages; everything connects out to it. | No |
| Nodes | Your machines, joined once with nfltr node join. Each machine's owner sets what agents there may do. | No |
| Agents | Claude Code processes a node starts for one task, each in its own directory. | Yes |
The key point: the hub is an MCP server, so a program can call its tools exactly as Claude does. Your service calls spawn_agent when a user submits a task and wait_for_agents to receive results. Claude runs only in the agents.
1. Join your machines, each with its own policy
On every machine that should run agents, install nfltr and Claude Code and join it. The agents' Claude login comes from the machine's environment; your service never needs it.
$ curl -fsSL https://nfltr.xyz/install.sh | sh
$ nfltr config add-api-key # your nfltr key; prompts for it
$ export CLAUDE_CODE_OAUTH_TOKEN=... # or ANTHROPIC_API_KEY
# a build box for the payments team: agents may run commands and push to these repos
$ nfltr node join --name build-1 --max-agents 2 --harness claude-code --allow-all-tools \
--labels team=payments --allow-repo 'github.com/acme/payments-*' \
--serve-hubs platform-payments
# a production box: only tasks that name it; writes need the owner's signed grant
$ nfltr node join --name prod-1 --max-agents 1 --harness claude-code --allow-all-tools \
--labels team=payments,env=prod --serve-hubs platform-payments --require-pin \
--tool-policy-file prod-policy.json --approval-public-key nfltr-approval-pub-v1:...
--labelslets your service target a team's machines without naming each one.--serve-hubskeeps the machine to its team's hub. It sorts your own account's hubs; it is not an authentication boundary (see Teams).--require-pintakes only tasks that name this machine.--tool-policy-filedenies tool calls by pattern (for examplekubectl apply); rules markedunless_approvedlift only for a task carrying the owner's grant for exactly its prompt, checked by the node itself.--harness claude-codepicks the agent CLI when the machine has more than one.- Agents' commits name the git identity of the node's OS user (its
user.nameanduser.email, orGIT_AUTHOR_*andGIT_COMMITTER_*in the node's environment), read at join. Set one, for example a team bot; with none, the node sets none and an agent picks a name itself.
2. Drive the hub from code
Your service starts one hub per team with the official MCP SDK and keeps the session open. The relay address and key come from the environment (NFLTR_API_KEY) or the saved nfltr config, never from the command line. Pass the environment through: the SDK otherwise gives the server a minimal one.
params = StdioServerParameters(
command="nfltr",
args=["mcp", "--toolset", "hub", "--hub-id", hub_id,
"--task-store-file", os.path.join(state_dir, "nfltr", "tasks.json")],
env=dict(os.environ))
read, write = await stack.enter_async_context(stdio_client(params))
session = await stack.enter_async_context(ClientSession(read, write))
await session.initialize()
async def call(tool, args=None, timeout=120):
r = await session.call_tool(tool, args or {}, read_timeout_seconds=timeout)
text = r.content[0].text if r.content else ""
if r.is_error:
raise HubError(text) # the hub's reason, e.g. an unknown constraint
return json.loads(text)
Starting a task is one call. Everything about where and how it runs is a constraint your service passes through: machine, labels, workspace (a repository the machine allows), model, timeout_ms, budget, approval and so on. Nothing is filled in for you.
agent = await call("spawn_agent", {"prompt": prompt,
"constraints": {"labels": {"team": "payments"}}})
# {"agent_id": "agent-...", "status": "pending" | "running", ...}
A machine with no free slot keeps the task pending, with a waiting fact saying what it waits for; it starts when a slot frees.
Receiving results, without losing any
One background loop per hub calls wait_for_agents. It returns every finished turn with the agent's answer, machine and reported usage. Pass ack: the seqs of the results you have stored. A result you received but did not acknowledge comes again on the next call, also after your service restarts, so a crash between receiving and storing loses nothing; you recognise a repeat by its seq.
acked = []
while True:
work.clear()
r = await call("wait_for_agents", {"timeout_ms": 600000, "ack": acked}, timeout=660)
acked = []
for c in r.get("completions", []):
store_result(team, c) # status, text, machine, usage.tokens, usage.cost_usd
acked.append(c["seq"])
if not r.get("completions") and not r.get("pending"):
await work.wait() # nothing running: wait for our own next spawn
Agent text arrives framed as untrusted content (between <<<untrusted-content>>> markers), because the hub's results are also read by models. Strip the frame, and keep treating the text as data: show it as text, never as HTML or as instructions.
On restart your service reopens the same hub ids. Agents kept running on the nodes meanwhile; their results are waiting.
3. The web service
The rest is an ordinary web app. The example uses FastAPI and SQLite:
| Endpoint | Hub tool |
|---|---|
POST /tasks {prompt, machine?, labels?, repo?, approval?, constraints?} | spawn_agent; a refusal becomes HTTP 400 with the hub's reason |
GET /tasks, GET /tasks/{id} | Your own table; a running task adds list_agents' live status and progress |
POST /tasks/{id}/messages | send_message: steer a running agent, or continue a finished one in the same session |
POST /tasks/{id}/stop | stop_agent |
GET /machines | list_nodes: machines, labels, free slots |
POST /approvals {machine, prompt} | request_approval: the digest and command for the machine's owner |
POST /hooks/{team}/{name} | spawn_agent with the hook's configured prompt and the event as data |
Each team is a hub id in the config, with its users' tokens (stored as SHA-256), webhooks and an optional budget:
{"teams": {"payments": {
"hub_id": "platform-payments",
"budget": {"max_tokens": 5000000, "max_usd": 25},
"users": {"alice": "<sha256 of alice's token>"},
"hooks": {"ci-failure": {
"secret_sha256": "<sha256 of the hook's secret>",
"prompt": "A CI job failed. Find the cause from the log and the repository and say what fix you would make. Change nothing.",
"constraints": {"labels": {"team": "payments"}}}}}}}
The budget becomes the hub's own cap (--hub-max-tokens, --hub-max-usd): past it, the hub stops that team's running agents and refuses new ones.
4. Run it and submit a task
$ python3.12 -m venv .venv && .venv/bin/pip install -r requirements.txt
$ export NFLTR_API_KEY=... # or rely on the saved nfltr config
$ .venv/bin/uvicorn app:app --port 8080
$ curl -s localhost:8080/tasks -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"prompt":"Find why the payments-api tests are flaky and push a fix branch.",
"repo":"https://github.com/acme/payments-api","ref":"main",
"labels":{"team":"payments"},"constraints":{"mutates_workspace":true}}'
$ curl -s localhost:8080/tasks/agent-... -H "Authorization: Bearer $TOKEN"
A Dockerfile and compose file in the example run the service in a container; it needs your nfltr key and no Claude credentials. With your own relay, set NFLTR_SERVER, NFLTR_PROXY_URL and the relay's certificate pin (or run nfltr config set-relay once) for the service and the nodes alike.
A small real run: two machines, two tasks over the API and one from the CI webhook, all three answered in about 40 seconds for $0.27 together, each with its machine and cost on the task page.
5. Risky actions: the owner signs
- Your service calls
POST /approvals {"machine":"prod-1","prompt":"..."}; the hub'srequest_approvalreturns the prompt's digest and the owner's command. - The owner reads the prompt and signs it on their own machine, with a key only they hold:
$ nfltr orch hub approve --key-file ~/.config/nfltr/approval.key \ --node prod-1 --prompt-file fix.txt --ttl 15m --expect-digest <prompt_sha256> - Your service submits the same prompt to
prod-1with"approval": "nfltr-approval-v1...". The node checks the signature, its own name, the prompt and the expiry, and only then lifts itsunless_approvedrules for that one task. Your service and the hub can pass a grant on; they cannot make or widen one. The node spends the grant on that task and refuses it to any later one, even after it restarts. Until then a grant is a production change for whoever holds it, so the example passes it to the hub once and never stores, logs or returns it: team members see only"approval": "attached".
The key pair comes from nfltr orch hub approval-key --out ~/.config/nfltr/approval.key; its public half goes to the node's --approval-public-key.
Teams, audit and cost
- Teams. One hub id per team (or per user) gives each its own agents, results and budget. Your service decides who belongs to which team. Within one nfltr account every hub can reach the account's machines, so teams that must not trust each other get separate nfltr keys, or separate relays.
- Audit. The example records who submitted what (API or webhook), where it ran, its answer, and the tokens and dollars the agent reported, in a task table and an append-only JSON-lines log.
- Cost. Each result carries the usage the agent's Claude Code reported. A task stopped before it finished reports tokens but no cost.
What you'd build without nfltr
| Part | Build it yourself | With nfltr | Still yours |
|---|---|---|---|
| Agent processes: start, stop, workspace | A daemon on each machine that starts Claude Code per task, gives it a directory, stops it and reaps its children | nfltr node join; spawn_agent and stop_agent; each agent in its own directory and process group, with a clean Claude config | Isolation beyond a directory: OS users, containers, VMs |
| Reaching machines behind NAT | Inbound ports, a VPN or a bastion, and reconnects | Nodes, hubs and your service connect out to the relay (nfltr.xyz or your own); no inbound port | Running your own relay, if you choose to |
| Results that survive restarts | A queue, result storage, de-duplication, resume after a crash | Agents keep running while your service is down; wait_for_agents with ack delivers each result once, and again only if you did not acknowledge it; reopening the same hub id picks up what you missed | Storing results; retrying a spawn safely (a spawn has no idempotency key yet) |
| Placement | An inventory of machines, labels and free slots, and a wait queue | machine and labels are hard constraints; a task stays pending until an eligible slot frees; list_nodes shows free slots | Deciding which machine or label each task needs |
| Per-machine tool policy and owner approvals | Rules the agent cannot bypass, a signing scheme, single-use grants | --tool-policy-file rules checked by the node; unless_approved rules lift only for the owner's grant for exactly that prompt, and the node spends it once | Who owns each machine and its key, and what the rules say |
| Audit and cost per agent | Collecting each agent's usage and tying it to a machine and a request | Each result carries the machine and the usage and cost the agent's Claude Code reported; nfltr orch task list and the dashboard list every agent; --hub-max-tokens and --hub-max-usd cap a hub | Which person in your organisation asked (your service's records) |
| Everything else | Not provided by nfltr: identity and SSO, organisation roles, secrets management, sandboxing and containers, high availability, billing, the UI | All of it, either way | |
A workflow engine such as Temporal can sit above this: it decides what happens and when (steps, timers, waiting on people); nfltr decides where and how each agent runs. A workflow step would call spawn_agent and finish on wait_for_agents; mind the missing idempotency key if the engine retries the step. We haven't built that combination.
What the example does not do
- Single sign-on, organisations and quotas: put your SSO gateway in front and map identities to teams.
- A container per task: agents get their own directory and a clean Claude config; for stronger isolation, run nodes as separate OS users, in containers or VMs.
- High availability: one service instance owns each team's hub at a time.
- Spawn retries: a spawn call has no idempotency key yet, so a service that crashes mid-spawn may leave an agent it does not know about (its result still arrives and is logged).
- Other agent runtimes: the nodes run Claude Code here.
Get the example
Clone github.com/nfltr/agent-platform-example (Apache-2.0). It contains app.py, hubclient.py, the GitHub integration, the task page, an example config, a Dockerfile, a compose file and the team presets (join-node.sh, tool policies, a .mcp.json for developers' own Claude Code); its README covers each one.