Build Your Own Agent Platform

Hosted agent products give you a box to type a task into, an agent that works on it in the cloud, and a result. This tutorial builds the same thing on your own machines: a small web service where your teams submit tasks over HTTP or from a webhook, Claude Code agents run them on machines you choose, and the answer, the machine and the cost come back to your service. nfltr supplies the machines, the transport, durable delivery and per-machine policy; your service makes every decision, with no model in it.

The complete service is at github.com/nfltr/agent-platform-example (Apache-2.0): copy it and adapt it.

A developer's day on this service (2 min 37 s, captions in the picture, no sound): a developer offloads two tasks from her own Claude Code, a failed CI run becomes an investigation comment on the pull request, a /agent comment pushes a fix branch, a production restart runs only with the owner's signed grant and the same grant replayed is refused, and the platform shows who started what and what it cost. Recorded on 2026-10-04 with nfltr v1.0.673 on nfltr.xyz: two cloud VMs as build and staging machines, a simulated production machine, a real GitHub repository and CI; $1.18 of model spend.

The shape

PartWhat it isRuns a model?
Your serviceAn ordinary web app (the example is a few hundred lines of Python): users, tokens, a task table, an audit log, a page.No
The hubnfltr mcp --toolset hub, started by your service as a child process, one per team. An MCP server whose tools start, message, stop and list agents and return their results.No
The relaynfltr.xyz, or your own relay (beta). Carries messages; everything connects out to it.No
NodesYour machines, joined once with nfltr node join. Each machine's owner sets what agents there may do.No
AgentsClaude Code processes a node starts for one task, each in its own directory.Yes

The key point: the hub is an MCP server, so a program can call its tools exactly as Claude does. Your service calls spawn_agent when a user submits a task and wait_for_agents to receive results. Claude runs only in the agents.

1. Join your machines, each with its own policy

On every machine that should run agents, install nfltr and Claude Code and join it. The agents' Claude login comes from the machine's environment; your service never needs it.

$ curl -fsSL https://nfltr.xyz/install.sh | sh
$ nfltr config add-api-key                  # your nfltr key; prompts for it
$ export CLAUDE_CODE_OAUTH_TOKEN=...        # or ANTHROPIC_API_KEY

# a build box for the payments team: agents may run commands and push to these repos
$ nfltr node join --name build-1 --max-agents 2 --harness claude-code --allow-all-tools \
    --labels team=payments --allow-repo 'github.com/acme/payments-*' \
    --serve-hubs platform-payments

# a production box: only tasks that name it; writes need the owner's signed grant
$ nfltr node join --name prod-1 --max-agents 1 --harness claude-code --allow-all-tools \
    --labels team=payments,env=prod --serve-hubs platform-payments --require-pin \
    --tool-policy-file prod-policy.json --approval-public-key nfltr-approval-pub-v1:...

2. Drive the hub from code

Your service starts one hub per team with the official MCP SDK and keeps the session open. The relay address and key come from the environment (NFLTR_API_KEY) or the saved nfltr config, never from the command line. Pass the environment through: the SDK otherwise gives the server a minimal one.

params = StdioServerParameters(
    command="nfltr",
    args=["mcp", "--toolset", "hub", "--hub-id", hub_id,
          "--task-store-file", os.path.join(state_dir, "nfltr", "tasks.json")],
    env=dict(os.environ))
read, write = await stack.enter_async_context(stdio_client(params))
session = await stack.enter_async_context(ClientSession(read, write))
await session.initialize()

async def call(tool, args=None, timeout=120):
    r = await session.call_tool(tool, args or {}, read_timeout_seconds=timeout)
    text = r.content[0].text if r.content else ""
    if r.is_error:
        raise HubError(text)          # the hub's reason, e.g. an unknown constraint
    return json.loads(text)

Starting a task is one call. Everything about where and how it runs is a constraint your service passes through: machine, labels, workspace (a repository the machine allows), model, timeout_ms, budget, approval and so on. Nothing is filled in for you.

agent = await call("spawn_agent", {"prompt": prompt,
                                   "constraints": {"labels": {"team": "payments"}}})
# {"agent_id": "agent-...", "status": "pending" | "running", ...}

A machine with no free slot keeps the task pending, with a waiting fact saying what it waits for; it starts when a slot frees.

Receiving results, without losing any

One background loop per hub calls wait_for_agents. It returns every finished turn with the agent's answer, machine and reported usage. Pass ack: the seqs of the results you have stored. A result you received but did not acknowledge comes again on the next call, also after your service restarts, so a crash between receiving and storing loses nothing; you recognise a repeat by its seq.

acked = []
while True:
    work.clear()
    r = await call("wait_for_agents", {"timeout_ms": 600000, "ack": acked}, timeout=660)
    acked = []
    for c in r.get("completions", []):
        store_result(team, c)        # status, text, machine, usage.tokens, usage.cost_usd
        acked.append(c["seq"])
    if not r.get("completions") and not r.get("pending"):
        await work.wait()            # nothing running: wait for our own next spawn

Agent text arrives framed as untrusted content (between <<<untrusted-content>>> markers), because the hub's results are also read by models. Strip the frame, and keep treating the text as data: show it as text, never as HTML or as instructions.

On restart your service reopens the same hub ids. Agents kept running on the nodes meanwhile; their results are waiting.

3. The web service

The rest is an ordinary web app. The example uses FastAPI and SQLite:

EndpointHub tool
POST /tasks {prompt, machine?, labels?, repo?, approval?, constraints?}spawn_agent; a refusal becomes HTTP 400 with the hub's reason
GET /tasks, GET /tasks/{id}Your own table; a running task adds list_agents' live status and progress
POST /tasks/{id}/messagessend_message: steer a running agent, or continue a finished one in the same session
POST /tasks/{id}/stopstop_agent
GET /machineslist_nodes: machines, labels, free slots
POST /approvals {machine, prompt}request_approval: the digest and command for the machine's owner
POST /hooks/{team}/{name}spawn_agent with the hook's configured prompt and the event as data

Each team is a hub id in the config, with its users' tokens (stored as SHA-256), webhooks and an optional budget:

{"teams": {"payments": {
  "hub_id": "platform-payments",
  "budget": {"max_tokens": 5000000, "max_usd": 25},
  "users": {"alice": "<sha256 of alice's token>"},
  "hooks": {"ci-failure": {
    "secret_sha256": "<sha256 of the hook's secret>",
    "prompt": "A CI job failed. Find the cause from the log and the repository and say what fix you would make. Change nothing.",
    "constraints": {"labels": {"team": "payments"}}}}}}}

The budget becomes the hub's own cap (--hub-max-tokens, --hub-max-usd): past it, the hub stops that team's running agents and refuses new ones.

4. Run it and submit a task

$ python3.12 -m venv .venv && .venv/bin/pip install -r requirements.txt
$ export NFLTR_API_KEY=...          # or rely on the saved nfltr config
$ .venv/bin/uvicorn app:app --port 8080

$ curl -s localhost:8080/tasks -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
    -d '{"prompt":"Find why the payments-api tests are flaky and push a fix branch.",
         "repo":"https://github.com/acme/payments-api","ref":"main",
         "labels":{"team":"payments"},"constraints":{"mutates_workspace":true}}'
$ curl -s localhost:8080/tasks/agent-... -H "Authorization: Bearer $TOKEN"

A Dockerfile and compose file in the example run the service in a container; it needs your nfltr key and no Claude credentials. With your own relay, set NFLTR_SERVER, NFLTR_PROXY_URL and the relay's certificate pin (or run nfltr config set-relay once) for the service and the nodes alike.

A small real run: two machines, two tasks over the API and one from the CI webhook, all three answered in about 40 seconds for $0.27 together, each with its machine and cost on the task page.

5. Risky actions: the owner signs

  1. Your service calls POST /approvals {"machine":"prod-1","prompt":"..."}; the hub's request_approval returns the prompt's digest and the owner's command.
  2. The owner reads the prompt and signs it on their own machine, with a key only they hold:
    $ nfltr orch hub approve --key-file ~/.config/nfltr/approval.key \
        --node prod-1 --prompt-file fix.txt --ttl 15m --expect-digest <prompt_sha256>
  3. Your service submits the same prompt to prod-1 with "approval": "nfltr-approval-v1...". The node checks the signature, its own name, the prompt and the expiry, and only then lifts its unless_approved rules for that one task. Your service and the hub can pass a grant on; they cannot make or widen one. The node spends the grant on that task and refuses it to any later one, even after it restarts. Until then a grant is a production change for whoever holds it, so the example passes it to the hub once and never stores, logs or returns it: team members see only "approval": "attached".

The key pair comes from nfltr orch hub approval-key --out ~/.config/nfltr/approval.key; its public half goes to the node's --approval-public-key.

Teams, audit and cost

What you'd build without nfltr

PartBuild it yourselfWith nfltrStill yours
Agent processes: start, stop, workspaceA daemon on each machine that starts Claude Code per task, gives it a directory, stops it and reaps its childrennfltr node join; spawn_agent and stop_agent; each agent in its own directory and process group, with a clean Claude configIsolation beyond a directory: OS users, containers, VMs
Reaching machines behind NATInbound ports, a VPN or a bastion, and reconnectsNodes, hubs and your service connect out to the relay (nfltr.xyz or your own); no inbound portRunning your own relay, if you choose to
Results that survive restartsA queue, result storage, de-duplication, resume after a crashAgents keep running while your service is down; wait_for_agents with ack delivers each result once, and again only if you did not acknowledge it; reopening the same hub id picks up what you missedStoring results; retrying a spawn safely (a spawn has no idempotency key yet)
PlacementAn inventory of machines, labels and free slots, and a wait queuemachine and labels are hard constraints; a task stays pending until an eligible slot frees; list_nodes shows free slotsDeciding which machine or label each task needs
Per-machine tool policy and owner approvalsRules the agent cannot bypass, a signing scheme, single-use grants--tool-policy-file rules checked by the node; unless_approved rules lift only for the owner's grant for exactly that prompt, and the node spends it onceWho owns each machine and its key, and what the rules say
Audit and cost per agentCollecting each agent's usage and tying it to a machine and a requestEach result carries the machine and the usage and cost the agent's Claude Code reported; nfltr orch task list and the dashboard list every agent; --hub-max-tokens and --hub-max-usd cap a hubWhich person in your organisation asked (your service's records)
Everything elseNot provided by nfltr: identity and SSO, organisation roles, secrets management, sandboxing and containers, high availability, billing, the UIAll of it, either way

A workflow engine such as Temporal can sit above this: it decides what happens and when (steps, timers, waiting on people); nfltr decides where and how each agent runs. A workflow step would call spawn_agent and finish on wait_for_agents; mind the missing idempotency key if the engine retries the step. We haven't built that combination.

What the example does not do

Get the example

Clone github.com/nfltr/agent-platform-example (Apache-2.0). It contains app.py, hubclient.py, the GitHub integration, the task page, an example config, a Dockerfile, a compose file and the team presets (join-node.sh, tool policies, a .mcp.json for developers' own Claude Code); its README covers each one.