A faster, leaner relay, measured against frp, rathole, chisel and bore

Every byte that goes through nfltr goes through the relay. A share link to a dev server, an SSH session through nfltr tcp, a hub talking to its agents on other machines: each of them is a connection that the relay accepts, routes and copies. If the relay is slow, everything built on it is slow; if it drops a byte, something above it breaks in a way that is hard to trace back.

On October 6 we put the relay in one Docker harness next to four popular open-source reverse tunnels, frp, rathole, chisel and bore, and measured them the same way. nfltr lost most of the comparisons, and one of them showed a bug that lost response bytes. This post covers the three days that followed: how we measured, how we found the causes, what we changed in seven rounds, what we tried and threw away, the conformance suites we ran and the bugs they found, one benchmark we almost got wrong, and where nfltr stands now, including where it still trails.

Why, and what we were aiming for

The relay is transport. It runs no model and makes no decisions; it carries connections between clients and the agents (nfltr http, nfltr tcp, nodes and hubs) that dialled out to it. That makes it the one component every nfltr feature depends on, and the one we can compare directly with tools whose only job is the same.

The goal we set was plain: on the same hardware limits, match or beat frp and the others on

  • throughput: bulk TCP, small and large HTTP requests, with and without keep-alive, and many requests in flight through one tunnel;
  • scaling: hundreds of tunnels on one relay, and memory per tunnel;
  • reliability: no lost bytes, no crashes under load, a fast return after the relay restarts;
  • footprint: idle CPU and memory, and the size of the binary you download.

frp is the reference point. It is the most widely used of the four, it encrypts its own transport by default, and it has a public HTTP routing mode like nfltr's share links. rathole and bore are small Rust tools, chisel tunnels over SSH. All four are open source; nfltr is closed source, which is why this post describes mechanisms rather than code.

How we measured

The harness

Everything ran on one host: a 10-CPU arm64 machine running Docker. Each tool got the same four containers on one internal Docker network with no route out:

ContainerRoleLimit
relaythe server: nfltr relay serve, frps, the rathole server, the chisel server, the bore server1 CPU, 512 MiB
exposerthe tunnel client next to the service: nfltr tcp or nfltr http, frpc, the rathole client, chisel's client, bore local1 CPU, 512 MiB
targetthe service behind the tunnel: nginx serving 1 KiB and 1 MiB files with keep-alive, and an iperf3 server1 CPU, 512 MiB (see fairness)
clientthe load generators, oha 1.16.0 and iperf3 3.12, plus the client half of a tool that needs one (nfltr tcp-connect)4 CPU, 2 GiB
A fifth container starts many tunnel clients for the tunnel-count test. Versions: frp 0.71.0, rathole 0.5.0, chisel 1.12.0, bore 0.6.0, all upstream release binaries; nfltr built from source at the commit under test.

One CPU per relay and exposer is deliberate. It turns the benchmark into a measure of how much work each tool does per byte and per request, which is what decides cost on a small VM, rather than how many cores it can spread across.

The tests, per tool:

  1. Bulk TCP: iperf3 for 10 s through the tunnel with 1 and with 8 parallel streams, three runs, median.
  2. HTTP through a TCP tunnel: oha for 10 s against the 1 KiB file with 1 and 50 connections and keep-alive, 50 connections without keep-alive (a new tunnel connection per request), and the 1 MiB file at 50 connections.
  3. Public HTTP URL (a virtual host on the relay, routed by host name): the same 1 KiB requests, for the tools that have one.
  4. Concurrency: 1, 100 and 1,000 connections through one tunnel, keep-alive and not. A step is skipped once errors pass 1 %.
  5. Reconnect: docker restart of the relay while the tunnel clients keep running, then the time until a request through the tunnel succeeds again.
  6. Tunnel count: 1, 50, 200 and later 500 client processes with one tunnel each on a fresh relay: the time until all of them answer, and the relay's memory and CPU while they idle.
  7. Idle and load cost: CPU from the cgroup's counters over a 60 s idle window and over the load runs; memory as resident size and cgroup anonymous memory, sampled every second.
  8. Binary size: the server and client binaries, and the official image where one exists.

For nfltr only, a profiling pass built the same source with a profiling endpoint added at build time (no product code changed) and took 30 s CPU profiles, heap and goroutine dumps of the relay, the agent and tcp-connect during the bulk and HTTP loads.

Fairness choices

The tools do not do the same work, so we wrote down what each one does and made a few choices to keep the comparison honest.

  • Encryption is recorded per tool. nfltr's TCP tunnel is TLS from the client to the relay and TLS from the relay to the agent; the relay sees the bytes in between. frp encrypts the frpc-to-frps leg (its default) but not the client-to-relay leg. rathole ran twice, with its Noise encryption and with none (the 0.5.0 release has no TLS). chisel runs SSH over WebSocket. bore encrypts nothing. When nfltr beats an unencrypted tool or loses to one, the table says so.
  • Like for like on encrypted public HTTP. nfltr's public HTTP URL is end-to-end encrypted by default: the client's TLS goes through the relay unopened and ends at the agent. frp's plain HTTP virtual host is not encrypted at all. So we added frp-https: frp's HTTPS virtual host, which routes by server name and passes TLS through to the target nginx. Each new client connection then pays a full TLS handshake with the far side, as it does with nfltr. The plain frp row stays in the tables for reference.
  • Moved work is not free. frp-https ends TLS in the target container, nfltr ends it in the agent. In the first runs the target had no CPU limit, so frp's TLS work was outside the limit and nfltr's inside it. From round 4 on, the target runs under the same 1 CPU and 512 MiB, and every HTTP row records CPU-seconds per 1,000 requests: the CPU used by relay, exposer, target and the client-side tunnel process (tcp-connect), but not the load generator.
  • nfltr ran with the per-agent rate limit off. nfltr relay serve limits each agent to 100 requests per second by default, which would cap every HTTP test at 100. The benchmark passes --agent-rate-limit 0. The end-to-end encrypted path is not rate limited either way.
  • Same-window comparisons. The host was shared with other work. The same nfltr binary moved by up to 2× between runs an hour apart, and frp's own 8-stream number went from 5.5 to 9.9 Gbit/s. From round 3 on, every before-and-after comparison ran interleaved in one window: old build, new build, frp, frp-https, repeated, medians taken. Numbers from different runs are compared only where the table says so.
  • Ratios across runs, not absolute numbers. The run on the round 5 release (v1.0.676) had a faster host than the earlier final run (v1.0.675): frp itself roughly doubled, from 3.98 to 7.45 Gbit/s on one stream and from 36.4k to 73.3k small requests per second. So across full runs we compare nfltr's number as a share of frp's in the same run, never absolute numbers. The earlier runs stay below as history.
  • Power recorded per tool. The host is a laptop, and on battery it runs the same build 40-45 % slower. From the final run on, the harness records the power source and Low Power Mode for every tool and flags a run where they changed (see the benchmark we almost got wrong).

Caveats that apply to everything below: one host, a Docker bridge, no network latency or loss. Over a real WAN, round trips (connection setup, reconnect) matter more and CPU per byte matters less. Single 10 s oha runs carry a few percent of run-to-run noise.

The baseline

The first full run, on the code as it was on the morning of October 6:

Testnfltrfrprathole (Noise / none)chiselbore
iperf3, 1 stream (Gbit/s)1.993.971.02 / 20.711.8220.74
iperf3, 8 streams (Gbit/s)1.875.481.01 / 13.661.7911.06
1 KiB, c=50, keep-alive (req/s)32,66135,85931,512 / 60,31846,17757,546
1 KiB, c=50, new connection each (req/s)1,3384,4822,000 / 2,6647,2242,380
1 MiB, c=50, keep-alive (req/s)389923108 / 900376913
Public URL, 1 KiB, c=50, keep-alive (req/s)8,269 (encrypted end to end)7,915 (plain HTTP)no public URL mode
Public URL, c=50, new connection each1,728, 32.1 % errors6,676
Back after relay restart (s)5.412.727.67 / 3.901.21never (exits)
200 tunnels registered118 of 200 in 303 s200 in 1.9 s200 in 4.8 s200 in 1.9 s200 in 0.6 s
Idle relay memory (RSS, MiB)38.419.24.3 / 3.77.41.3
Binary (MiB)107.117.4 / 14.11.010.32.6
Baseline, 2026-10-06. TCP rows go through each tool's TCP tunnel. nfltr's binary is the harness build (linux/arm64, symbols kept) and holds the relay and the client in one file; frp has a server and a client binary. The target had no CPU limit in this run.

What it showed:

  • Bulk throughput was half of frp's, and 8 streams were slower than 1, which should not happen.
  • A new connection per request was the slowest of the field, about 4× slower than frp.
  • The public URL lost response bytes when connections closed: 32 % of requests failed without keep-alive. That is a correctness bug, not a speed problem.
  • At 1,000 concurrent connections through one tunnel, the agent's container hit its 512 MiB limit and the kernel killed it. No other tool was killed at any step.
  • With 200 agents, only 118 ever registered, while every other tool registered 200 in under 5 seconds.
  • Reconnect took 5.4 s, against 1.2 s for chisel.
  • The binary was 107 MiB.

It was not all bad. Small keep-alive requests were level with frp and rathole, nfltr had the lowest 99th-percentile latency of all tools at 1,000 connections (28.7 ms against frp's 67.5), and its end-to-end encrypted public URL matched frp's unencrypted one at 50 connections.

How we found the causes

Each problem went through the same steps:

  1. The benchmark says what is slow and by how much.
  2. A static read of the data path: we drew every hop a byte takes for each kind of tunnel, with its buffers, goroutines, queues and copies, and wrote down sixteen candidate causes ranked by expected payoff over effort. Each was a hypothesis, not a finding.
  3. Profiles under load: CPU, heap and goroutine profiles of each process, and later execution traces and the cgroup's throttling counters, to confirm or kill each hypothesis.
  4. A focused harness that reproduces the problem in seconds: the relay's tunnel dispatcher and the agent's tunnel loop in one test process, talking over a real loopback gRPC connection with the production transport settings, with a target that can upload, download, echo or stop reading. A second harness runs the real binaries on one host with each process limited to one Go processor, to mimic the container limit.
  5. A test that fails on the old code, then the fix, then a Docker rerun.

Some of what that turned up:

The response that lost its tail

The 32 % errors on the public URL came from a race between data and close. When a service writes its last bytes and closes the connection, the agent sends the relay a few data frames and then a close frame, back to back. The relay's per-connection loop waited on two things at once, "data arrived" and "session closed", and when both were ready Go picks one at random. Often it picked the close, and the data still queued was thrown away. Every response from a server that closes its connection was at risk: requests without keep-alive, Connection: close, and any TCP service that writes and hangs up. Keep-alive traffic hid it, because the connection stays open.

The harness reproduced it at once: 38 of 100 connection-per-request requests failed with an unexpected end of file, and 1 MiB downloads lost 32 KiB at the end. The rule now: a close marks a session's queue closed, and the writer delivers everything queued before it closes the client connection.

The 100-agent wall

With 200 agents, registration stopped near 100 while the relay sat at 0.5 % CPU: nothing was busy, agents were being turned away. The cause was a default. A limit on connected agents defaulted to 100; our hosted relay sets it to 0 (no limit), but nfltr relay serve and the local relay set nothing and inherited 100. Agent 101 got an answer that every agent treats as transient, so it printed "Connected", lost its stream at once, and retried forever with a backoff growing to about 30 s, never saying why. Because the count was taken before a session registered, a burst of agents overshot it, which is why the benchmark saw 118 rather than 100.

Now there is no limit unless the operator sets one with nfltr relay serve --max-agents. When the limit is hit, the agent exits once with a message naming it and the flag, and is not retried. The count and admission happen under one lock, so 32 agents connecting at once into room for 2 admit exactly 2.

The configuration lookup that copied everything

The public URL mode where the relay ends TLS and reads HTTP (we call it inspectable; the default is end-to-end encrypted) ran at 1,700 requests per second with the relay at 93-98 % CPU, about 0.6 ms of relay CPU per request. The CPU profile blamed system calls. The allocation profile told the real story: the relay allocated 306 KiB per request, and 86 % of all its allocated bytes came from one function in our configuration library.

The way the relay registered its command-line flags with that library meant that every read of a dotted setting, such as the share domain or a timeout, copied all of the relay's ~211 flags into a fresh map, formatting each value as text: about 35 KB per read. The request path read settings 10 to 20 times per request. The system call time in the CPU profile was the kernel handing out fresh pages for all that garbage.

We changed how flags are loaded (an explicitly set flag becomes an override, every other flag a default) without changing what any read returns, which a test checks key by key against the old behaviour. Allocation per request fell from 306 to 21 KiB, and relay CPU per request from 0.28 to 0.05 ms. It also fixed every other per-request settings read in the relay at once, which is why it was fixed where the flags are loaded rather than at each caller.

Why 8 streams were slower than 1

It was not CPU throttling (the cgroup counted 1-4 throttled periods of 322), not flow-control windows, and not lock contention. An execution trace of the relay during the 8-stream test showed a queue of queues in front of one serial stage:

  • the 8 session readers spent most of their time blocked because the agent's outbound queues were full;
  • the stream senders behind those queues were blocked on gRPC's per-stream write quota;
  • all of an agent's streams share one HTTP/2 connection, so one goroutine encrypts and writes everything for that agent. It was 38 % of the relay's CPU samples, 60 % of that in write system calls, one per 16 KiB TLS record;
  • and that writer was ready to run but not running for 2.4 of the 3 traced seconds.

More streams only added hand-offs in front of the one writer. That shaped rounds 3 to 5: fewer system calls per byte, then more than one connection for an agent that is busy, then taking those connections off gRPC.

A buffer pool that cleared 1 MiB for every 64 KiB

In the bulk-transfer profiles, clearing memory was 11 % of the relay's CPU and 15 % of the agent's. gRPC's default buffer pool zeroes every buffer it hands out, and its size classes jump from 32 KiB straight to 1 MiB. Every 64 KiB tunnel frame therefore took a 1 MiB buffer and cleared all of it, on send and again on receive. This also explained a puzzle from round 1: making tunnel frames smaller (256 to 64 KiB, to bound memory) made 1 MiB HTTP responses slower, from 389 to 330 requests per second, because the clearing per byte grew fourfold.

An agent that kept dialling its own service

On the public URL, 20-25 % of the agent's CPU went into opening TCP connections to the local service. The agent's reverse proxy used a client that keeps only 2 idle connections per host. With 50 sessions in flight, nearly every request opened a new connection to the service and left one in TIME_WAIT. In the harness, 32 sessions of 10 requests each made 226 connections to the service.

Concurrent sends, and a leak per session

Two bugs on the agent showed up only with several sessions at once. Every session's reader sent on one gRPC stream concurrently, which gRPC does not allow. The race detector stayed silent; instead the 8-session harness test deadlocked, with readers parked on the stream's write quota and both receive loops idle. And every end-to-end encrypted session left its HTTP server goroutine behind, because the one-connection listener it served from never returned from its second accept: 100 sessions left 97 goroutines.

Other findings from the same read: one slow client could stall every other session of the same agent (they shared one receive loop that wrote to each connection inline); the relay could queue up to 4,096 frames of 256 KiB per agent; a channel was closed while senders could still write to it, with the resulting panic hidden by a recover; and one client that sent nothing could stall every accept on the encrypted port for up to 5 s, because the relay read each TLS ClientHello on its accept loop.

What changed, round by round

The work ran in seven rounds, each merged and measured before the next. The table follows nfltr through the full Docker runs up to round 5; rounds 3 to 6 were also measured in same-window comparisons, shown in their own sections, and round 7 was about correctness rather than speed. The final run, after round 7, is in where nfltr stands now.

nfltrBaselineAfter round 1After round 2v1.0.675frp, same runRound 5 (v1.0.676)frp, same run
iperf3, 1 stream (Gbit/s)1.992.063.434.923.989.287.45
iperf3, 8 streams (Gbit/s)1.871.892.853.785.6110.539.66
1 KiB, c=50, new connection each (req/s)1,3382,8107,5587,4744,34011,4916,972
1 MiB, c=50, keep-alive (req/s)3893304344868871,0711,654
Public URL, 1 KiB, c=50, keep-alive (req/s)8,2697,20119,37037,92136,120 (frp-https)66,70970,833 (frp-https)
Public URL, c=50, new connection each1,728 (32.1 % errors)2,840 (0 %)3,7184,0352,477 (frp-https)7,6624,762 (frp-https)
Inspectable URL, c=50, keep-alive (req/s)1,7051,6866,7267,1277,886 (plain HTTP)11,29515,442 (plain HTTP)
Back after relay restart (s)5.412.692.750.842.761.172.24
200 tunnels registered118 in 303 s100 in 305 s200 in 3.2 s200 in 2.9 s200 in 2.0 s200 in 2.6 s200 in 1.7 s
Binary, harness build (MiB)107.175.675.675.717.4 + 14.175.817.4 + 14.1
Full Docker runs on 2026-10-06 (three) and 2026-10-07 (two: v1.0.675 after round 4, v1.0.676 after round 5), relay and exposer at 1 CPU each. The host's speed moved between runs: the round 2 run was taken on a busier host than the others (frp's numbers rose up to 1.9× between runs of the same frp binary), and the round 5 run on a faster one (frp roughly doubled against the v1.0.675 run). Compare each nfltr column with the frp column of its own run, not across columns. The target container was CPU-limited only in the last two runs.

Round 1: correctness first

The first round fixed what was wrong before making anything faster.

  • No byte lost on close. Each session got its own byte-bounded queue and writer; a close is delivered after the data queued before it. The public URL went from 32 % errors to none.
  • One sender per gRPC stream. The agent's session readers hand frames to one sender goroutine through a 64-frame queue. We chose a queue over a lock so a reader can read its next chunk while the previous one is sent: single-session downloads went from about 400 to 515-713 MB/s against the lock variant in the harness.
  • Sessions isolated from each other. Every session on both sides writes through its own queue (1 MiB) and goroutine, so a client or service that stops reading no longer stalls its siblings. Before, the stuck-target test stalled the other session and the tunnel could not even shut down.
  • Memory bounded in bytes, not frames. 1 MiB per session one way, 64 frames of 64 KiB per agent the other way, 64 KiB read buffers. Peak heap of relay and agent together, 8 sessions uploading, went from 139-148 to 51-70 MiB.
  • The leak and the panic. Each encrypted session now releases its server when its connection closes (97 leaked goroutines per 100 sessions to none), and the queue that senders write to is never closed under them.
  • gRPC flow-control windows sized for latency. At 40 ms round-trip time the old windows gave 72 MB/s; turning on gRPC's own bandwidth estimator gave 64-77, no better; larger fixed windows gave 224-250. Each side now grants the window that bounds what the other side can park in its memory. Uploads at 40 ms went to 254-258 MB/s, downloads to 132-134.
  • Tunnels as HTTP/2 streams. nfltr tcp-connect used to open a new TCP connection, a full TLS handshake and a WebSocket upgrade to the relay for every local connection; a relay profile put 49 % of its CPU in TLS server handshakes and 21 % in three log lines per tunnel. Now it keeps one pooled TLS connection to the relay and opens each tunnel as an HTTP/2 stream, sending its first bytes with the open. Setting up a tunnel fell from 743 to 65 microseconds and from 958 to 145 allocations. Per-tunnel log lines moved to debug; metrics still count and time every tunnel. The WebSocket path remains for relays and proxies that speak only HTTP/1.1.
  • One redial schedule. Every client waited a fixed 5 s before redialling a relay that went away. Now they share one schedule: 250 ms, doubling, capped at 1 s during the first minute of an outage, with jitter. The same loop had a second bug: when a reconnect gave up, it ran a connection cycle on a stream that did not exist and crashed.
  • Accepts never wait on one client. Each TLS ClientHello on the encrypted port is read in its own goroutine.
  • A smaller binary (below).

Docker after round 1: new connections 1,338 to 2,810 requests per second, reconnect 5.4 to 2.7 s, the public URL error-free. Bulk throughput did not move, and 1 MiB responses got slower (389 to 330), which round 2 explained.

The binary: 107 MiB to 76

nfltr is one binary: the CLI, the relay (for nfltr relay serve and the local relay that nfltr orch starts), WireGuard, peer-to-peer file transfer and the docs site. Most of its weight was relay backends that only our hosted relay uses. We measured each one by removing it from an otherwise unchanged build:

CutSaved (stripped linux/amd64)
Embedded policy engine, moved to the hosted relay's build12.2 MB
Postgres driver, same3.8 MB
Message-broker client, same3.5 MB
Object-storage client, same3.4 MB
Demo recordings stored compressed3.83 MB
Vendored browser scripts stored compressed, one copy instead of two0.33 MB
Build paths trimmed0.27 MB
Total88.4 MB to 61.2 MB (−30.8 %)
The release build, stripped. In the harness's unstripped build the same change shows as 107.1 to 75.6 MiB. No user-facing feature was removed: the built-in backends (in-memory broker and cache, SQLite, file storage) stay, and a setting that asks for a backend this binary does not contain fails at startup with a message naming the setting, never a silent fallback.

A size budget now fails the build when the release binary grows past 62,000,000 bytes, printing the modules that grew. With Go 1.27 (below) the release binary is 60.85 MB.

Round 2: bulk throughput, the HTTP path, many agents

Bulk. Measured with real binaries on one host, each relay and exposer limited to one Go processor, interleaved:

Load (6 s)BeforeRound 2Change
Upload, 1 stream (Gbit/s)6.27 / 6.899.69 / 9.56+45 %
Upload, 8 streams4.82 / 6.399.12 / 10.08+70 %
Download, 1 stream6.39 / 6.958.71 / 9.18+34 %
Download, 8 streams6.17 / 6.768.53 / 8.60+32 %
1 MiB HTTP, c=50 (req/s)675 / 692903 / 989+38 %
Two interleaved rounds each, macOS loopback. Absolute numbers are not comparable with Docker; the ratios are what count. With 8 streams the old build was slower than with 1, as in Docker; round 2 is level.
  • A buffer pool that does not clear. A pool with a size class for every power of two from 1 KiB to 1 MiB, which does not zero what it hands out (gRPC overwrites every byte), used by both relay and agent. Together with the next item, +5-10 % per relay CPU-second on macOS, where clearing was a smaller share than on Linux.
  • Tunnel frames sent without copying their data. The frame header is encoded and the payload goes to the transport as is, byte-identical to what the protobuf library would produce, which a test checks.
  • 256 KiB HTTP/2 frames between tcp-connect and the relay. Every TLS record is a write system call, and every body read on the relay cost a flow-control update write. Those flushes were 36 % of relay CPU with 16 KiB frames, 15 % after.
  • Several streams per agent. The relay offers an agent 4 tunnel streams in a response header; an agent that understands it opens them, and the relay hands new sessions to them in turn. No message format changed: an older agent ignores the header and keeps one stream. +7-8 % with 8 sessions. The new registry also fixed a reconnect race in which an agent's old stream ending removed the agent although its new stream had already registered.
  • One WebSocket implementation instead of two hand-written ones, masking 8 bytes at a time instead of one: masking went from 1.6 to 9.9 GB/s, and a 32 KiB client write from 40 KiB allocated to none.

The HTTP path. Besides the configuration fix above:

  • No loopback hop. A request for an agent connected to the same relay process went through a reverse proxy to the relay's own dispatcher port: a second HTTP parse, two more socket copies and a loopback connection per request. At 50 connections on macOS that exhausted ephemeral ports and 6.5 % of requests failed. It is now handed over in process, with the same header handling as the proxy, which a test compares.
  • One backend client on the agent, shared by both HTTP modes and the encrypted sessions: up to 1,024 idle connections per host and a 4 KiB read buffer instead of 256 KiB, which only pinned memory (1,000 idle connections took 73 MiB with a 64 KiB buffer, about 4 MiB with the default). Connections to the service in the 32-session test fell from 226 to at most 32, and agent CPU per request from 0.33 to 0.02 ms.
  • 1 MiB chunks instead of 8 MiB for large HTTP bodies. For 64 MiB downloads at 16 at once, relay peak memory went from 397 to 121 MiB and the agent's from 1,104 to 187 MiB, time to first byte from 90 to 36 ms, throughput unchanged.
Real binaries, 1 KiBBeforeAfter
Inspectable URL, c=50 keep-alive (req/s)3,048 / 2,60016,416 / 12,279
Inspectable URL, relay CPU per request0.28 ms0.05 ms
Encrypted URL, c=50 keep-alive (req/s)2,339 with 20-24 % errors45,930 / 45,920
Encrypted URL, c=1000 keep-alive (req/s)289 / 489, mostly dial errors29,167 / 35,881
macOS loopback, relay and agent at one Go processor each, two interleaved rounds. The "before" errors on the encrypted path are the agent's dials to its service running out of ports.

Many agents. Besides removing the 100-agent default, we cleaned up state that outlived agents. About 39 metric series carried an agent label but only one was ever deleted, so series grew with every agent ever seen; now all of them are dropped when the agent disconnects, and a test fails if a new labelled metric is registered any other way. A cache of per-agent request producers held 100 entries, so above 100 busy agents it closed a producer on almost every request, sometimes while another request was still sending on it; the cache now holds connected agents only, and a producer in use is never closed. In process, 500 agents now register in 90 ms; before, 465 of 500 got in.

Round 3: fewer system calls, less memory per agent, faster reconnect

  • One write per TLS batch. Go's TLS layer hands each 16 KiB record to the connection as a separate write. A thin wrapper now collects all records of one write and sends them in one system call, on the gRPC leg of relay and agent. That made a larger gRPC write buffer worth having (32 to 256 KiB on both ends).
  • A backlog goes out in one write. A session writer that fell behind writes everything queued, up to 256 KiB, with one vectored write.
  • No pipes on the tunnel path. Two in-memory pipes, each with a goroutine pumping it, were replaced by reading the session queue directly: two goroutine wake-ups fewer per request and per response on the agent.
  • Extra streams only when needed. Four eager streams cost the relay about 122 KiB and 9 goroutines per idle agent. An agent now opens its extra streams only once two sessions run at once.
  • Less per agent. Four housekeeping goroutines per agent session became one; heartbeats are sent from the stream's own sender; the relay dropped gRPC's 32 KiB read buffer per connection, which under TLS only duplicated what the TLS layer already holds.
  • Faster reconnect. The relay slept 0.7-1.2 s at startup before listening, waiting before even its first attempt to set up its internal message broker; it now waits only between failed attempts (start to listening: 1.1 s to 45 ms, which also shortens every deploy of the hosted relay). The redial schedule kept doubling while the relay was coming back; now the first 8 redials stay at 250 ms. gRPC's own reconnect backoff could swallow a redial; clients now reset it before each attempt. And a connection that arrived while its agent was reconnecting polled every 500 ms; it now wakes when the agent registers.

Same-window Docker comparison, three interleaved rounds:

Testnfltr beforenfltr afterChangefrp
iperf3, 1 stream (Gbit/s)6.287.69+22 %7.49
iperf3, 8 streams (Gbit/s)5.737.35+28 %9.90
1 KiB, c=50, keep-alive (req/s)52,65151,567noise73,454
1 MiB, c=50, keep-alive (req/s)756906+20 %1,662
Encrypted URL, 1 KiB, c=50 (req/s)35,94743,750+22 %70,316 (frp-https)
2026-10-06, 19:34-19:58, relay and exposer 1 CPU each, medians of the rounds' medians. The host was quieter than in the round 2 run, so every tool's absolute numbers are higher; compare within the table only. The target container was not yet CPU-limited, which flattered frp-https (see round 4).

nfltr now beat frp at 1 stream, and 8 streams no longer fell below 1 by more than noise. Memory per idle agent, in process: from 24 goroutines and 237 KiB of heap and stacks to 10 goroutines and 105-109 KiB at 100 and 500 agents. In Docker, with 100 agents, relay memory per agent went from 366 to 308 KiB (−16 %; the Docker figure includes the Go heap's headroom). Reconnect, with real binaries on one host and the relay restarted at once: from 2.63 s to 0.44 s between stopping the relay and the first request through the tunnel.

Round 4: a fair benchmark, a faster encrypted path, more connections when busy

Round 3 left the encrypted public URL at 43.8k requests per second against frp-https's 70.3k, with nfltr's agent at 98 % of its CPU. Then we noticed that frp-https ended TLS in the target container, which had no CPU limit. So the first change was to the benchmark: the target got the same limit, and every row got CPU-seconds per 1,000 requests.

  • A stream path for encrypted HTTP. After ending TLS, the agent used to run each session through a full HTTP server and reverse proxy. Now an HTTP/1 session to an http:// service gets its own connection to the service: requests are parsed only to set the Host header the service expects, everything else goes on as the client sent it, and responses are copied back as bytes. Upgrades (WebSocket) become a plain byte copy. A session whose first request is also its last uses the pooled connections instead, so clients without keep-alive do not pay a dial each. The full HTTP path stays for features that must read requests (password protection), for https:// services and for HTTP/2 clients.
  • nfltr http --mode passthrough, a new value of an existing flag. The relay routes by server name and the agent hands the TLS bytes on to your service, which holds the certificate. Neither the relay nor the agent can read the traffic. This is the mode that compares directly with frp-https.
  • More connections for a busy agent. All of an agent's streams shared one gRPC connection, and so one writer and one reader goroutine at each end. Each tunnel response now carries a random 128-bit token; an agent that is serving two or more sessions at once opens up to three more connections with its own credentials and that token, and the relay adds them to the agent's rotation only if the token names a live stream of the same authenticated agent. Extra connections never count as agents, end with the primary one, and are retired after 30 s without a session. tcp-connect likewise spreads tunnels over up to 4 HTTP/2 connections. Each half needed the other: spreading only one side moved the serial stage to the other.
Encrypted URL, 1 KiBnfltr beforenfltr afternfltr passthroughfrp-https
c=1, keep-alive (req/s)10,13712,60515,46918,973
c=50, keep-alive (req/s)43,33658,766 (+36 %)66,79849,270
c=50, keep-alive, CPU-s per 1k0.0380.0290.0250.045
c=50, new connection each (req/s)7,0955,18513,2954,560
c=50, new connection each, CPU-s per 1k0.1960.2830.1440.399
Five interleaved Docker rounds, relay, exposer and target at 1 CPU each, medians. Another benchmark used 2-4 of the host's cores: frp-https at c=50 ranged from 43.6k to 69.5k across rounds (67-69k in its two quiet rounds).

With the target under a limit, frp-https spends more CPU per request than nfltr. Agent CPU per keep-alive request fell 45 %. Without keep-alive the medians put the new build below the old one, but the paired rounds disagreed in sign (−34 % to +17 %), and back-to-back runs put it at or above the old one. Both builds spend most of that test in the TLS handshake (key exchange arithmetic alone is about a third of the agent's samples), so we claim neither a gain nor a loss there. For reference, in our earlier local measurements a plain Go HTTPS server on one CPU did 4.1k requests per second without keep-alive, and the agent 4.2-7.0k: at or above that floor.

The extra connections, measured with real binaries where each process could use 4 CPUs (which exposes serial stages): 8-stream downloads from 13.03 to 17.25 Gbit/s (+32 %) and 1 MiB HTTP from 1,395 to 1,919 requests per second (+38 %), with relay CPU per unit of work unchanged or better.

Round 5: a thin data plane, less memory per agent, quieter logs

After round 4, nfltr still trailed frp on 8 streams and on 1 MiB responses. Every tunnel byte between agent and relay was a protobuf message on a gRPC stream: HTTP/2 framing and flow control, gRPC's per-stream write quota, and one writer goroutine per connection behind a hand-off from each stream's sender. frp moves its bytes as multiplexed frames over one TLS connection. A second limit showed once that was out of the way: each agent stream had one receive loop for all its sessions, and it waited whenever one session's 1 MiB queue was full. With 50 concurrent 1 MiB downloads every session waited behind the fullest one, and neither the relay (70 % of its core) nor the agent (46 %) was busy. That is the head-of-line problem round 1 bounded but could not remove.

  • Data connections. The extra connections a busy agent opens (round 4) become plain TLS connections to the relay's existing port that carry the same session frames, length-prefixed, without gRPC or HTTP/2. The relay tells them from gRPC by their first four bytes after TLS, so no new port or listener is needed. Senders collect frames in a pooled 256 KiB buffer and write everything queued in one TLS batch and one system call; an idle connection holds no buffer. The agent's first stream stays on gRPC: it carries authentication, share declarations and heartbeats, and an idle agent, or one serving one session at a time, never opens a data connection.
  • Per-session credit. Each side announces a window, the most it queues for one session (4 MiB on the relay, 16 MiB on the agent). A sender takes credit for each chunk of at most 64 KiB before queuing it; a receiver returns credit as bytes leave the session's queue, in one window frame per half window, so a request and response of a few KiB sends none. A receive loop never waits on a full queue any more: a client or service that stops reading stalls only its own session, and memory stays bounded by the window per session.
  • Negotiated, with a fallback. The relay offers the data plane in a response header, and only when its listener takes data connections. An agent that cannot open one (an HTTPS proxy in its environment, or a load balancer that is not a plain TCP pass-through) uses round 4's gRPC connections for the rest of that tunnel's life. Old agents with new relays and new agents with old relays work as before; no message format changed.
  • Bound to the agent by a token. A data connection's only credential is round 4's 128-bit token, sent only to the authenticated agent and valid while the agent's first stream lives. At most three data connections per agent, they never count as agents, and they end with the first stream. The share-policy rules above still hold: a data connection enforces what its agent's first stream declared, a protected share reaches it only when the agent declared that share, and the check is repeated there. A data connection opens no session itself, so it reaches nothing the existing ownership checks did not already allow.
Testnfltr beforenfltr afterChangefrp
iperf3, 1 stream (Gbit/s)8.119.28+14 %7.45
iperf3, 8 streams (Gbit/s)7.3710.54+43 %9.74
1 KiB, c=1, keep-alive (req/s)14,87214,948+1 %21,529
1 KiB, c=50, keep-alive (req/s)53,75458,248+8 %73,408
1 KiB, c=50, new connection each (req/s)11,45711,4860 %6,957
1 MiB, c=50, keep-alive (req/s)8631,052+22 %1,623
Three interleaved Docker rounds, relay, exposer and target at 1 CPU each, medians. The encrypted public URL did not change (67,110 against 67,054 at 50 keep-alive connections): an nfltr http agent serves one session per connection, which stays on its first stream unless several overlap.

8 streams are faster than 1 again, and both are ahead of frp in the same window. The 1 MiB gap narrowed but did not close: the relay now spends 44 % of its CPU in the HTTP/2 writer towards tcp-connect, one TLS record and one write system call per frame, and 37 % reading the data connections. That leg is next.

Less memory per agent. The v1.0.675 run measured 352 KiB of relay memory per idle nfltr http agent at 500 agents, against 132 for frp. A profile of one such agent found 12 goroutines and about 80 KiB of live heap, plus gRPC write buffers left in a shared pool after the burst of registrations; Go keeps about twice the live heap. The changes:

  • One share client. Issuing a share URL kept its HTTPS connection to the relay open for 90 s, two relay goroutines and about 16 KiB per agent at exactly the time the benchmark measures. There were three copies of that client; there is now one, which closes its connection. It is built after the pinned certificate of a relay you run yourself is installed, which also fixed nfltr http failing to get its share URL from nfltr relay serve --tls-self-signed.
  • One goroutine fewer per session: the session's housekeeping runs on the goroutine that was only waiting for the session to end.
  • A receive queue that grows with use instead of a 1,000-slot queue allocated at connect: an idle session holds none of its 9.5 KiB.
  • Relay write batches of 64 KiB instead of 256. gRPC takes a write buffer from a shared pool for every flush, heartbeats included; right after 500 agents registered the pool held 380 buffers, 97 MiB. With bulk bytes on the data plane the smaller batch costs about 2 % on one iperf stream, within noise elsewhere.
  • A memory limit by default. Without GOMEMLIMIT in its environment, the relay now sets a soft limit at 90 % of the memory limit of its container or service unit, so it collects harder near the limit instead of being killed. An explicit setting wins; without a limit nothing changes.

With real binaries on one host, 500 agents: relay memory per agent from 231-281 to 154-188 KiB, 12 goroutines to 9. In the Docker run on v1.0.676: 173 KiB per idle tunnel at 500, from 352 in the v1.0.675 run, against frp's 131.

Registration time is the agent starting, not the relay. 500 agents still take longer to register than 500 frp clients. A CPU profile of the relay while 500 nfltr http agents registered at once found 0.74 s of relay CPU for all of them, about 1.5 ms each, with no step blocked. The agent processes spent 24.7 s, about 50 ms each: 96 % of the CPU of registering a fleet is the agents, mostly starting the process (package initialisation, and two garbage collections during it). In the benchmark 500 agents start in one container with 4 CPUs, and 500 × 50 ms on 4 CPUs is about 6.2 s, close to the 6.3 s the v1.0.675 run measured. frp's client is a small dedicated binary. So nothing on the relay side was changed for speed; round 6 is about how fast nfltr starts.

Quieter production logs. The hosted relay had been logging at debug: the runtime configuration it reloads set the level to debug over its start-up flag. Each agent-to-agent round trip wrote about 8 lines, so one canary run of about 380 round trips wrote about 2,000 lines in a minute or two and showed as a CPU spike; an idle relay wrote about 26 lines a minute. Production now runs at info, the relay's default level is info, and per-request success lines moved to debug: a round trip logs nothing at info (4 lines before), and an idle relay with two agents logs nothing (4 lines a minute before). Errors, refusals and agent connects and disconnects stay at info and above. The relay host's watchdog still checks health every minute but checks certificates hourly, so a failing certificate keeps the check failed and is logged once an hour instead of every minute. The end-to-end canary against the hosted relay runs hourly.

Round 6: the client leg, and a faster start

Round 6 ran as two lanes in parallel: one on the leg between nfltr tcp-connect and the relay, one on how long an nfltr process takes to start.

The client leg on the data plane. After round 5, 44 % of the relay's CPU at 1 MiB went to the HTTP/2 writer towards tcp-connect: each response write passed through two goroutines, and each HTTP/2 data frame was its own TLS record and its own write system call. Round 6 gives that leg the same framing and per-session credit as the agent leg.

  • One upgraded connection. tcp-connect opens one TLS connection to the relay's HTTP port and turns it into a data connection with an HTTP/1.1 upgrade, the mechanism WebSockets use. Every tunnel is then a session on it, and its first bytes go out right after the open, with no round trip. The upgrade request carries the API key and is refused before any upgrade without a valid one.
  • The same checks for every tunnel. Each session open runs the one connection check the HTTP/2 and WebSocket paths already used: the credential again (a key revoked since the upgrade opens nothing more), the owner-only rule from who may reach an agent, and the port. A refused session is closed exactly like one to an agent that does not exist.
  • Writes from the goroutine that has the bytes, with one yield. On a data connection, on both legs, a sender that finds no write in progress becomes the writer, yields the CPU once so others about to send can add their frames, and writes everything buffered in one system call; the others only buffer. On a 1-CPU relay, without the yield, it wrote one record per response. On an idle CPU the yield returns at once.
  • A fallback. A relay that predates this, a proxy that refuses upgrades, or a relay without TLS keeps the old paths (HTTP/2 streams, then WebSockets). The cost on an old relay is one refused request, once per process. No message format changed.
Same window1 KiB, c=50, keep-alive1 KiB, c=50, new connection each1 MiB, c=501 KiB, c=1,000, keep-alive
Relay CPU per request, before (µs)29.5101.21,73023.1
Relay CPU per request, round 6 (µs)11.136.41,4039.2
Relay CPU per request, frp (µs)30.4176.71,11430.0
Requests per second, before / round 6 / frp19,006 / 32,150 / 28,1155,769 / 8,679 / 3,644462 / 629 / 83938,044 / 60,360 / 32,477
Three interleaved Docker rounds, relay, exposer and target at 1 CPU each, medians. Other work loaded the host in this window (frp did 28k small requests per second here against 73k in round 5's), so compare within the table; relay CPU per request moves less with host load than throughput does.

Relay CPU per small request fell by more than half, to about a third of frp's. At 1 MiB the relay is still CPU-bound, at 1.26× frp's CPU per byte: each received frame is a fresh allocation, and an agent's backlog is joined into one buffer before it is written.

A faster start. Go runs the initialisation of every package linked into a binary in every process, whatever the subcommand, so every agent paid for the relay's, the hub's and the docs site's start-up work. A profile of nfltr version found regular expressions compiled at start-up (13 % of its CPU), about 96 metrics registered one at a time, each described on its own goroutine (6 %, and most of the process's context switches), page templates parsed whether or not a page was ever served, and a check of the command list that only a test needs. All of it now happens on first use or in tests; no metric, pattern or command changed.

nfltr versionBeforeAfter
CPU time3.2 ms2.0-2.3 ms (about 30 % less)
Context switchesabout 450about 80
Allocations during initialisation22.8k (2.50 MB)9.3k (1.43 MB)
CPU for one nfltr http agent to register (median of 30)23.0 ms19.2 ms
Linux arm64 container on a quiet host: 100 runs of nfltr version, 30 agents. A test now fails when our own packages allocate more than a fixed budget while initialising, so new start-up work has to be lazy or justified.

The profile also showed where an agent's remaining 19 ms go: about 36 % is loading the system's certificate authorities, parsing some 150 root certificates once per process to verify the relay's certificate. frpc does not verify its server's certificate by default, so it does not pay this. We keep the cost on purpose: an agent that skipped verification could be talked into handing its API key and its traffic to anyone in the path. Cheaper ways that keep verification, such as a pinned certificate or a smaller trusted set given by the operator, need their own security decision and are not in this release.

Tried and rejected

Not everything that sounded right measured right. What we built or measured and did not keep:

TriedMeasuredWhy not kept
A lock around each gRPC sendCorrect, but each frame's read and send serialised; the queue was faster (about 400 against 515-713 MB/s)Throughput
Killing a session whose queue fillsBounds head-of-line blockingBreaks legitimate slow consumers (an upload to a slow disk, a download to a slow client); a full queue applies backpressure instead
gRPC's automatic flow-control windows64-77 MB/s at 40 ms, against 224-250 for fixed larger windowsSlower
Per-session credit carried in existing message fieldsNot builtA wire change without a schema; built in round 5 with its own framing instead
Larger gRPC write buffers alone (round 2)No gain while every TLS record was its own system callKept only after round 3's TLS batching made them pay
Read coalescing below TLS on the relayRelay CPU per Gbit/s unchanged within noiseAn idle agent would pin the buffer; removed
Small frames sent from the reader's own goroutineAll frames: iperf 1 stream −30 %, 1 MiB −15 %; frames up to 16 KiB only: within noiseNot worth the ordering subtlety
tcp-connect over 4 connections, alone (round 3)8 streams −10 % (6.4 against 7.1 Gbit/s)The agent's one connection stayed the serial stage; kept in round 4 once the agent side spread too
tcp-connect capped at 2 connections1 MiB: 877 against 1,027 req/s with 4Slower
TLS session resumption instead of HTTP/2 streams for tunnelsNot builtKeeps the TCP connect, a handshake round trip and the upgrade per tunnel
A separate multiplexer over one WebSocketNot builtA second framing layer and flow control where HTTP/2 already has both
Batching TLS records on the encrypted stream path1 MiB downloads 1,297-1,323 against 1,272-1,673 MB/s without; 38 % fewer allocations, 9 % more bytesNo faster
Keeping the agent-limit refusal retryable with a long backoffNot builtAgents would still loop without telling the operator
Data connections without credit (round 5)8 streams +25 %, but 1 MiB unchanged (787 against 771 req/s)Every session of a connection still waited behind the fullest queue
Credit with a 1 MiB window1 MiB +9 %, but one iperf stream −9 % to −32 %The relay sat waiting for credit; the final windows are 4 and 16 MiB
New sessions rotated over all streams, the gRPC one included8 streams 7.94 against 9.32 Gbit/s for data connections firstSlower
HTTP/2 extended CONNECT, or raw bytes on a gRPC streamNot builtKeeps HTTP/2 flow control and the single writer per connection, which round 5 removes
Authenticating data connections with the API key againNot builtA second authentication path to keep in step; the token proves the authenticated agent handed it over, and the connection dies with that agent's stream
Opening an nfltr http agent's encrypted stream on first useWould save about 27 KiB per idle agentDeferred: needs a new relay-to-agent signal and adds a round trip to the first encrypted connection
A lower GC target on the relayNot builtTrades relay CPU for memory on every relay; the default memory limit bounds the headroom where it matters, near the limit
Caching key lookups, batching registration writesThe relay spends about 1.5 ms of CPU per registering agentNot where registration time goes (round 5)
Credit returned at a quarter window instead of half (round 6)No difference outside the noiseReverted
The client leg on the gRPC port, or a new protocol name in the TLS handshake (round 6)Not builttcp-connect knows only the relay's HTTP address, and every relay TLS setup would need the new name; an HTTP upgrade does the same on the port it already uses
Skipping certificate verification at agent start, as frpc does by default (round 6)About a third of an agent's start-up CPUAn agent must know it is talking to its relay before it sends its key
Compressing the binary (UPX)Not measuredSlower start, antivirus false positives, and it breaks macOS code signing

Reliability and security found along the way

Reading every path a byte takes also means reading every path a request takes. Three problems were not about speed. The hosted relay at nfltr.xyz runs v1.0.678, which has every relay fix in this post; the hub fixes ship in the same release of the CLI.

We can't say whether anyone used these paths before the fix: the relay's logs from before then are no longer available, because they lived only on the VM we retired when we moved regions. We are not claiming there was no misuse; we simply have no record either way.

Protected share links on encrypted URLs

A share link can carry an access policy: a password, a bearer token, a required header, or an IP allowlist. On a relay that passes encrypted share URLs through to the agent, which the hosted relay does, that policy was not enforced on the encrypted path. The relay checked policies only on the path where it ends TLS itself; on the encrypted path it looked up which agent the link belonged to and dropped the policy, and the agent, the only hop that can read the requests, never knew it. So a protected link behaved like an unprotected one. The same audit found that --basic-auth on the agent itself was not checked on encrypted URLs either.

The rule now is that a share's policy applies on every path that serves the share, and protected shares stay end-to-end encrypted:

  • the relay checks the IP allowlist when it accepts the connection, before anything reaches the agent;
  • the agent, which received the policy when the share was issued, checks credentials on every request in constant time and never logs them; a failed check never reaches your service;
  • the agent tells the relay which shares it enforces, and the relay refuses to forward a protected share to an agent that has not declared it, for example an older agent, so a missing check fails closed;
  • nfltr http --mode passthrough, where nobody reads the requests, refuses share policies at startup instead of serving a link that cannot be protected.

Who may reach an agent

The share-link fix raised a wider question, so we wrote the rule down and checked every path against it: an agent can be reached without a login only through a share link it created; every other path needs the agent's owner. Several paths fell short. Some checked only that the caller was authenticated, or did not check the caller at all, rather than that the caller owned the agent; one host-name form reached an agent directly with no share involved; one older route was open by documented design; and messages between agents did not reliably check that both belonged to the same account. One handler turned out to be unreachable in the production chain and was deleted rather than fixed. All of them now go through one ownership check, and a refusal looks exactly like an agent that does not exist, so agents of other accounts cannot be discovered by probing.

A completion that lost to a timeout

This one was in the hub, not the relay, and the relay work exposed it. When a hub's agent finishes, its worker reports the result and exits. A watch in the hub, after noticing a worker had gone, waited 2 s for the outcome and then recorded the agent as lost with its node. During a relay restart, the result could already be recorded while the wait was still resuming its stream, and "lost" won anyway. Faster redials and multi-stream sessions shifted the timing into that window. It showed up once in three runs of our hub fault gate. The rule now: an outcome the runtime has already recorded for an agent's task wins over a later "worker gone". A unit test reproduces the exact order; the relay-kill soak ran 96 of 96 goals clean after the fix (89 of 89 before; the race is rare, which is why the unit test is the guard).

Smaller reliability fixes

  • An agent turned away at the agent limit said "Connected" and then retried forever; it now exits once with the reason.
  • A reconnect loop that gave up crashed on a stream it never had.
  • One silent client could stall accepts on the encrypted port for 5 s.
  • A channel was closed under live senders, and the panic hidden.
  • Concurrent agents could pass the agent limit together; admission is now atomic.
  • Producers could be closed while a request was sending on them.

Round 7: conformance suites

Our own tests check what we thought of. Round 7 ran the open-source suites other people wrote for the protocols the relay speaks, against a local relay, agent and client on an internal Docker network, plus a read-only TLS scan of the public endpoints, the same scan any visitor can run. After every suite the harness checks that every process is still up and every share still answers; that check is how the most serious bug below surfaced.

SuiteWhat it checksResult
Autobahn TestSuiteWebSocket (RFC 6455), the relay's server and the CLI's client112 and 107 failed of 247 before; 0 after
h2specHTTP/2 (146 cases), the relay's HTTPS front, a share, and its gRPC portThe same results as plain Go servers built with the same versions, apart from one timing case explained in the report; before the fix below, it crashed the relay
smuggler, http2smuglHTTP request smuggling and desync on both kinds of public URL134 of 134 mutations handled, none potentially smuggled; no smuggled HTTP/2 header parsed
gRPC interop testsgRPC through an nfltr tunnel, plaintext and TLS end to end28 of 28
testssl.shTLS on both relay ports, the agent's encrypted endpoint and the public endpointsTLS 1.2 and 1.3 only, authenticated encryption with forward secrecy, hybrid post-quantum key exchange on the relay; two old suites on the public endpoint, fixed
govulncheckKnown vulnerabilities reachable from our codeOne reachable, fixed; one allow-listed until a dependency upgrade (below)
Go fuzzing14 targets over the parsers of untrusted bytes: WebSocket frames, data-plane frames and handshakes, TLS server names, share codesOne bug found
SchemathesisThe REST API against its published schema, with and without credentialsNo server error; no secured operation answered without a credential; two bugs and some schema drift
Client TLS refusalsThe CLI against expired, self-signed, wrong-host, untrusted and weak-key certificates, on every path that dials the relayNone accepted without a pin or a trusted authority; the refusals were not actionable, fixed

Seven bugs found and fixed, each with a test that fails without the fix:

  1. A relay crash over HTTP/2, without a login. When a request to a share link outlived its proxy timeout while its body was still being read, a late read touched HTTP/2 state that had already been released, and the relay process crashed. Any client of any share could trigger it; it needed no account. Over HTTP/1 the same late read could reset the deadline of the next request on the connection. The fix ends that work when the handler returns. It shipped in v1.0.677, and the hosted relay has run it since.
  2. WebSocket framing. A ping between the fragments of a message swallowed the tunnel bytes in it; frames that RFC 6455 forbids were accepted; every violation closed with the "normal" code; and a fragmented message had no total size limit. Messages are now assembled and checked in one place for the relay and the CLI, violations close with the right code, and a message is capped at 1 MiB.
  3. Zero bytes the client never sent. When a TLS ClientHello arrived cut short, the relay's server-name router passed on the whole zero-padded buffer instead of the bytes it had read, corrupting the stream of a slow or truncated client. A fuzz target found it on its first run.
  4. Old cipher suites in production. The certificate path the hosted relay uses offered two CBC-SHA1 suites in TLS 1.2 that the other path did not. Both paths now offer the same modern set.
  5. A reachable advisory in the tracing exporter, which could log endpoint URLs. Upgraded.
  6. Key management. Creating a key accepted agent IDs the key file cannot hold, which could stop a relay you run yourself from restarting, and deleting a key without naming the agent answered "not found". Both need the admin token; both are now refused with a clear error.
  7. Refusals with no way forward. A refused relay certificate printed the raw error. It now prints the certificate's fingerprint and what to do: pin it, add the authority, renew it, or use a name it holds.

Written up and not fixed in this release:

  • The agent's end-to-end encrypted TLS offers no post-quantum key exchange, though the relay's does. That is the path meant to stay confidential end to end, so it is the first on the list.
  • One advisory in a token library is reachable only through the optional Apache Pulsar broker with OAuth2, with tokens from the operator's own provider; it is allow-listed until the Pulsar client is upgraded.
  • The published REST schema has drifted from the code: errors are plain text where it promises JSON, and some routes are missing from it.
  • A data-plane peer may grant up to 4 GiB of credit in one window frame, and a data connection has no total credit limit across its sessions. Neither is a parser fault.
  • tcp-connect does not use the relay certificate pin saved by nfltr config set-relay, and its own pin flag takes a different format.

The last gates. Before the release we ran our fault soak, 30 minutes of relays, nodes and hubs killed and paused while a scripted stand-in for Claude Code does the agents' work, and the hub ladder, which walks a hub from readiness through spawn, steer, stop, reattach and nodes. The soak caught one more bug, in the hub rather than the relay: after a hub came back from being paused, its acknowledgement of a finished agent's result could be lost in flight, because the session cancelled the stream as soon as it had sent it. The node then kept that agent's process and task directory until its next reconnect, 6 minutes later. The session now waits, up to 5 s, for the worker to end the call after the acknowledgement.

How we tested it

Every fix in this post started with a test that failed on the old code. The kinds of test that now guard the relay:

  • Data integrity under concurrency: 8 sessions echoing, uploading and downloading on one agent, run under the race detector; a service that writes and closes must lose nothing; 100 encrypted requests on a connection each must all arrive whole, with goroutines back to baseline.
  • Head-of-line blocking: a 768 KiB burst to a service that never reads must not stall a sibling session; on a data connection, an unread session takes exactly its window while another session on the same connection still gets its data.
  • Connection setup: many tunnels on one relay connection; the HTTP/1.1 fallback losing nothing; one silent TLS client not blocking accepts.
  • Reconnect timing: an agent back within 0.8 s of the relay starting again, after downtimes of 0, 0.5 and 1 s (the old code took 1.05-2.54 s); the first redial after a reset reaching a relay that is back.
  • Budgets: exactly 9 relay goroutines and at most 120 KiB per idle agent, at 100 and 500 agents, for a plain agent and for an nfltr http agent with its share URL; a dotted settings read allocating under 1 KiB; the release binary under 62,000,000 bytes.
  • Interop: current agents with older relays and the reverse, for the multi-stream header, the extra-connection token and the data plane, plus an agent whose data connection fails, since agents and relays are upgraded at different times.
  • Authorization with two accounts: for each path that can reach an agent, the owner gets through and another account gets the same answer as for a missing agent; protected shares refuse missing and wrong credentials and never touch the service; extra connections and data connections cannot be bound by another agent, with an unknown token, past three per agent or after the agent's stream ended; a protected share reaches a data connection only when its agent declared the share.
  • Fault and soak runs: relays killed with SIGTERM and SIGKILL mid-transfer and mid-agent, nodes and hubs killed and paused past their heartbeats, checking that every agent's result arrives exactly once.

Each merge ran the race detector on the touched packages and every package that imports them, the relay package three times over (400-500 s), the cross-platform build for every shipped target, the dead-code and binary-size checks, and our $0 end-to-end hub run with a scripted stand-in for Claude Code.

Flaky tests were fixed, not retried. Three that surfaced during this work: a test whose cancel also stopped the relay it ran against (10 failures in 80 runs before, 0 in 300 after); a test that modelled a source cancelling by also disconnecting the target (2-5 in 500); and a guard that counted connections through one shared HTTP client, which could open extra ones while requests raced. In each case the product behaviour was right and the test modelled the wrong thing. The Docker smoke tests that reached agents by a host name the new ownership rule no longer routes were migrated to issue a share link and use it, as a user would.

Go 1.27

We moved to Go 1.27.1 during the same week. Measured on the same source, with the old and new toolchains:

MeasureGo 1.26.6Go 1.27.1
Release binary (stripped linux/amd64)61,264,034 B60,792,992 B (−0.8 %)
JSON unmarshal of hub task status and eventsbaseline−51 % and −55 % time
JSON marshal of the samebaseline−11 % and −17 % (not statistically clear)
Relay, TCP push, 1 / 8 streams (Gbit/s)6.00 / 6.035.83 / 5.85
Relay, 1 MiB HTTP, c=50 (req/s)652686
Relay, 1 KiB HTTP, c=50, keep-alive (req/s)29,81731,175
Relay rows: real binaries, one Go processor each, medians of 8 interleaved rounds on a shared host; round-to-round spread on the same binary was ±15-30 %, so the relay rows are within noise. JSON rows: 6 runs each.

The honest summary: Go 1.27 made the hub's JSON handling about twice as fast to decode and the binary a little smaller, and did nothing measurable for relay throughput. Its new goroutine-leak profile now backs two of our leak tests, which can assert exactly which goroutines are stuck forever without a settling delay.

A benchmark we almost got wrong

The first final run on October 8 showed nfltr 30-45 % slower than the run before it on bulk transfer (5.07 and 7.08 Gbit/s on one and eight streams, against frp's 7.26 and 9.95), with nothing in the code to explain it. The run's own profiling pass, a few minutes later, gave the same build 8.80 and 12.9. A same-window A/B of the same build and macOS's power log explained it: the host is a laptop, it was on battery while nfltr ran and on AC by the time frp ran, and on battery macOS turns on Low Power Mode, which made the same build 40-45 % slower and every tool with it (frp went from 7.45 to 4.00 Gbit/s on one stream on battery). The relay's profile had the same shape on both, only every operation cost about twice as much.

Had we compared that run's numbers, this post would report a regression that did not exist. The harness now records the power source and Low Power Mode before and after every tool, and the results flag a run where they changed or were on battery, so it gets rerun. The numbers below are from a rerun with every tool on AC power.

Where nfltr stands now

The final full run, on the released code (v1.0.678, after round 7), on 2026-10-08, with every tool on AC power. Relay, exposer and target at 1 CPU and 512 MiB each. Hosts differ between runs (see fairness), so the first table compares nfltr with frp as a ratio within each run; the two earlier columns are history:

nfltr ÷ frp, same runv1.0.675v1.0.676v1.0.678 (final)
iperf3, 1 stream1.241.251.13
iperf3, 8 streams0.671.091.26
1 KiB, c=1, keep-alive0.720.740.75
1 KiB, c=50, keep-alive0.860.780.96
1 KiB, c=50, new connection each1.721.652.10
1 MiB, c=500.550.650.83
Encrypted URL, c=50, keep-alive (against frp-https)1.050.940.94
Encrypted URL, c=50, new connection each (against frp-https)1.631.611.74
Relay memory per idle tunnel at 500 (lower is better)2.671.321.39
Time to register 500 tunnels (lower is better)2.101.772.28
Throughput rows: above 1, nfltr is faster. Single 10 s runs; a few percent either way is noise. The two scaling rows moved against nfltr between the last two runs mostly because frp's numbers did (500 tunnels in 2.6 s, then 2.15 s; nfltr 4.6, then 4.9), and each is one run.
Bulk TCP (Gbit/s)Encryption1 stream8 streams
nfltrTLS client to relay, TLS relay to agent8.4212.47
frpTLS frpc to frps only7.489.88
rathole, NoiseNoise1.871.84
rathole, plainnone34.1223.30
chiselSSH1.721.64
borenone39.8525.31
no tunnelnone102.60114.98
HTTP through the TCP tunnel1 KiB, c=1, keep-alive1 KiB, c=50, keep-alive1 KiB, c=50, new connection each1 MiB, c=50
nfltr16,279 (0.055)69,841 (0.027)15,165 (0.140)1,388 (1.87)
frp21,623 (0.042)73,057 (0.032)7,227 (0.253)1,674 (1.27)
rathole, Noise11,892 (0.057)50,635 (0.033)2,856 (0.472)213 (9.15)
rathole, plain14,214 (0.042)108,660 (0.016)2,972 (0.304)failed (100 % errors)
chisel24,175 (0.034)79,994 (0.024)10,775 (0.151)481 (2.22)
bore14,099 (0.043)104,662 (0.016)3,176 (0.383)107, 0.3 % errors (2.26)
Requests per second, CPU-seconds per 1,000 requests in brackets (relay, exposer, target and nfltr's tcp-connect). No errors in any other cell.
Public HTTP URL, 1 KiBc=1, keep-alivec=50, keep-alivec=50, new connection each
nfltr, encrypted end to end (default)12,979 (0.055)66,637 (0.025)8,270 (0.170)
nfltr, passthrough14,377 (0.052)72,010 (0.024)16,016 (0.122)
nfltr, inspectable (relay ends TLS)6,229 (0.145)11,705 (0.112)3,402 (0.327)
frp-https (TLS passed through)20,010 (0.045)70,675 (0.032)4,751 (0.376)
frp, plain HTTP14,686 (0.060)15,029 (0.146)11,954 (0.169)
Requests per second (CPU-seconds per 1,000 requests). rathole, chisel and bore have no public URL mode. No errors in any cell.
1,000 connections, one tunnelKeep-aliveNew connection each
nfltr, TCP tunnel132,637 req/s, p99 15.2 ms21,180, p99 91.3 ms
nfltr, encrypted URL97,410, p99 33.8 ms5,500, p99 286 ms
frp, TCP tunnel63,449, p99 51.2 msskipped: 25.6 % errors at 100 connections
frp, plain HTTP URL9,528, p99 482 ms5,868, p99 297 ms
rathole, plainfailed at 1 connection (100 % errors)
rathole, Noise353, p99 1,487 msfailed at 1 connection (100 % errors)
chisel79,514, p99 44.4 ms10,140, p99 165 ms
bore26,670, p99 0.1 ms1,243, p99 754 ms
frp-https was not part of the concurrency steps. Relay memory at 1,000 connections through the TCP tunnel: nfltr 44 MiB, frp 172, chisel 144; nothing was killed for memory. rathole's plain mode worked in this run's other small-request tests but failed every concurrency step. bore did 125.6k with keep-alive in the v1.0.676 run and 26.7k in this one; we did not investigate why.
Restart and scalenfltrfrpOthers
Relay restart to working tunnel (s)1.202.27chisel 1.14, rathole 0.67 (Noise); plain rathole did not come back within 120 s; bore exits
200 tunnels registered (s)2.31.7baseline: chisel 1.9, rathole 4.8, bore 0.6
500 tunnels registered (s)4.92.15not run
Relay memory with 500 idle tunnels (MiB)9171not run
Relay memory per idle tunnel at 500 (KiB)175 (was 352)126baseline, at 200: chisel 144-174, rathole 502-519, bore 18-20
Relay idle CPU with 500 tunnels0.9 %0.0 %not run
Client process memory per tunnel (MiB)8.53.0baseline: chisel 3.2, rathole 4.1, bore 0.3
Restart is one docker restart per tool per run, timed from the restart command to the first request through the tunnel. Relay memory is the cgroup's anonymous memory.
Footprintnfltrfrpratholechiselbore
Idle relay CPU0.1 %0.0 %0.0 %0.0 %0.1 %
Idle relay memory, RSS (MiB)43.019.53.7-4.311.91.7
Exposer under load, peak anonymous memory (MiB)42.0 (was 205.2)125.55.0-8.813.73.3
Binary (MiB)75.9 harness build (linux/arm64); about 58 release (stripped, linux/amd64)17.4 server, 14.1 client1.010.32.6
Load is the 1 KiB, c=50, keep-alive run through the TCP tunnel. Exposer memory is the cgroup's anonymous memory, sampled every second. nfltr is one binary that also contains the CLI, nfltr orch, WireGuard, peer-to-peer transfer and the docs site.

Where nfltr leads

  • Bulk throughput, on one stream and on eight. 8.42 Gbit/s on one stream and 12.47 on eight, against frp's 7.48 and 9.88, with both of nfltr's legs encrypted. On the baseline, eight streams ran slower than one, at a third of frp's rate. Only the unencrypted tools are faster (plain rathole and bore, 23-40 Gbit/s).
  • New connections. 15,165 requests per second through the TCP tunnel without keep-alive, 2.1× frp (7,227) and ahead of chisel (10,775), at 55 % of frp's CPU per request. At 1,000 connections, 21,180, where frp passed 1 % errors already at 100.
  • Many connections through one tunnel: 132.6k requests per second at 1,000 keep-alive connections, the most of any tool in this run (chisel 79.5k, frp 63.4k), with a p99 of 15.2 ms against frp's 51.2 and a quarter of frp's relay memory.
  • Encrypted public URLs without keep-alive: 8,270 against frp-https's 4,751, 1.74×, at under half its CPU per request.
  • Passthrough mode: 16,016 requests per second without keep-alive, 3.4× frp-https, and 72.0k with keep-alive, level with frp-https's 70.7k at three quarters of its CPU per request.
  • Reconnect: 1.20 s from relay restart to a working tunnel, level with chisel (1.14) and about twice as fast as frp (2.27). In this run rathole with Noise was faster (0.67); plain rathole did not come back, and bore exits.
  • The agent's memory under load: 42 MiB at its peak, against 205 at the start and 126 for frp's client.

Where it is level

  • Small requests with keep-alive through the TCP tunnel: 69.8k at 50 connections, 96 % of frp's 73.1k, at less CPU per request (0.027 against 0.032 CPU-seconds per 1,000). It was 78 % in the v1.0.676 run, before round 6. chisel (80.0k) and the unencrypted tools (over 100k) are faster.
  • Encrypted public URLs with keep-alive: 66.6k at 50 connections against frp-https's 70.7k, 94 %, using 22 % less CPU per request.
  • Idle CPU: 0.1 % or less for every tool.

Where it still trails, and why

  • Latency of a single request: 16.3k requests per second at one connection, 75 % of frp's 21.6k; chisel does 24.2k.
  • Large responses: 1,388 requests per second for 1 MiB, 83 % of frp's 1,674, at 1.9 CPU-seconds per 1,000 against 1.3. Rounds 5 and 6 took it from 55 % of frp's rate; the relay is now CPU-bound on a copy and an allocation per frame.
  • Registration: 500 tunnels take 4.9 s against frp's 2.15. That time is the agent processes starting, and about a third of an agent's registration CPU is loading the system's certificate authorities to verify the relay (round 6).
  • Memory per tunnel and per agent: 91 MiB of relay memory with 500 idle tunnels against frp's 71 (175 against 126 KiB per tunnel; it was 352 in the v1.0.675 run), and 8.5 MiB per agent process against 3.0.
  • Idle relay memory: 43 MiB resident against frp's 19.5, and far more than rathole, chisel or bore.
  • Size: about 58 MiB for the stripped release binary against 17, 10, 2.6 and 1 MiB.
  • The inspectable URL mode, where the relay ends TLS and reads HTTP, runs at about a sixth of the encrypted mode's rate. It does strictly more work per request, on the relay's one CPU.

The reasons are structural more than accidental:

  • An extra encryption leg and an extra process. nfltr's TCP tunnel encrypts the client-to-relay leg as well as the relay-to-agent leg, so the relay decrypts and re-encrypts every byte, and the client side is its own process (tcp-connect). frp's visitor leg is plaintext in this setup; bore and plain rathole encrypt nothing. That costs latency on every request and CPU on every byte.
  • Certificate verification at startup. Every agent verifies the relay's certificate against the system's authorities before it sends its key; frpc skips that by default. We keep it.
  • gRPC for every agent's first stream. Every agent holds a gRPC connection with its transport state, a stream with its goroutines, and an HTTP session for its control traffic. That is most of the per-tunnel memory left.
  • One binary for the whole product. The relay, the CLI, the orchestration hub, WireGuard and the docs site ship together, so you download one file for everything and pay for all of it in size, idle memory and start-up time.

What's next

  • Round 8: more conformance and security suites: ssh-audit for the SSH paths, WireGuard interoperability, STUN and TURN conformance, promtool on the metrics, OWASP ZAP against the web surfaces, MCP conformance for the hub's tools, gosec and semgrep on the code, and scanning the images we publish.
  • The gaps above: a copy and an allocation per frame for large responses, the cost of verifying the relay's certificate at start without giving up verification, memory per idle agent, and the round 7 items written up but not fixed, post-quantum key exchange on the agent's encrypted path first.
  • One nfltr process for many services (a new up command): one process that runs every service in a config file, so a machine with five services pays for one process start and one certificate check instead of five. It is designed and coming; it is not in this release.
  • Size: moving the relay into its own binary would save about 13 MB from the CLI, at the cost of a second download for people who run their own relay. That is a product decision we have not made.

We will rerun the same harness after each round and update the scoreboard.

Reproduce it

Our harness lives in our closed repository, but nothing in it is special. To run the same comparison with your own tools:

  1. Containers with hard limits. Relay, exposer and target at 1 CPU and 512 MiB each (Docker's cpus and mem_limit), a client container with more (we used 4 CPUs and 2 GiB), all on one internal network with no egress. Pin every tool's version and record it.
  2. A target with keep-alive. nginx serving a 1 KiB and a 1 MiB file, plus iperf3 -s, in the target container under its own limit. For the passthrough comparison, nginx also serves HTTPS with a certificate for the share host name.
  3. The loads. iperf3 -c <tunnel> -t 10 -P 1 and -P 8, three runs, median. oha -z 10s -c 1 and -c 50 against each file, and -c 50 --disable-keepalive. Then -c 1, 100 and 1000 through one tunnel, stopping at the first step over 1 % errors.
  4. CPU and memory from the cgroup, not from the tools: read cpu.stat's usage_usec before and after each run for relay, exposer, target and any client-side tunnel process, divide by successful requests for CPU-seconds per 1,000, and sample memory.stat's anon every second.
  5. Reconnect: docker restart the relay while the clients keep running, then time until a request through the tunnel succeeds.
  6. Tunnel count: start N client processes with one tunnel each against a fresh relay, time until all N answer, then read the relay's memory and CPU over 20 s idle.
  7. Write down the encryption of every leg for every tool, and give each tool's like-for-like variant its own row (frp's HTTPS virtual host for end-to-end encrypted public URLs).
  8. On a shared machine, interleave. Run old build, new build and the comparison tools in turn, several rounds, and compare only within a window.

For nfltr, the agent side is nfltr tcp <port> or nfltr http <port> (add --mode passthrough to compare with frp-https), the client side of a TCP tunnel is nfltr tcp-connect, and on nfltr relay serve pass --agent-rate-limit 0, or every HTTP test stops at 100 requests per second per agent. If you get numbers that differ from ours, tell us; that is how we found most of what is in this post.

← All posts