Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Performance

Four benchmark documents live in the repository, and this page condenses them. Every table below is reproduced from a committed measurement; the raw JSON, the version manifests and the diagnostics are in-tree so any of it can be re-derived without re-running anything.

DocumentThe question it answers
bench/thesis/RESULTS.mdWhat does a configuration change cost, against ingress-nginx? This is the project’s actual thesis.
bench/thesis/RESULTS-EC2.mdDoes that thesis survive on real Linux? Same opponent, EC2 and k0s, no VM
bench/RESULTS.mdRaw HTTP/1.1 forwarding throughput, against nginx
bench/engine/RESULTS.mdDoes the io_uring engine get under the syscall floor?
bench/PROFILE.mdWhere does a request actually go?

Read this first

Most of these are macOS Docker Desktop VM numbers, on Apple Silicon under a linuxkit guest. They are valid relative to each other under identical conditions; they are not Linux bare-metal absolutes and should not be quoted as such. Two sections are the exception and say so where they appear: the thesis re-run on EC2 and k0s, and the uring engine’s real-Linux check. Both carry their own caveat — a shared, burstable four-vCPU box with the load generator on it — which compresses ratios rather than inflating them.

The competitor is not understated on purpose. ingress-nginx does not reload for every change — endpoint updates go through its Lua balancer without touching nginx at all — and a report that measured only the changes which force a reload would be describing a system that does not exist. The endpoint-churn arm is measured and reported alongside the rest.

The losses are on this page. Idle-connection memory, the kubectl apply write path, upstream connection reuse, and an unexplained 9% gap at high concurrency all belong to the other side.

The thesis: what a configuration change costs

Both controllers installed into the same Docker Desktop cluster at the same time, separate namespaces, separate IngressClasses, one replica each. Load is never sent to both at once. ingress-nginx is the more generously provisioned of the two: it runs with no memory ceiling while ramjet-ingress runs under the 256Mi cap its own chart imposes.

Load reaches the pods through a NodePort on the node’s own bridge address — identically for both. kubectl port-forward was tried and rejected, because a port-forward is one multiplexed stream through a Go proxy on the host and becomes the bottleneck long before either contender does.

Two kinds of churn are measured, every two seconds for 100 seconds:

  • Ingress-spec churn adds a differently-named path to a churn Ingress, which forces a reload. Verified from ingress-nginx’s own log, not assumed.
  • Endpoint-only churn moves one running pod in and out of a Service’s selector, which goes through the Lua balancer without reloading. The mutation flips a label on a running pod rather than scaling, because scaling would have measured the scheduler’s latency and reported it as the controller’s.

Throughput and latency under churn (oha, c64, 2 runs per cell)

ContenderArmRPS (median)vs own baselinep50p99p99.9HTTP errors
ramjetbaseline104,8440.35 ms3.8 ms12.5 ms0
ramjetspec107,536+2.6%0.35 ms3.6 ms10.3 ms0
ramjetendpoint103,298-1.5%0.35 ms3.8 ms12.7 ms0
nginxbaseline87,8650.43 ms4.6 ms13.0 ms0
nginxspec78,368-10.8%0.45 ms5.4 ms18.3 ms1722
nginxendpoint64,305-26.8%0.57 ms6.7 ms21.1 ms0

Idle keep-alive connections that survived the window

ContenderArmHeldSurvivedLostConfig events applied
ramjetbaseline10010000
ramjetspec100100098
ramjetendpoint100100098
nginxbaseline10010000
nginxspec100010066
nginxendpoint10010000

This is the cleanest result in the whole report, because it does not depend on how fast the machine was. Under spec churn ingress-nginx ended every single run with 0 of 50 idle connections surviving — 0/100 across both rounds, and 0/50 again in each of the two contended replicate rounds. ramjet-ingress kept 100 of 100.

Controller cost of the churn window

ContenderArmPod CPU-secondsCPU per requestvs own baselinePod memory at end
ramjetbaseline300.9 s28.7 µs17.9 MiB
ramjetspec310.1 s28.8 µs+0%18.3 MiB
ramjetendpoint302.3 s29.3 µs+2%16.5 MiB
nginxbaseline402.8 s45.8 µs128.1 MiB
nginxspec395.5 s50.5 µs+10%115.7 MiB
nginxendpoint367.7 s57.2 µs+25%127.8 MiB

Recompiling and republishing a route table 49 times in 110 seconds cost the data plane nothing this benchmark can find. The CPU column is the version of the claim that survives, because it normalises out how fast the machine was.

Where ingress-nginx is fine: endpoint churn. It did not reload once — 0 reloads across every endpoint-churn run, confirmed from its own log — and it kept all 50 idle connections and served zero errors. Its Lua balancer does what it claims, and any characterisation of ingress-nginx as “reloads on every change” is wrong.

The +25% below has since been retracted. This run reported that ingress-nginx’s non-reloading path was the more expensive one for CPU — +25% against its own baseline, where the reloading path cost +10%. It was measured under the arm-ordering confound described in the next section, and when the EC2 run reversed the arm order the figure came back as +1%. The Lua balancer is close to free; treat the +25% as an artifact and not as a finding.

What benchmark 1 does not establish

The -26.8% endpoint-churn throughput figure is not solid. Arms always ran in the order baseline, spec, endpoint within a round, so the endpoint arm was always last and always held whatever drift the machine had accumulated. Rounds 3 and 4 were run specifically to test that, with the arm order reversed — and they were invalidated by the docker daemon: another agent started their own six-container proxy benchmark partway through, and throughput for both contenders fell to roughly a quarter, varying between 27k and 96k rps within a single round.

So the ordering confound on the throughput number is unresolved. The CPU-per-request figure is the version that survives; the connection-survival and reload counts reproduced identically in all four rounds regardless of contention.

The contended rounds are kept in the repository rather than discarded, with their own warning label: they report ramjet-ingress losing 62% of its throughput to spec churn, from a contender whose CPU per request did not move at all under the same churn on a quiet machine. That number is the other agent’s benchmark, not this one’s.

On real Linux, same opponent

Everything above is VM numbers, so the thesis was re-run against the same kubernetes/ingress-nginx 4.15.1 on an EC2 t3.xlarge running k0s v1.36.3 — real Linux, no VM in the path, three controllers standing in one cluster. Full report: bench/thesis/RESULTS-EC2.md.

The thesis transfers, at the same magnitudes.

Under Ingress-spec churnramjetingress-nginx
Idle keep-alive connections surviving100 / 1000 / 100
HTTP errors01,829
CPU per request vs own baseline+1%+12%

Steady-state forwarding on the same box, median of three interleaved 30-second runs at c64, with the backend Service’s own NodePort as a no-proxy baseline:

ContenderRPS% of baselineCPU per requestMemory
ramjet (uring)12,35555.6%69 µs15.2 MiB
ramjet (hyper)10,85848.8%100 µs10.5 MiB
ingress-nginx7,97135.9%187 µs67.6 MiB

Propagation of a new Ingress is 372–384 ms at the median against 2,762 ms, and the spread matters more than the median: twenty ramjet trials spanned 370–426 ms with no slow mode, while ingress-nginx’s ten spanned 398–2,813 ms in three clusters set by its --sync-rate-limit default of 0.3.

Two claims move, and both move ingress-nginx’s way. Reversing the arm order resolved the confound flagged above: endpoint-only churn costs it +1% CPU per request, not +25%, and −6.3% throughput rather than −26.8% — the same contention tax ramjet-ingress paid in the same arm. And the kubectl apply write path is now a draw, 191 ms against 189, with ingress-nginx still doing nginx -t validation through an admission webhook in that time.

Same caveat as the uring section below, and it is load-bearing. The load generator, the proxies and the upstreams all share four burstable vCPUs, so the proxy is never the sole bottleneck and every ratio on this page is compressed toward 1. Steal stayed under 1.1%. The connection-survival result is the one that does not care: 100 against 0, in both environments.

Propagation latency

kubectl apply to the first request the data plane answers correctly, polled every 20 ms, no other load running. Ten trials of each shape per contender, interleaved with the order flipping every trial. kubectl runs inside the same container as the poller, so “applied” and “served” are two readings of one clock.

ContenderChangeTrialsMedianp95MinMaxMedian kubectl apply
ramjetnew Ingress10363 ms566324566159
ramjetbackend swap10354 ms556322556150
nginxnew Ingress101,151 ms3,6383023,638138
nginxbackend swap10459 ms3,6573843,657188

~3x faster at the median and ~6x at p95, and the more useful half of that is the spread: ten new-Ingress trials from 324 to 566 ms, against 302 to 3,638 ms.

ingress-nginx’s slow trials cluster just under 3.5 seconds and alternate with fast ones. That shape is a rate limiter, not a queue, and the controller names the number itself: --sync-rate-limit defaults to 0.3, one sync per 3.33 seconds. Raising that flag would shorten this tail — it is a default, not a limit of the design — but the default is what a cluster gets. ramjet-ingress has a fixed 200 ms debounce and no rate limit.

A loss: the write path

The admission webhook is not the reason, and this is where ingress-nginx ties or wins. Median kubectl apply was 138 ms for ingress-nginx against 159 ms for ramjet-ingress — the write path including nginx -t validation of the whole generated configuration is faster than ramjet-ingress’s plain unvalidated write.

500 routes

500 Ingresses with distinct hosts, applied as one batch, one contender at a time.

ContenderCreatedkubectl apply wall timeApply → last route servedController CPUController memory before → afternginx reloads
ramjet500/50010.7 s10.9 s1 s21.1 → 20.8 MiB
nginx500/50058.5 s61.7 s67 s115.8 → 214.0 MiB19

Both reached 500. Neither choked. That is worth saying first, because the brief allowed for reporting the number at which one of them fell over.

  • 5.7x faster convergence, of which most of ingress-nginx’s time is the write path: roughly 117 ms per Ingress against 21 ms. The admission webhook that was free at one Ingress is not free at 500.
  • Controller CPU differs by a factor of 100: 0.66 CPU-seconds against 66.7.
  • Memory is the sharper result. 21.1 → 20.8 MiB: 500 compiled routes are, within measurement noise, free. ingress-nginx grew 98 MiB, roughly 200 KiB per route. Under the 256Mi limit ramjet-ingress’s own chart ships, ingress-nginx would have been within 40 MiB of being OOM-killed at 500 routes; it survives because its chart ships no limit at all.

Propagation with the routes loaded:

ContenderTrialsMedianp95Median on an empty cluster
ramjet5507 ms723 ms363 ms (1.4x)
nginx55,006 ms5,964 ms1,151 ms (4.4x)

Every ingress-nginx trial at scale was slower than its own worst trial on an empty cluster.

The throughput row of this benchmark is the weakest number in the whole document and is deliberately not reproduced here. 47,535 rps against 16,109 is a 3x gap, but the two measurements were taken minutes apart at 38% and 25% VM CPU idle respectively, on a machine that was doing someone else’s work. It is one run each under unequal conditions and should not be quoted as a throughput result.

Deleting 500 Ingresses took about 105 seconds for each. A tie, and API-server bound rather than controller bound.

Idle-connection memory: the loss

ingress-nginx wins this one decisively, and it is the most important negative result in the report.

10,000 idle keep-alive connections, no Kubernetes, both proxies on the same docker bridge with the same upstream and nginx’s own tuning. Two passes, order reversed between them.

Originally

ContenderPassIdle beforeAt 10kAfter closePer connectionRetained
ramjet11.5 MiB266.1 MiB229.6 MiB27.1 KiB+228.1 MiB
ramjet2229.6 MiB329.0 MiB292.3 MiB10.2 KiB+62.7 MiB
nginx116.3 MiB58.9 MiB16.5 MiB4.4 KiB+0.2 MiB
nginx216.5 MiB58.8 MiB16.5 MiB4.3 KiB+0.0 MiB

An idle connection cost nginx 4.4 KiB and ramjet-ingress 27.1 KiB — 6x — and ramjet-ingress did not give the memory back, growing monotonically across connect/disconnect cycles. At 266 MiB peak it would have been OOM-killed by the 256Mi limit its own Helm chart ships, with no traffic flowing.

After the fix

ContenderPassIdle beforeAt 10kAfter closePer connectionRetained
ramjet12.5 MiB200.7 MiB11.5 MiB20.3 KiB+9.0 MiB
ramjet211.5 MiB201.5 MiB11.5 MiB19.5 KiB+0.0 MiB
nginx116.3 MiB58.8 MiB16.5 MiB4.4 KiB+0.2 MiB
nginx216.5 MiB58.8 MiB16.5 MiB4.3 KiB+0.0 MiB

The retention problem is gone: a second full cycle peaked at 201.5 and settled at 11.5 again — the same number, not a higher one. ramjet-ingress now also idles lower than nginx does, 11.5 MiB against 16.5.

The per-connection cost improved by a quarter and ingress-nginx still wins it. 27.1 KiB to 20.3 is real, and 20.3 against 4.4 is still 4.6x. The gap is structural.

The original table is left in the repository exactly as it was — a benchmark that overwrites the evidence it was judged against cannot be checked afterwards.

Where the remaining 20.3 KiB goes

Measured at 2,000 connections:

What the connection has donePer connection, cgroupPer connection, VmRSS
Accepted, never sent a byte6.1 KiB1.7 KiB
One request, answered by the proxy itself20.1 KiB16.2 KiB
One request, forwarded to the upstream20.8 KiB16.9 KiB

The ~4.4 KiB gap between the columns on a merely-accepted connection is kernel socket memory, which cgroup v2 charges to the container. That is very nearly nginx’s entire per-connection cost, which is the sharpest way to state the difference: nginx’s 4.4 KiB is, to a first approximation, the socket and nothing else. It hands a connection’s request buffers back to its pool when the connection goes idle and keeps only the connection object. There is no equivalent in hyper.

Two checks that pin it down: sending 6 KiB of request headers instead of 90 bytes moved the figure by two bytes (16,927 against 16,929) — the read buffer is resident whether or not anything is read into it. And patching hyper’s INIT_BUFFER_SIZE from 8192 down to 1024 gave 11.3 KiB cgroup and 7.3 KiB RSS; two 8 KiB buffers becoming two 1 KiB ones accounts for 9.6 KiB, and nothing else moved.

There is no public API that lowers it: max_buf_size caps how far the read buffer may grow, and hyper refuses to set it below INIT_BUFFER_SIZE. So 16 KiB per idle keep-alive connection is this engine’s floor until hyper’s initial allocation follows its configured maximum instead of a constant. The patched measurement is what that change would be worth: roughly 2.5x nginx instead of 4.6x. It is a one-line change in a dependency, and the right place to make it is upstream.

The experimental uring engine was measured on the same harness and is not cheaper: 23.2 KiB per connection, because it allocates per-connection buffers of its own.

What this means for the chart

resources.limits.memory: 256Mi stays, and the values file now carries the arithmetic instead of leaving it to be rediscovered — about 20 KiB per idle keep-alive connection, so 256Mi is roughly twelve thousand of them. Raising the default to make room for a per-connection cost that is still 4.6x nginx’s would have hidden the finding rather than fixed it.

Raw forwarding throughput vs nginx

A forwarding-engine drag race: one route, one host, 128-byte plaintext responses, static configuration. Both proxies pinned to the same two cores with --cpuset-cpus=0,1 (a CPU quota is invisible to sched_getaffinity, so pinning is what makes both see 2 CPUs and start 2 workers), oha 1.16.0 at HTTP/1.1 keep-alive, a discarded 10s warmup, then 3 × 30s at c64 and 1 × 30s at c256, interleaved. The c64 rows are the median-throughput run, not a per-column average, so every number in a row comes from one real 30-second measurement.

Concurrency 64 (median of 3 × 30s runs)

ContenderRPSp50p90p99p99.9
ramjet-ingress85,908666 µs921 µs2,528 µs6,236 µs
nginx86,670671 µs873 µs2,314 µs5,902 µs
baseline (no proxy)229,400223 µs356 µs1,219 µs4,421 µs

Concurrency 256 (single 30s run)

ContenderRPSp50p90p99p99.9
ramjet-ingress82,5242,975 µs3,617 µs6,396 µs14,185 µs
nginx89,6362,652 µs3,559 µs7,683 µs17,180 µs
baseline (no proxy)247,077918 µs1,233 µs3,554 µs9,111 µs

At c64 the two are level. 85,908 against 86,670 is a 0.9% difference, and nginx’s own three runs spread 4.5% — the gap is smaller than the noise in the measurement, which means this benchmark can no longer tell them apart at this concurrency. It does not mean ramjet-ingress is faster; the honest statement is “the same”. Divide two cores by throughput and both spend 23.6 µs of CPU per request.

At c256 nginx is still 9% ahead, and that gap is outside the noise. nginx’s throughput barely moves between c64 and c256 (+3%) while ramjet’s drops 4%. Latency runs the other way — ramjet’s p99 at c256 is 6,396 µs against nginx’s 7,683 µs — so what this looks like is ramjet trading a little throughput for shorter queues under saturation, not falling over.

Where it started

The first measurement of the same benchmark had nginx 45% ahead, and it is kept in the repository unchanged. What closed it:

MeasureBeforeAfterChange
c64 throughput61,56885,908+39.5%
c256 throughput59,64482,524+38.4%
c64 p50967 µs666 µs-31%
c64 p993,107 µs2,528 µs-19%
CPU per request32.5 µs23.6 µs-27%
vs nginx at c6469% of it99% of it
vs nginx at c25667% of it92% of it
requests per upstream connection~5908,17914x
memory under load19.2 MiB33.1 MiB+72%

That last row is a real cost and is reported as one: one runtime per core means one connection pool, one timer wheel and one set of hyper buffers per core rather than per process. On a 2-core replica that is 14 MiB; on a 64-core node with no CPU limit it would be considerably more, which is an argument for setting --worker-threads deliberately rather than letting it follow the host. (The later memory work brought this to 24.9 MiB.)

Zero errors across all twelve runs, both before and after: 49,114,324 requests, every one a 200, no transport errors from any contender at either concurrency.

Method honesty on this benchmark

Both head-to-head tables were taken with a fixed within-round order. run.sh has since been changed to rotate which contender leads each round, and to wait a 15s cooldown before each warmup — because plain interleaving assumes the machine is steady within a round, and on a laptop it is not: the package heats up as the round proceeds, so a fixed order hands whoever goes first a systematically cooler machine in every round. Neither table has been re-measured under the rotated protocol, so read the numbers as carrying that bias in ramjet’s favour at c64, bounded by the within-round drift (the baseline’s 13.3% spread is the visible upper bound; the contenders’ 1.7–5.3% the likelier scale).

Other stated unfairness:

  • The upstream keepalive pools were not equal in the first measurement. nginx held 128 idle upstream connections, ramjet 64 — the edge was nginx’s, and it was left alone rather than patched, because changing the product to win its own benchmark is not a measurement.
  • nginx tuning choices were tested, not assumed. reuseport was measured both ways and kept because it is better for nginx. access_log off removes nginx’s default per-request write, which ramjet does not have. proxy_cache was deliberately not enabled: ramjet has no response cache, and serving from nginx’s memory would compare two different jobs.
  • Both are round-robin, matching nginx’s default, rather than ramjet’s leastConn.
  • A shared docker daemon, and it bit. The reported 30s runs absorb it, and the contender spread is the evidence.

What this does not test

This is a forwarding-engine drag race on the narrowest possible workload: one route, one host, 128-byte responses, plaintext HTTP/1.1, static configuration. It says nothing about the project’s actual thesis — that a config change is a pointer swap rather than an nginx reload. Nothing here exercises TLS termination, HTTP/2, large or streaming bodies, thousands of routes, or configuration churn under live traffic.

That paragraph predates the optimization and still stands unchanged. It is the more important one on the page.

Where a request actually goes

Profiling asked where the 10 µs gap lived, and the answer was not in the forwarding code. Route matching, header rewriting, URI building and the metrics counters together account for about 2% of a request.

Own-code functionInclusive CPU
upstream::endpoint_uri (builds and parses a URI per request)0.80%
headers::apply_forwarded (X-Forwarded-*)0.27%
headers::strip_hop_by_hop0.20%
headers::upgrade_protocol0.14%
forward::select_backend (the router match itself)0.13%

The router’s 25 ns match is 0.1% of that. There was no hot function to find.

What the profile found instead was the runtime moving each request’s work between cores:

WorkersThroughputProxy CPUCPU per request
147.1k rps88%18.7 µs
268.4k rps183%26.7 µs

The same code costs 43% more CPU per request on two threads than on one, held at 33–43% across three interleaved rounds. That is the shape of a work-stealing scheduler under a request that ping-pongs between workers; nginx does not pay it, because its workers are shared-nothing processes.

So the data plane became one current_thread runtime per core with nothing shared between them. Everything else that was tried was inside the noise and is recorded as such:

TriedResultKept?
Removing the per-request header clone+0.9%No — the clone buys endpoint failover for less than the noise floor
Flattened writes instead of vectored-0.2%No — the syscall dominates; iovec handling is free either way
Raising tokio’s event_interval from 61 to 512-1.3%No
Per-core sharding / cache-line padding of the metrics counters-0.9%No, and this is the useful negative: there is nothing to win, so the sharding was never written

The floor

After the change the profile reads:

CostSelf CPU
writev31.2%
read28.2%
kevent9.1%
clock_gettime2.2%
everything in ramjet_proxy~1%

59.4% of a request is the four unavoidable syscalls, and another 9.1% is finding out a socket is ready. That is the floor for this design, and it is not a hyper problem or a tokio problem — it is the I/O model.

Getting under it means fewer syscalls per request, which on Linux means io_uring.

The remaining 9% gap at c256 has not been profiled. Every measurement was taken at c64, and the native harness cannot hold c256 steady enough to be worth reading. Whether it is queueing, the per-runtime pool split, or something else is an open question.

The uring engine

A second data plane on a completion-based reactor, selected with --engine uring. Docker on Linux, --cpuset-cpus=0,1, the same pair of upstreams, oha at c64, three rotated rounds of 30 seconds each. Both ramjet rows are the same image with one flag different.

Concurrency 64 (median of 3 runs)

ContenderRPS% of baselinep50p90p99p99.9
ramjet (hyper)80,68235.1%687 µs1,030 µs2,905 µs7,562 µs
ramjet (uring)116,92750.9%483 µs702 µs1,941 µs5,569 µs
nginx80,79035.1%696 µs1,002 µs2,853 µs7,687 µs
baseline (no proxy)229,902100.0%227 µs384 µs1,160 µs3,868 µs

Concurrency 256 (single run)

ContenderRPS% of baselinep50p90p99p99.9
ramjet (hyper)65,45828.2%3,130 µs5,217 µs15,918 µs52,723 µs
ramjet (uring)110,05747.5%1,962 µs3,070 µs8,122 µs28,387 µs
nginx84,48036.4%2,680 µs3,936 µs8,465 µs22,245 µs
baseline (no proxy)231,837100.0%924 µs1,552 µs3,866 µs10,629 µs

+44.7% over nginx at the median, and the proxy hop costs 255 µs where nginx’s costs 469 µs — it keeps 51% of the no-proxy throughput where the other two keep 35%.

The claim that survives the drift

The machine would not sit still: the baseline, which has no moving parts and nothing under test, spread 15.1% across three rounds. Drift makes a median shaky. It does not touch a rank-order claim:

Comparisonworst uring roundbest rival roundverdict
uring vs ramjet (hyper)111,25085,640uring ahead by 30% at worst
uring vs nginx111,25084,862uring ahead by 31% at worst

Every measured uring round beat every measured round of both rivals; the ranges do not overlap. “At least 31% ahead of nginx” is the claim that survives, and +44.7% is the median’s reading of the same thing. report.py makes this check itself and refuses the run if the ranges ever overlap.

The hyper row is not under-measured relative to nginx: here it is 0.13% from nginx, so +44.9% for uring over hyper is the same result as +44.7% over nginx rather than an artifact of a cold hyper.

And a cross-day check that cuts the other way

Comparing this session against the committed head-to-head runs — taking both cells from the same run, which an earlier version of the engine document failed to do — puts nginx’s row here on the low side:

engine session4f58bd7, after optimizationd1c08c6, first measurement
nginx, absolute80,79086,67089,593
baseline, absolute229,902229,400247,875
nginx as % of baseline35.1%37.8%36.1%

The ratio travels across days at the few-percent level — 2.6 points against the nearer comparator, 1.0 against the older one — which is enough to trust this session’s ordering, and not enough to swap an absolute row for another day’s.

The slowdown was not uniform, and that is the part worth carrying: this session’s baseline is within 0.2% of 4f58bd7’s, while its nginx is 6.8% lower and its hyper engine 6.1% lower. Whatever cost the two TCP proxies those points did not cost the no-proxy baseline anything.

So +44.7% is the optimistic end of the margin rather than the middle of it. Against the best committed nginx median, uring’s worst round is +28%; against this session’s own best nginx round, +31%. At least 28% ahead is the figure that survives every pairing, and +44.7% is what you get comparing contenders measured in the same session on the same host.

Why: the syscall counters

cqes_per_waiting_enter    = 21.7 … 39.5   (typically 28–37)
enter_share_of_thread_cpu = 0.81

Between 22 and 40 completions are harvested per trip into the kernel. A request is four operations, so that is roughly seven to ten requests per syscall, against the hyper engine’s four syscalls per request plus a kevent to learn a socket was ready. The 81% is the share of the serving thread’s CPU spent inside io_uring_enter, which is where the kernel actually does the reads and writes — it is not overhead, it is the work.

Cost per request

requestsCPUCPU per requestmemoryreqs per upstream conn
ramjet (hyper)855,179199.1%27.9 µs23.5 MiB6,681
ramjet (uring)1,273,310198.1%18.7 µs10.8 MiB9,947
nginx986,645174.4%21.2 µs3.9 MiB61,665

49% more requests for the same CPU as the hyper engine, and less than half its memory. Two honest readings alongside that:

  • nginx did not saturate its cores in this pass (174% of an available 200%), so its 21.2 µs is a fair figure for what it spent but its throughput here may have been limited by something other than CPU.
  • A loss: nginx reuses upstream connections six times better — 61,665 requests per connection against 9,947. Per-core pools are the reason, and the price was named when they were introduced: a connection returned to a full pool on one core cannot be reused by another. It is not costing throughput here, but it is a real difference and it is nginx’s win. nginx is also, by a wide margin, the most memory-frugal of the three.

The macOS negative result

The same binary with the flag flipped, on the native macOS harness:

  A (hyper)  median 50,738 rps   spread 30.8%
  B (uring)  median 53,024 rps   spread 18.1%
  B vs A: +4.5%   (inside the noise)

+4.5% against an 18–31% spread is not a result, and the harness says so itself. That is the prediction, not a disappointment: on macOS the reactor’s backend is kqueue, which performs each syscall eagerly at submission. There is no ring, no batch, and nothing to collapse. The whole benefit measured on Linux is io_uring’s, so the platform without io_uring measures none of it.

Caveats on the uring numbers

This is a macOS Docker Desktop VM, and that matters more here than for any other benchmark in the repository. The whole result is about the cost of entering the kernel, and a syscall in a virtualised linuxkit guest is dearer than one on bare metal. io_uring’s advantage is the cost it avoids, so a more expensive syscall flatters it. Treat the margin as an upper bound and the direction, not the size, as the transferable claim.

  • nginx runs under Docker’s stock seccomp profile; both ramjet containers run under a pinned one (moby v24.0.7’s default plus the three io_uring syscalls). It is a superset of what nginx needs so it cannot disadvantage nginx, but it is a difference between contenders.
  • The two engines were not feature-equivalent when this was measured. The uring engine served HTTP/1.1 plaintext and nothing else. None of the missing features is exercised by this workload, so the comparison is like for like for this traffic. It was not a claim that the engines are interchangeable. Most of that gap has since closed — see TLS, and a tunnel below and Engines.
  • c256 is a single run, not a median, and is reported as such.
  • The correctness gate compares the two engines’ response headers field by field, so neither can be fast by doing less.

On real Linux

The caveat above says to treat the margin as an upper bound and the direction, not the size, as what transfers. This is the check on that, and it comes out exactly as the caveat predicted. A t3.xlarge EC2 instance running k0s — real Linux, no VM in the path — with the same two engines and the same flag difference:

ramjet (hyper)ramjet (uring)uring vs hyper
RPS, median10,61011,221+5.8%
spread across runs4.1%2.1%
% of baseline47.4%50.2%
p504.59 ms3.76 ms−18.1%

Baseline with no proxy in the path was 22,363 rps.

+5.8%, against +44.9% in the VM. The direction transferred and the size did not, which is the whole of what the caveat asked to be believed.

The share-of-baseline column is where the mechanism shows:

% of baselineDocker Desktop VMt3.xlarge, k0s
ramjet (uring)50.9%50.2%
ramjet (hyper)35.1%47.4%

uring held its share almost exactly; hyper gained twelve points. So the VM was not flattering the reactor so much as it was punishing the other engine. The hyper engine pays four syscalls per request plus a kevent, a virtualised syscall is dearer than a native one, and taking the VM away refunds most of that to the engine making the most calls. The reactor, which was already avoiding those calls, had little to be refunded.

Per busy CPU the gap is smaller still, roughly +2.4%, where busy is 100 − idle:

busyRPSRPS per busy point
ramjet (hyper)9110,610116.6
ramjet (uring)9411,221119.4

The us/sy split underneath is uring 35/47 against hyper 41/44 — the reactor spending more of the machine in the kernel and less in userspace, which is the same shape as the io_uring_enter share above and means the same thing: for this engine the kernel time is the work. Those two do not sum to busy; the remainder is wa, st and rounding, which is why the busy column is taken from idle rather than by adding them up.

Every figure in that table counts the load generator and the upstreams as well as the proxy, because all three were on the one box. So +2.4% is a whole-machine efficiency number rather than the engine’s own. Getting the engine’s own needs proxy-only CPU seconds, which this run did not capture — the same reason the run is not the one to quote from at all:

The caveat that matters more than any of the numbers: this run was CPU-contended, and the contention structurally compresses the gap between the engines. Four shared vCPUs, with the load generator on the same instance as the proxy and the upstreams. The proxy was therefore never the sole bottleneck, and a benchmark where the thing under test is not the limiting factor understates every difference between two versions of it. A rerun on pinned, isolated cores with the load off-box is the number worth quoting, and this is not that run. Treat +5.8% as a floor under contention rather than the engine’s ceiling.

The p99 inversion, unexplained

On the same real-Linux run the reactor wins throughput and the median and loses the tail: p99 is about 6.8% worse than the hyper engine’s, and it was worse at both c64 and c256. Consistent enough not to be noise, and not currently explained.

It is not the shape the VM measurements had, where uring led at every percentile including p99.9. Until someone can say why, no tail-latency claim is made for the reactor — see Limitations.

What is not claimed

  • Not that io_uring beats epoll in general. This is one proxy workload, one kernel, one VM.
  • Not that the uring engine is ready to deploy. At the time of this measurement it had no TLS, no HTTP/2, no upgrades, no Kubernetes mode and no graceful drain. That list is now down to HTTP/2, which is served by dispatch.
  • Not that the hyper engine is badly written. Profiling took it to the syscall floor, and this is what is underneath that floor. The difference is the I/O model, which is what was being tested.

TLS, and a tunnel

The measurements above are plaintext HTTP/1.1, which is what the uring engine could serve when they were taken. It terminates TLS now. Same machine, same topology, same pinning, same rules — one certificate added.

Both sides resume sessions, which is the setting that decides a TLS benchmark: nginx ships ssl_session_tickets on and every deployment turns on ssl_session_cache, so a run against a ramjet with resumption off would be measuring a configuration nobody deploys. ECDSA P-256, one certificate generated per run and mounted into all three containers, HTTP/1.1 on all three.

Keep-alive, 30s per run, three rounds, median by throughput:

c=64rpsp50p99
ramjet (hyper)82,7350.68 ms2.48 ms
ramjet (uring)107,9200.51 ms2.38 ms
nginx76,5310.76 ms2.40 ms
c=256rpsp50p99
ramjet (hyper)81,2652.96 ms6.24 ms
ramjet (uring)104,1882.14 ms7.02 ms
nginx68,7583.38 ms8.92 ms

A new connection per request, so every request pays for a handshake — which is what a rolling deployment or a reconnecting CDN does to a replica:

conn/sp50p99
ramjet (hyper)12,4125.09 ms12.62 ms
ramjet (uring)14,4364.25 ms11.25 ms
nginx5,73610.68 ms25.86 ms

The margin over hyper survives TLS almost intact: 1.30x at c=64, 1.28x at c=256, against 1.30x on the plaintext run. Crypto is added work, but it is added to both ramjet contenders equally, and what separates them is still how the bytes reach the record layer rather than what happens inside it.

The handshake row narrows to 1.16x between the engines, and it should — a handshake is arithmetic, not syscalls. The 2.52x over nginx there is ring against OpenSSL as much as it is one proxy against another, and should be read as a statement about two TLS stacks.

WebSocket tunnels are level, and that is the expected result

64 tunnels, 128-byte payloads, one echo in flight per connection:

echo/sp50p99
ramjet (hyper)103,422595 µs1,509 µs
ramjet (uring)105,271585 µs1,374 µs
nginx102,843571 µs1,619 µs

All three within 2.4%. A reader who saw 1.30x on the TLS table would expect a gap here, and there is none, for a reason worth stating: after a 101 there is no request, no routing and no header rewriting — one read and one write per echo on each side, with nothing to batch and nothing to overlap. The rate is bounded by round-trip latency rather than by how many times the kernel is entered, and submission batching is exactly the advantage that has nothing to work on.

What does separate them is steadiness: both ramjet engines hold a 1.2–1.4% spread across runs against nginx’s 7.1%, and the uring engine has the best p99.

Full protocol, raw JSON and the fairness notes are in bench/engine/RESULTS.md.

Route matching

Not a system benchmark — a microbenchmark of the matcher against a table of 1,000 hosts and 10,001 routes, on an Apple M2 Pro, criterion, 100 samples.

CaseTimeWhat it costs
deep_prefix_hit25.2 nsexact host, four-segment prefix — the normal request
exact_hit22.5 nsexact rules sort first, so this is the cheapest hit
host_miss_default_backend20.6 nstwo failed hashes, then the default backend
wildcard_hit29.7 nsa failed exact hash plus a parent-domain hash
uppercase_host_fold31.8 nsthe only path that copies, into a stack buffer
regex_hit42.8 nsfull scan past every prefix, then a regex
root_prefix_hit47.3 nsworst case: scans every prefix rule before matching /

For scale, a single uncached main-memory reference is roughly 80 ns — matching a route costs less than one cache miss.

These are laptop numbers taken on a machine that was not otherwise idle, so treat them as an order of magnitude rather than a regression baseline. What the benchmark is really for is the shape: matching does not get slower with table size, because host selection is a hash and a host carries a handful of rules.

match_request performs no heap allocation, and that is enforced rather than asserted in a comment: tests/no_alloc.rs installs a counting global allocator and checks every path through the matcher, including the mixed-case fold, canary resolution, and SNI lookup.

Where ingress-nginx won or tied

Result
Idle-connection memoryWon, heavily. 4.4 KiB/connection against 27.1, and it returns all of it on close while ramjet-ingress retained and grew. Still won after the fix, by less: 4.4 against 20.3, and both now return what they took
kubectl apply write path (single Ingress)Won in the VM, drew on real Linux. 138 ms median against 159 there; 191 against 189 on EC2. Either way it is doing strictly more work in the time — nginx -t validation through an admission webhook that ramjet-ingress does not have
Endpoint-only churn: connection safetyTied. 50/50 idle connections survived, zero errors, zero reloads, in both environments. Its Lua balancer does exactly what it claims and the reload argument does not apply to endpoint changes
Endpoint-only churn: CPUWon a retraction. The VM run’s +25% CPU-per-request penalty on that path did not survive the EC2 rerun with the arm order reversed: +1%. The non-reloading path is close to free and the earlier figure was an ordering artifact
Deleting 500 IngressesTied. ~105 s each; the API server is the bottleneck, not either controller
Reaching 500 routes at allTied. Both converged; neither fell over
Stall severityTied-ish. Neither contender produced a stall over one second attributable to churn. ingress-nginx’s reload is visible in the tail, but it is tens of milliseconds, not seconds

Plus, from the raw-forwarding and engine benchmarks: nginx is still 9% ahead at c256, reuses upstream connections 3–6x better, and is by a wide margin the most memory-frugal of everything measured.

Reproducing any of it

./bench/run.sh                                   # ~15 minutes, cleans up after itself
python3 bench/report.py                          # re-render tables from committed JSON
IMAGES="ramjet:before ramjet:after" python3 bench/ab.py

bench/thesis/run-all.sh                          # the whole cluster suite
python3 bench/thesis/report.py
bench/thesis/teardown.sh                         # remove everything, and verify it

bench/thesis/ec2/setup.sh                        # the same thesis on a real k0s node
bench/thesis/ec2/run-all.sh                      # ~50 minutes
python3 bench/thesis/ec2/report.py
bench/thesis/ec2/teardown.sh

./bench/engine/run.sh                            # ~25 minutes
python3 bench/engine/report.py

cargo bench -p ramjet-router                     # the matcher microbenchmark

Raw data lives beside each harness: bench/results/ (current), bench/results/4f58bd7/ and bench/results/before/ (kept verbatim), bench/thesis/results/ including b4/ and b4-after/, bench/thesis/results-ec2/, and bench/engine/results/. Each archived run keeps its own versions.txt, diagnostics.txt and table.md.

Two harness rules worth knowing before you re-run anything:

  • Do not shorten WARMUP below 10s, including in the smoke test. A 2s warmup measured 39,000 rps at c64 and 72,549 rps at c128 in the run immediately after — throughput rising with concurrency is the signature of a first run that had not finished warming.
  • Check the host before trusting any number. On a quiet host the baseline measures ~248,000 rps; below ~230,000 the host is busy and the run is not worth starting. Host-side contention reaches the guest through vCPU preemption, so container --cpuset pinning does not protect against it. Overriding any tunable redirects output to results/scratch/ so a smoke run cannot overwrite a real measurement.