y-http-bench
A zero-dependency HTTP benchmarking CLI built on the Go standard library
Stack
- Go(纯标准库)
- 单文件
- 交叉编译
- CLI
Overview
y-http-bench is an HTTP load-testing (benchmark) tool written in pure Go standard library: a single file, zero dependencies, works out of the box. A build produces one executable — no runtime to install, no third-party packages.
It targets the "quickly find out how much an endpoint can take" scenario — pick fixed request count, fixed duration, or fixed QPS — and prints a latency percentile report when the run ends.
Quick Start
# Fixed request count: 20000 requests at 100 concurrency
y-http-bench -url https://example.com -c 100 -n 20000
# Fixed duration: 50 concurrent workers for 30 seconds
y-http-bench -url https://example.com -c 50 -d 30s
# Rate limited: steady 1000 QPS for 1 minute
y-http-bench -url https://api.test/users -c 50 -d 1m -qps 1000
# POST endpoint: custom headers + body
y-http-bench -url https://api.test/login -m POST \
-H "Content-Type: application/json" \
-H "Authorization: Bearer xxx" \
-body '{"user":"tom","pwd":"123"}' -c 20 -n 5000
# Read request body from file + export raw latency data for plotting
y-http-bench -url https://api.test/order -m POST \
-body-file ./payload.json -c 30 -d 20s -dump-latency latency.csv
At least one of -n and -d must be specified (otherwise the test never ends). If both are given, whichever condition is reached first ends the test. Pressing Ctrl+C during a run stops the benchmark and still prints a report based on the samples collected so far.
Flags
| Flag | Default | Description |
|---|---|---|
-url |
none (required) | Target URL |
-m |
GET |
HTTP method |
-c |
10 |
Concurrency (number of workers) |
-n |
0 |
Total number of requests; 0 means unlimited (use with -d) |
-d |
0 |
Test duration, e.g. 10s / 2m; 0 means unlimited (use with -n) |
-qps |
0 |
Global rate limit (requests/second); 0 means unlimited |
-timeout |
10s |
Per-request timeout |
-H |
none | Custom header; may be repeated |
-body |
empty | Request body as a string |
-body-file |
empty | Read the request body from a file |
-keepalive |
true |
Reuse connections; set to false to simulate short-lived connections |
-insecure |
false |
Skip TLS certificate verification (self-signed certificates) |
-dump-latency |
empty | Stream per-request latencies to a CSV file (columns: latency_ms,status,bytes,error); uses no extra memory |
-v |
false |
Print all error categories (at most 5 are shown by default) |
Output
During the run, a live progress line refreshes every 500ms (elapsed / completed / current QPS / failures). When the run finishes, a report is printed:
- Overview: total requests, failures, success rate, average QPS, throughput (with
-qpsset, the target QPS and achieved-rate ratio are also shown) - Latency distribution: Min / Avg / P50 / P90 / P95 / P99 / P99.9 / Max (milliseconds, measured until the response body is fully read)
- Status code distribution
- Error distribution (network-layer errors aggregated by type)
Latency accounting: only requests that received a response are counted (including 4xx/5xx); network-layer errors only contribute to the failure count and error distribution. When -qps is set, the timer starts at the request's "ideal send time", so queueing time is included in the latency.
Implementation Notes
- The rate limiter does not use
time.Ticker: on Windows the minimum ticker interval is limited by system timer resolution (~15.6ms), capping throughput at 60–90 req/s (a configured 5000 only achieved 88 in testing). Instead it uses an "absolute schedule": the ideal send time of the n-th request isstart + n×interval, computed fromtime.Now()only. Measured: configured 200 / 1000 / 5000 achieved 188 / 985 / 4982 req/s - Latency is measured from the "ideal send time" (a coordinated-omission fix): queueing time from rate limiting or from workers slowed by earlier requests is included in latency — otherwise P99 looks too optimistic under heavy load (same rationale as wrk2's fixed-rate mode)
- Latency statistics use a histogram, not full samples: logarithmic-linear bucketing, 1024 buckets per magnitude across 32 magnitudes (a fixed 256KB regardless of request count), with a relative error bound of ~0.1%; quantiles take the bucket midpoint clamped to
[Min, Max] -dump-latencystreams to disk: samples are written into abufio.Writeras they are merged, instead of holding everything in memory- HTTP/2 is enabled explicitly: when you hand-write an
http.Transport, Go does not enable HTTP/2 automatically —ForceAttemptHTTP2: trueis required - Table headers are aligned by display width: CJK characters occupy 2 terminal columns, so
padRight/displayWidthpad spaces manually to keep tables from misaligning
Known Limitations
- Still a closed-loop benchmark: each worker waits for the previous response before sending the next. With
-qpsset, queueing time is added back into latency, but true fixed-rate sending unaffected by server slowdowns requires an open-loop model - Single machine, single process; no multi-URL / multi-phase scenarios, ramp-up, think time, assertions or SLA threshold checks, and no distributed mode
- One allocation per request (
Request.Clone+ a fresh body reader); at extreme throughput it can't match epoll-based implementations with buffer reuse such as wrk - Client-side metrics are not broken down: connection reuse counts and DNS / TCP / TLS handshake timings are not reported separately
Benchmarking Tips
- You are load-testing the target service, not your local network — run from a neutral machine in the same data center / intranet as the target
- First figure out where the bottleneck is — watch the target machine's CPU, GC, and connection pool saturation
- Focus on P99 / P99.9 rather than averages; averages are easily skewed by masses of fast requests and hide the long tail
- On Windows, ephemeral ports are recycled slowly; under heavy short-connection load you may hit
connectexrejections — a limitation of the target side or the OS, not the tool
License
Last updated