Repository navigation
v0.3.19
·
122 commits
to master
since this release
変更(速さとメモリの改善)
v0.3.18 と同じマシンで交互に測った負荷テストの結果を添えます(4 vCPU の GitHub のランナー、名前空間の中の仮想の回線。実際の NIC では差が小さくなる見込みです)。
- L4 の TCP の転送を作り直しました。 読めるデータがあるときだけ 32 KiB のバッファを借りるようにし(何もしていない接続はバッファを持たない)、大きな転送が続く向きは splice(2) でカーネルの中だけで受け渡すようにしました(plain の TCP だけ。TLS 終端・STARTTLS・L7 は対象外)。
- TCP 1 本 5.0 → 12.9 Gbit/s、8 本 11.9 → 43.6 Gbit/s。大きな転送の CPU あたりの量は約 7 倍。
- TCP の接続 1 本あたりのメモリ 17.7 / 25.7 KiB(待機中 / 転送中)→ 7.5 KiB。
- 小さなやり取りの遅延・新しい接続の数は変わりません。
- L7(HTTP/2)の CPU を減らしました。 転送先への待機中の接続を多く使い回し(サーバごとに最大 1024 本、4 秒使わなければ閉じる)、リクエストごとの余分な処理を省きました。
- HTTP/2 の小さなリクエスト +50〜60%、HTTP/1.1 +11%。遅延(p50 / p99)も 30% 前後短くなりました。
- HTTP/2 over TLS に高い負荷をかけている間は、ピークのメモリが増えます(負荷が終われば戻ります)。
- UDP をまとめて読み書きするようにしました(recvmmsg / sendmmsg、受信バッファを大きく)。64 バイトの全力の送信で、届く量 +65%、取りこぼし 57% → 21%。セッションあたりのメモリは変わりません。
- メモリのアロケータに mimalloc を選べるようにしました(ビルドで
--features alloc-mimalloc)。既定は今までどおり musl の malloc です。
設定の項目は変わっていません。UDP の待ち受けのソケットの数・splice の切り替えなどの設定の形は v0.4.0 で決めます。
Changed (speed and memory)
Measured against v0.3.18 on the same machine, alternating builds (4-vCPU GitHub runner, virtual links in network namespaces; expect smaller gaps on real NICs).
- L4 TCP forwarding was rebuilt. A 32 KiB buffer is borrowed only while data is ready (idle connections hold no buffer), and directions with sustained large transfers move to splice(2), staying inside the kernel (plain TCP only; not TLS termination, STARTTLS or L7).
- TCP 1 stream 5.0 → 12.9 Gbit/s, 8 streams 11.9 → 43.6 Gbit/s; about 7× more data per CPU on large transfers.
- Memory per TCP connection 17.7 / 25.7 KiB (idle / busy) → 7.5 KiB.
- Small-message latency and new connections per second are unchanged.
- Less CPU for L7 (HTTP/2). More idle upstream connections are reused (up to 1024 per server, closed after 4 s unused) and per-request work was trimmed.
- Small HTTP/2 requests +50–60%, HTTP/1.1 +11%; p50/p99 latency about 30% lower.
- Peak memory is higher while HTTP/2 over TLS is under heavy load (it goes back down afterwards).
- UDP now reads and writes in batches (recvmmsg / sendmmsg, larger receive buffer): at full rate with 64-byte packets, +65% delivered and loss 57% → 21%. Memory per session is unchanged.
- mimalloc can be chosen as the allocator (build with
--features alloc-mimalloc). The default stays musl's malloc.
No settings changed. The shape of settings such as the number of UDP listener sockets and the splice switch will be decided in v0.4.0.