Skip to content

v0.3.19

Choose a tag to compare

@github-actions github-actions released this 06 Oct 11:09
· 122 commits to master since this release
6795d75

変更(速さとメモリの改善)

v0.3.18 と同じマシンで交互に測った負荷テストの結果を添えます(4 vCPU の GitHub のランナー、名前空間の中の仮想の回線。実際の NIC では差が小さくなる見込みです)。

  • L4 の TCP の転送を作り直しました。 読めるデータがあるときだけ 32 KiB のバッファを借りるようにし(何もしていない接続はバッファを持たない)、大きな転送が続く向きは splice(2) でカーネルの中だけで受け渡すようにしました(plain の TCP だけ。TLS 終端・STARTTLS・L7 は対象外)。
    • TCP 1 本 5.0 → 12.9 Gbit/s、8 本 11.9 → 43.6 Gbit/s。大きな転送の CPU あたりの量は約 7 倍。
    • TCP の接続 1 本あたりのメモリ 17.7 / 25.7 KiB(待機中 / 転送中)→ 7.5 KiB。
    • 小さなやり取りの遅延・新しい接続の数は変わりません。
  • L7(HTTP/2)の CPU を減らしました。 転送先への待機中の接続を多く使い回し(サーバごとに最大 1024 本、4 秒使わなければ閉じる)、リクエストごとの余分な処理を省きました。
    • HTTP/2 の小さなリクエスト +50〜60%、HTTP/1.1 +11%。遅延(p50 / p99)も 30% 前後短くなりました。
    • HTTP/2 over TLS に高い負荷をかけている間は、ピークのメモリが増えます(負荷が終われば戻ります)。
  • UDP をまとめて読み書きするようにしました(recvmmsg / sendmmsg、受信バッファを大きく)。64 バイトの全力の送信で、届く量 +65%、取りこぼし 57% → 21%。セッションあたりのメモリは変わりません。
  • メモリのアロケータに mimalloc を選べるようにしました(ビルドで --features alloc-mimalloc)。既定は今までどおり musl の malloc です。

設定の項目は変わっていません。UDP の待ち受けのソケットの数・splice の切り替えなどの設定の形は v0.4.0 で決めます。


Changed (speed and memory)

Measured against v0.3.18 on the same machine, alternating builds (4-vCPU GitHub runner, virtual links in network namespaces; expect smaller gaps on real NICs).

  • L4 TCP forwarding was rebuilt. A 32 KiB buffer is borrowed only while data is ready (idle connections hold no buffer), and directions with sustained large transfers move to splice(2), staying inside the kernel (plain TCP only; not TLS termination, STARTTLS or L7).
    • TCP 1 stream 5.0 → 12.9 Gbit/s, 8 streams 11.9 → 43.6 Gbit/s; about 7× more data per CPU on large transfers.
    • Memory per TCP connection 17.7 / 25.7 KiB (idle / busy) → 7.5 KiB.
    • Small-message latency and new connections per second are unchanged.
  • Less CPU for L7 (HTTP/2). More idle upstream connections are reused (up to 1024 per server, closed after 4 s unused) and per-request work was trimmed.
    • Small HTTP/2 requests +50–60%, HTTP/1.1 +11%; p50/p99 latency about 30% lower.
    • Peak memory is higher while HTTP/2 over TLS is under heavy load (it goes back down afterwards).
  • UDP now reads and writes in batches (recvmmsg / sendmmsg, larger receive buffer): at full rate with 64-byte packets, +65% delivered and loss 57% → 21%. Memory per session is unchanged.
  • mimalloc can be chosen as the allocator (build with --features alloc-mimalloc). The default stays musl's malloc.

No settings changed. The shape of settings such as the number of UDP listener sockets and the splice switch will be decided in v0.4.0.