Skip to content

v0.6.0: no proxy needed, and read for what a stranger can do

Choose a tag to compare

@nevindra nevindra released this 25 Sep 05:09
· 75 commits to main since this release

0.6.0 is the release where nilo stopped needing a proxy in front of it, and was then read line by line for what a stranger can do without one. It serves HTTPS itself, answers gRPC, compresses its answers and refuses a cross-site write, each behind an option or a build flag, so a program that asks for none of it is the size it was. Then every file under http/ was read for what a client on the socket could make it do, and what that found is under Fixed: a JSON body that could take the process down, a content type, cookie or event that could forge a header or an event of its own, a cached page that could reach the wrong user.

Needs Zig 0.16, as 0.5.0 does. Four things in it, in the order you will meet them:

  • Nothing in front, if you want. listen(.{ .tls = … }) is HTTPS on a build that asks for it with -Dtls; .also answers on more than one address; .grpc = true serves unary gRPC on the routes you already have; app.compress gzips text answers; nilo.csrf.sameOrigin refuses a cross-site write with no token; and the session secret changes without signing anybody out.
  • Faster where it was slow. Every executor accepts (short connections, 660K to 1.97M requests a second on eight cores), a response is flushed before the connection waits rather than on every send (sixteen pipelined requests, 3.05M to 12.3–13.6M), a route is found in a tree (125 ns to 27 ns on a 276-route table), and a server that is not busy spends a third less CPU on each request.
  • A Row that carries more. Its parent through a foreign key, its children in one more statement for the whole page, or a sum by group, all read by the db.select you already call. Beside it: $n spelled for SQLite, rawPage, rawExactlyOne, statements composed at run time, and a job that says how urgent it is. Most of the SQL half came from the first application built on 0.5.0, a SQLite one, and the eight things it sent back.
  • What a stranger can do, closed. A JSON body nested past 64 levels, header and event-stream injection, cookie attributes, Cached beside a session, the origin null, an unpinned script on /docs, four framing holes in the HTTP/1.1 parser, and work in nilo.spawn writing into somebody else's request.

Eleven entries ask something of you, under Read this before you deploy with the fix beside each. The one most programs meet first is an insert that leaves out a column nothing fills, which no longer compiles.

zig fetch --save 'git+https://github.com/nevindra/nilo?ref=v0.6.0#221e1b3eaed531efe13de7ab39dedc4091a8775c'

Keep the #commit: the tag is annotated, and Zig 0.16's zig fetch does not peel it, so ?ref=v0.6.0 alone is whatever main is that day.


Read this before you deploy

Five of these the compiler finds for you. The other six change what a running program does, and those are the ones to read: a queue needs a column added before the new binary claims from it, and an origin list with null or a trailing slash in it now stops the server at startup.

The compiler will catch these

  • A handler that takes a Cached(…) and also takes who the caller is no longer compiles. A kept answer goes to whoever asks next, so a route reading a Session(T), an Authorization, a Verified(…) or a FromHeader of Cookie or Authorization served the first caller's answer, and any cookie it set, to every caller after them. A resolved type of your own that stands for the caller declares pub const nilo_reads_caller = true; to be refused the same way. To upgrade: drop the Cached from a route that answers per caller, or move what is the same for everybody to a route of its own (ADR 188).
  • An insert that leaves out a column nothing fills no longer compiles. insert, insertMany, insertOrIgnore and insertOrUpdate name the columns and refuse, where the same insert used to fail with NotNullViolated the first time it ran. What an insert may leave out is the integer key a sequence fills, a column with a .default, an optional one, and one named in the new marker word .filled, which says the database fills it by means of its own (a DEFAULT written in a step, gen_random_uuid(), a trigger) and renders nothing. A Row with .managed = false is not checked. To upgrade: for each refusal, write the column, move its default into .default, or name it in .filled (ADR 181).
  • Store.claim takes the kinds the program can run: claim(scope, comptime kinds: []const []const u8, now, lease_until). job.Jobs passes its own kind_names; a store written outside nilo takes the parameter and narrows its claim by it. The names are bound, not spelled into the statement, so a kind may go on being named whatever it is named (ADR 215).
  • static.load takes a fifth argument, Absent, saying whether a directory that is not there is reported in one line or handed back alone. Only a caller of the module directly sees it; the four static calls on the App pass it (ADR 207).
  • ?nilo.Status(code, T), ?nilo.Response(T), ?nilo.Redirect(code) and ?nilo.Versioned(T) as a handler's return type are compile errors naming the shape to write. The first two compiled and sent the wrapper struct itself as JSON, headers and all, then crashed; the ? goes inside, Status(201, ?T) (ADR 203).

These change what a running program does

  • cors.Origins.set and setSplit refuse null and anything that is not an origin, with the new error.OriginNotAnOrigin, and cors.with refuses them while compiling. null is what any site's sandboxed frame sends, so WEB_ORIGINS="https://app.x,null" let every site read credentialed answers and pass csrf.reading; an origin with a path or a trailing slash never matched anything. A setting that holds either now fails at startup: remove null, and write https://app.example.com rather than https://app.example.com/ (ADR 088).

  • nilo.Room.join returns error.AlreadySeated for a socket that has a seat in another room, which a switch on Room.Error with no else has to name (ADR 035).

  • nilo_jobs gains a priority smallint NOT NULL DEFAULT 1 column. job.Table creates nothing, so a caller adds it beside their own rows:

    ALTER TABLE nilo_jobs ADD COLUMN priority smallint NOT NULL DEFAULT 1;

    The default is not decoration. A column that may not be null and has no default is the one ADD COLUMN that fails on a table with rows, and during a rolling deploy an older binary's INSERT — or a sibling binary's, which is ADR 215's own case — never names the column. createMissing will not help here: it creates tables that are missing and leaves an existing one alone. A queue that has not been altered is not subtly wrong — the claim names the column, so it fails loudly on the first claim rather than quietly ordering by something else. The index stays (state, run_at); widening it to include priority measured 120x slower on a queue holding future-dated rows, which is bench/result/job.md (ADR 214).

  • A Db's schema check and version guard run from a new service hook, nilo_check, which the App calls after the work app.before registered and before the first request; nilo_start opens the pool and nothing more. Under an App nothing changes but the order, which is the fix: a first boot with createMissing in before and db.checking(schema) beside it used to fail on the tables the next line would have made. A program driving a Db with no App and relying on nilo_start to check calls db.nilo_check(io) after its own boot work, or db.checkSchema (ADR 180).

  • app.start(io) runs the work before registered, and every nilo_check after it, which its doc comment already promised. A test that registered createMissing with before and then made the tables itself makes them twice, harmlessly (same ADR).

  • read_buffer defaults to 16 KiB, up from 8. It is also the ceiling on a request head, and 8 KiB was inside what a browser behind a single sign-on sends in cookies on every request. An idle connection holds the same 4,669 bytes — the pages go back while it waits (ADR 062) — and a connection inside a request holds two pages more. A server that wants the old number passes .read_buffer = 8 * 1024 (ADR 196).


What else is new

Nothing in front, if you want

HTTPS, more than one address, gRPC, compression and a cross-site check, each behind an option or a build flag. A program that names none of them builds and runs as it did.

  • listen(.{ .tls = .{ .cert = "…pem", .key = "…pem" } }): HTTPS, TLS 1.3, on a build that asked for it with .tls = true on the dependency (-Dtls in this repository). The library behind it, ianic/tls.zig, is fetched and linked only behind that flag, so every other build is what it was, and refuses the option at listen() in one line rather than serving plain HTTP on the port. What it costs is stated where the option is: 560 KB of binary and a page per idle connection in the build that asked, about 300 µs of CPU per handshake with an ECDSA certificate and 2.6 ms with an RSA-2048 one, no session resumption, one certificate per listener, a restart to reload it, and no audit behind the library, which is why a proxy in front stays the recommendation for anything on the internet (ADR 212, amending ADR 027; deploying). The handshake is bounded by header_timeout_ms, and its signature is computed on the blocking pool rather than on the executor that accepted the connection, so a burst of new connections does not hold up the requests already open on that thread: on the benchmark arena's 8gbit shape that is a mean of 150 µs rather than about 1 ms (ADR 217). zig build bench-tls-server -Dtls and bench/mem.py --tls are the benchmark server and the idle reading for it. Until ianic/tls.zig#59 merges, the library is pinned to a fork that adds one commit to upstream's zig-0.16.x: it signs through the key's CRT form, which takes an RSA-2048 handshake from 13.7 ms of CPU to 2.6 (the run).
  • listen(.{ .tls = … }) refuses a key that is not its certificate's, before it takes the port, naming both files and the two openssl commands that print the two public keys. Both files parsing is not the same as their being a pair: until now a privkey.pem from the wrong letsencrypt/live/ directory bound the port, logged that it was listening, and then failed every handshake with the reason visible only to the client. What is compared is the leaf's public key against the one the private key carries — the SEC1 point for EC, the modulus for RSA, the 32 bytes for Ed25519 — and a scheme with no prong is accepted rather than refused, so a key type the library learns later cannot stop a server that was serving. Nothing on the request path, and nothing in a build without -Dtls, where the check is not compiled at all (ADR 212).
  • listen(.{ .also = &.{ .{ .port = 8081, .tls = … } } }): more addresses to answer on, from one process. An entry is an address, a port and a certificate and nothing else; everything else on listen() belongs to the server rather than to one of its addresses, and max_connections counts the sockets this process holds rather than the sockets a port holds. One route table, one thread pool, and a handler is not told which listener a request arrived on. A cleartext port beside a TLS one is what it is for. Two entries naming one address are refused at listen() naming both, rather than arriving as the kernel's AddressInUse. boundPort() still answers for port, the first listener. Costs 82 KB of resident memory per extra listener on sixteen threads and nothing per connection: an idle connection measured 9,300 bytes before and after (ADR 213, deploying).
  • A listener can answer gRPC, in a build that passes .grpc = true to the dependency (-Dgrpc in this repository). .grpc = true on an entry in also, or on listen()'s own options for a server that speaks nothing else, serves unary calls over h2c, or over TLS with ALPN h2 when the listener has .tls as well. A method is an ordinary route, app.post("/package.Service/Method", handler): c.body() is the message, unframed and gunzipped, c.send(200, "application/grpc", bytes) is the answer, and a route that fails goes back as the grpc-status its status means, with the failure's error as grpc-message. grpc-timeout is the request's deadline, and limits.request_deadline_ms no longer lengthens one a request brought. The codec is the caller's: a zig-protobuf type decodes c.body(). A build without the flag contains none of it; the one with it is 116 KB larger on examples/hello, and an idle gRPC connection costs under a page more than an HTTP/1.1 one. A connection is held to the limits an HTTP/1.1 one is. The calls still arriving hold one max_body between them, and past that a call waits on its window, which a client may not send past. A call still arriving is cancelled after body_grace_ms plus max_body at body_min_rate, and a connection that owes a call and sends nothing for body_timeout_ms is sent away. An answer stuck on a zero window is let go of write_timeout_ms after it first stuck, 1,000 frames in a row that move no call forward send the connection away, and a header block is counted against the list limit before any of it is copied. Streaming calls and HTTP/2 for ordinary routes are not served (ADR 220).
  • zig build fuzz -- --frames throws generated HTTP/2 connections at the gRPC listener, and zig build test replays its corpus.
  • app.compress(.{}): every answer that is text, at least min_bytes (1 KB) long and going to a client whose Accept-Encoding takes gzip goes out gzipped, per request, with Content-Encoding: gzip, Vary: Accept-Encoding and the compressed length; a client that did not ask gets the body as it is. level is .fastest, .default or .best. The compressor is borrowed from a pool of one per thread, ~288 KB each, built when the chains are resolved; a compressed request pays one arena allocation for the compressed body and nothing per connection. Streams, event streams and static files are not touched: files were gzipped once at load. nilo.compress.Options, zig build bench-compress (ADR 211).
  • nilo.csrf.sameOrigin refuses a request that changes something from a page this server does not serve, a 403 before the handler runs, naming the origin. The browser already says where a request came from, in Sec-Fetch-Site and Origin, and a page can forge neither, so there is no token, nothing in the session and nothing in the form. GET, HEAD and OPTIONS are not asked, and a request with neither header (curl, a webhook) goes through. It refuses the one case SameSite=Lax lets through, a page on another subdomain of the same site. csrf.with(.{ .origins = … }) names a front end on another origin, and csrf.reading(&origins) takes the same cors.Origins as cors.reading. Opt-in; no allocation, nothing per connection, and nothing in a program that does not name it (ADR 224).
  • The session secret can change without signing anybody out. listen(.{ .session_secret = new, .session_fallback_secrets = &.{old} }) seals every new session under new and still opens a cookie sealed under old. Once the longest max_age you seal with has passed (24 hours if you never set one), no cookie sealed under old can open anyway, and it can be dropped. On several instances it is two deploys: stage new as the fallback first, then swap. Up to three fallback secrets. listen() refuses one of the wrong length, fallback secrets with no session_secret, and a fallback secret that is the current one or listed twice. The cookie format is unchanged, so nobody is signed out by the upgrade. A cookie under the current secret costs what it did. One the current secret does not open pays one refused decryption, about 270 ns, per fallback. Not for a secret that leaked: drop that one outright (ADR 225).

Faster where it was slow

Each of these was found by a benchmark shape that nilo did badly on, and each says what it was before and after.

  • Every executor accepts. listen() used to take connections on one fiber, which capped a server at ~43,000 connections a second whatever its thread count — the arena's short-lived profile read 426K req/s on 18 of 64 cores, with a p99 that was the backlog divided by that rate. Now one acceptor sits in accept on each executor and the connection is dealt round-robin as before. Connections closed after ten requests: 660K → 1.97M req/s on eight cores, p99 8.2 ms → 1.5 ms; keep-alive throughput unchanged. One parked fiber per thread for the life of the server, nothing per connection, and one timer per server fewer per connection accepted. Nothing changes in what listen() takes (ADR 200).
  • A response is flushed before the connection waits, not before send returns. When the client's next request is already in the read buffer, c.send leaves the response in the write buffer and it goes out with the next one, so sixteen pipelined requests are answered in one write rather than sixteen; a client that sends one request and waits, which is every browser and every client by default, is answered on send exactly as before. The same for a WebSocket's send, print and json, and a room's posts. The Engine guarantees the hold is never for good: every socket read flushes what is pending first, the WebSocket's wait does too, and a closing connection flushes before its FIN. Sixteen pipelined /health on eight cores: 3.05M → 12.3–13.6M req/s, p99 1.8 ms → 370 µs; sixteen pipelined echoes: 3.6M → 37M messages a second. Keep-alive and one-frame-at-a-time shapes are unchanged. What a pipelining client gives up is that a fast answer queued behind a slow handler now arrives with the slow one, bounded by write_buffer (ADR 201).
  • listen() takes backlog: how many completed handshakes the kernel holds for accept. 4,096 — net.core.somaxconn's default, what Go listens with — up from zio's 128, which nilo had been passing without saying so. Past the backlog a SYN is dropped, not refused, and the client retries a second later with nothing in the server's log: a burst of a thousand connections against 128 put 623 of them on that one-second retry, against 4,096 none. A queue capacity, so it costs nothing on a quiet server. bench/burst.py is the regression check (ADR 198).
  • The accept loop waits out a descriptor shortage instead of returning. ProcessFdQuotaExceeded, SystemFdQuotaExceeded and SystemResources from accept now sleep 5 ms, doubling to a second, and try again, with one warning per shortage; before, any of them ended listen() with a clean "nilo stopping" — at about a thousand connections on a default ulimit -n, well short of max_connections. listen() also warns at startup when the process's descriptor limit is below max_connections, with both numbers and the ulimit -n / LimitNOFILE= to change. bench/fdlimit.py is the regression check (ADR 194).
  • A refused request is hung up on with a FIN before the close, so its answer reaches the client. A 431, a 400 or 415 with a body behind it, a 413 for a body past max_body, a shed 503 — each left the client's bytes unread on the socket, and closing over unread input sends a reset, which a Windows client answers by throwing the buffered 431 away. The send side is shut first and what arrives is discarded, bounded at 64 KiB and one second. An ordinary Connection: close is untouched. The Engine contract gains Waker.halfClose (ADR 195).
  • nilo's HttpArena entry subscribes to echo-ws-pipeline and echo-ws-limited, each held back until the server was right for it: the first waited on ADR 201, the second on ADR 200 and then ADR 202, because the first of those alone had made the shape three times worse (frameworks/nilo).

A Row that carries more

Most of this came from the first application built on 0.5.0, a SQLite one, and what it found the guide showing on one database and not the other.

  • A Row can carry the row its foreign key points at, the rows that point back, or a sum by group, and db.select, db.one, db.find, db.page, db.count, db.exists and db.stream read it with no new option (ADR 218). On a narrower Row, a field whose type is another table's Row is a parent, joined through the .references between the two tables (LEFT JOIN when it is ?P, which it has to be exactly when the column may be null); a []const C field is children, read by one more statement for every row at once, so a page of twenty is two statements rather than twenty-one; and pub const nilo_aggregate = .{ .n = .count, .owed = .{ .sum = .amount } } makes the Row one row per group, its other fields the keys. pub const nilo_via names the column when the schema has several references or none. Conditions and orders reach through a parent's field (.where = .{ .customer = .{ .name = "Acme" } }), a condition on an aggregate is a HAVING, and sql.Ordering takes a path, .{ .customer, .name }. New: db.exactlyOne and tx.exactlyOne for a Row whose every field is an aggregate, and sql.exactlyOneFor and sql.childrenFor. Twenty-five new Refusals hold the rules, including the type each aggregate has to be read as. A program with no such Row is the same size, within 352 bytes either way. The guide page is a Row with more in it.
  • The $n in a raw statement are respelled for the dialect while compiling, ?n on SQLite, so WHERE ($2 IS NULL OR x = $2) binds the second value on both databases; SQLite read $2 as a named parameter indexed by first appearance and took the first. A statement naming $3 and handed two values is a Refusal on both. exec sends its run-time text as written (ADR 204).
  • db.rawPage(Row, c, sql, values) and tx.rawPage: a raw statement read as a page, the Row's columns then count(*) OVER () as one more on the end of the list, answering the same Page(Row) db.page does. A list exactly the Row's width is a Refusal (ADR 205).
  • db.rawExactlyOne(Row, c, sql, values) and tx.rawExactlyOne: rawOne for a statement that has one row by construction, an aggregate with no GROUP BY or a RETURNING on a keyed write, answering the Row and error.QueryFailed for none (ADR 206).
  • sql.Composed, db.compose and db.composed / db.composedOne / tx.composed: a statement composed at run time from literals, checked identifiers and parameters — the pieces a query engine has — and from nothing that can carry a run-time string. Placeholders are spelled for the Db's dialect; the values are counted against them at run time. raw is unchanged (ADR 208).
  • A service may declare pub fn nilo_check(self: *T, io: std.Io) !void, run once after before and before the first request; a failure is a boot that does not happen. A wrong arity is a Refusal (ADR 180).
  • examples/sqlite/: two Rows on one SQLite file, the tables made at boot with createMissing in before and checked after, a list with a Query, a paged join through rawPage, a report through rawExactlyOne and raw, and a transaction. zig build run-sqlite; its tests run under test-sql.
  • pub const priority: job.Priority = .high; on a job kind, beside its timeout_ms: a free worker takes the most urgent due row, and among equals the one that has been due longest. .high, .normal (the default, and what every existing kind gets) and .low. The case it came from is a queue where a model backfill running for minutes stood in front of the cache revalidations a person was waiting on — not because the machine was busy, but because the backfill was in front. Per kind rather than per push, because how urgent a kind is belongs to the kind; three levels rather than a number, because 2 says nothing about whether it beats 1, and a kind that writes one is a Refusal naming the levels (ADR 214).
  • run.loop(): the Io a Run was made on, or null for a Run.init(gpa), for a job that writes a file or sleeps between attempts and has only the Run the worker handed it.

And the rest of the server

  • nilo.Gate serves its waiters in the order they came, and gate.enterWithin(ms) waits at most ms (ADR 222). A turn given back while someone waits is handed to the oldest of them by name, instead of going back where the caller that just left could take it again first; enterWithin gives up with error.TimedOut, holding nothing and owing no leave, and enterWithin(0) asks whether a turn is free now. Nothing to change: enter and leave keep their shape.
  • jwt.Keyring takes remember_tokens: how many verified tokens to remember by digest, so a bearer token seen again skips the signature arithmetic and keeps every claims check. A new key set forgets them (ADR 209).
  • app.health and app.metrics describe the routes they register, a 200 each, and are no longer counted in the "N of M routes hold the Ctx and return nothing" line, which is about the application's handlers (ADR 120).
  • app.failures(T): the body every failure goes out with, when nilo's {"error":…,"status":…} is not the one your clients already read. T is a struct whose fields are the JSON, with a pub fn nilo_failure(status: u16, message: []const u8) T that fills it; nilo writes it with the JSON writer a handler's answer goes through, and the API description's Failure schema comes from the same fields. A fail function's sentence, the 404 and 405 nilo answers itself, a 401's challenge, a 405's Allow and the CORS headers all survive it. The five answers written before there is a request to route — a malformed head, a head too long or too slow, an unreadable coding, a shed 503 — keep nilo's own. Nothing changes for an App that does not call it. Three refusals (ADR 024).
  • Every response carries a Date, second after the status line — RFC 9110 §6.6.1's MUST, which nilo had never met, and what a cache in front does its freshness arithmetic from. Formatted once a second per thread, lazily, from nilo_core's clock; no task, no atomic, no allocation. A Date a handler sets wins. In the same head, Connection: keep-alive is no longer written on an HTTP/1.1 response — persistence is what HTTP/1.1 means, and the line is now written only when it says something: keep-alive to an HTTP/1.0 client being kept, close to anybody being closed. The benchmark response goes from 1,110 bytes on the wire to 1,123, which is what every other server in bench/compare/ sends for the same body. Ctx.connection() is the new way to ask; Ctx.keepAlive() still answers the bool (ADR 197).
  • listen() takes request_deadline_ms: a deadline every request starts with, what nilo.deadline(ms) gives one route given to all of them. Every wait nilo owns is cut to it and c.overdue() reads it; a route's own nilo.deadline replaces it, and a request that takes the connection over — a stream, a WebSocket, bodyStream() — lets a default go and keeps a route's own. Off by default. The option had been rejected once, and why it is back is in the same ADR (ADR 105).
  • Every crafted request in the parser's tests — the framing conflicts, the strict chunk sizes, the absolute-form target, the head that never ends — is now also run split at every byte and trickled a few bytes a read, and has to come out identical to the same bytes arriving at once, down to where the next request starts. The parser's own tests only ever read from a buffer holding the whole input, so every seam that resumes across a read boundary was untested at exactly the boundary. http/http1.zig, one test.

Fixed

  • A content type carrying a line break is refused like any other header value. c.send, c.streamWith and a FileBody wrote their content type into the head unchecked, so a MIME type stored from an upload or passed on from upstream could add headers of its own: c.send(200, "text/plain\r\nX-Injected: 1", …) sent X-Injected. It is a 500 naming the problem now, the answer setHeader gives (ADR 029).
  • An event stream can no longer be made to deliver an event nobody sent. data was split on LF alone, but a browser also ends a line at a lone CR, so events.data("a\rdata: forged\revent: admin") delivered an event named admin to every subscriber. data and comment are split where the browser splits them, and a CR or LF in an event's name or id is error.EventFieldBreaksLine, with nothing written.
  • A cookie's path, domain and expiry are checked for ; and control bytes, as its value always was. A path built from the request's own turned /x;Domain=example.com into a cookie sent to every subdomain; setCookie and a session's cookie refuse it with a 500 now (ADR 029).
  • The /docs page loads one version of its viewer, held by its hash. It loaded @scalar/api-reference with no version and no integrity into the application's own origin, where the script runs beside the session cookie, so whatever the package published next would have run there. It is 1.72.0 with a sha384 now (ADR 016).
  • A JSON body whose type holds itself can no longer take the process down. std.json reads a type like a comment tree (c: []const Node) by recursing once per level, with only the fiber's stack to stop it, and {"c":[ twenty thousand times, 160 KB and inside max_body, was a segfault that ended every connection. For such a type a body nested more than 64 levels deep is a 400 now, found by one pass over the bytes before the parse; a type that cannot hold itself is bounded by its declaration, is not scanned, and costs nothing new.
  • A fail function in work started with nilo.spawn, or in a job worker, no longer writes into an unrelated request. Every request set the threadlocal fallback slot on its executor thread and left it there across its suspensions, which is exactly what the comment on it said must never happen: a spawned fiber, which has no slot of its own, read it. It is set now only when App is called with no Engine underneath, and setFallbackSlot refuses in Debug to be called from a fiber that has a slot (ADR 006).
  • A WebSocket that joins a second nilo.Room, or never leaves one, no longer leaves a seat ringing a finished connection. Joining a second room overwrote the socket's ticket, leave then gave up a seat by an index into the wrong room, and the first seat stayed taken with its bell in a frame that had returned, so the next say wrote into freed stack. Joining a second room is error.AlreadySeated now, joining the same one twice does nothing, leave on a room the socket is not in does nothing, and nilo gives up any seat still taken when the loop returns.
  • Routes are found by a tree rather than a scan, so a large table under one prefix no longer costs a third of a request. On a real application's 276 routes, all under /api, a match went from 125 ns to 27 ns on average and from 250 ns to 54 ns at worst; a one-route app went from 19 ns to 13 ns, because the same change stopped copying a 300-byte Match out of every match. Two answers change, both where a * meets a route of another length, and both now follow the rule ADR 012 always stated: /a/* wins over /:x/b/c for /a/b/c (the earlier literal), and /files wins over /files/* for /files whichever was registered first. zig build profile -- --routes <file> measures matching on a table of your own.
  • A slow outbound call through nilo_fetch or nilo_s3 is no longer reported as the handler holding its thread. The watchdog warned "handler … held its thread for …ms" and advised nilo.blocking for a call whose fiber was parked on the socket the whole time, the false report ADR 210 had already fixed for SQL. Every step of a call that waits (the permit queue, the head, each body read, the drain) now says so through the Limits the client was started with. A handler's own work between the pieces of a streamed body is still watched. Nothing is added to the Exchange on the handler's stack.
  • A statement cut off by a cancellation hands the cancellation back (ADR 223). nilo_sql turned error.Canceled into QueryFailed (or Disconnected from a pool wait) and the cancellation was gone, so a background loop whose nilo.sleep(..) catch return is its only way out logged the failure and slept on: SIGTERM during a statement left the process running for good. The cancellation is now re-armed, connections go back to the pool and transactions roll back with cancellation held off, and a rollback a cancellation made impossible is no longer logged as a failure. The errors a caller sees are unchanged.
  • nilo.testing.Client.send copies the request before handing it to the App, so a WebSocket test written with a string literal no longer crashes on macOS. A message is unmasked in place in the read buffer, and under LLVM, which a Mac builds with, a literal sits in a read-only page; the write was a bus error there and passed on Linux, whose Debug backend puts literals in writable memory.
  • A request body sent with Content-Encoding: gzip by a writer that flushes before it closes, the way Go's does, is read. Such a stream ends in an empty final block, and the decompressor asks for a byte of room before reading even that one, so every such body was refused as one that could not be read.
  • bucket.list against a real server read a key with a space in it back with a + in its place (b c.txt came back as b+c.txt), and against MinIO handed back an ETag that getIf never matched. Under encoding-type=url AWS and MinIO both write a space as + and a plus as %2B, and the key was decoded as a path; MinIO's XML is Go's, which writes a quote as &#34;, and only &quot; was read. Both were found the first time CI ran nilo_s3's live tests against a real MinIO rather than skipping them.
  • A WebSocket that lives for a few messages no longer maps and unmaps a message buffer. A message that arrives whole is handed over from the connection's read buffer; with more short-lived sockets open on a thread than its free list keeps, every connection used to mmap 16 KiB and munmap it again, and each munmap stopped every core the process runs on to flush its TLB. Message.data is borrowed until the next receive, as before. A busy socket whose messages fit the read buffer also holds no message buffer now, about 4 KiB less resident a connection (ADR 216).
  • pg.zig is pinned at ec8cf27, lalinsky's master rebuilt on karlseguin's. The old pin, 91d0705, had been left on no branch by that rebuild, so a cold -Dsql fetch depended on GitHub still serving an unreachable commit. The move brings upstream's fixes: an authentication error is copied before the reader lets go of it, a failed query always hands its pooled connection back, and a numeric with leading zero groups prints correctly (nilo reads numeric as text, so the last never reached a sql.Decimal).
  • The line a SQLite statement writes when it gives up waiting for a connection names the statement holding it, as in it is held by SELECT 1. It used to guess at a tx waiting for itself, and sent the one application that hit it looking for a transaction it did not have; what held the writer was a statement queued for a thread behind a slow read (ADR 107, zio#745).
  • A fail function called by work registered with app.before, a seed calling the same service functions its handlers call, lost its sentence: the boot said failed with Failed and nothing else. The line that stops the boot now carries the status and the message, and which of the registered pieces failed, on both listen() and app.start(io) (ADR 129).
  • nilo-dev started the binary left in zig-out before its build had finished, so after a change made with the loop stopped the old server ran first: one application had its SQLite file created and seeded with the schema it had just changed. The loop now builds once to the end before it starts anything, and when that build fails it removes the stale binary and starts the first one that compiles. Nothing changes in a dependent's build.zig (ADR 190).
  • A response type as wide as a detail page, a record holding lists of records of twelve fields or so, failed to compile inside http/json.zig with "evaluation exceeded 1000 backwards branches", whatever its depth, and the advice to raise the quota was not something the application could do. covers raises it where the walk starts, to 20,000, room for some 2,500 fields.
  • A worker spun forever on a job kind it did not know. A row pushed by another binary — an older deploy, a sibling service — was claimed, found to be of an unknown kind, and handed back at the run_at it already had; so the same worker claimed it again on the next turn, forever. One such row left behind by a removed kind ran up 88 210 claim/release pairs, kept a worker permanently busy and timed out /health. Worse than the spin: the claim counts an attempt, so a row nobody here could run was walking towards dead in the binary least able to judge it. The claim now asks only for the kinds the program knows, so the row is left queued, untouched and at nought attempts, for the binary that does know it (ADR 215).
  • Every binary with a static set in it, which is every binary with the API reader page, was 25 KB larger than it needed to be: the Accept-Encoding reader answered "is this q=0" with std.fmt.parseFloat(f32, …). A digit scan now; hello is −24,608 bytes stripped, rest −3,648 (ADR 211).
  • A Db on the default connect_on_init = 0 dials one connection at boot whether or not it has a schema check, so app.before — a migration, a key set — finds a pool with something to lend. An unchecked Db with a before hook got Disconnected on every cold boot with the database up (ADR 115).
  • The comptime count of a raw select list stopped at WITHIN GROUP, reading its GROUP as GROUP BY, so percentile_cont(0.5) WITHIN GROUP (ORDER BY v) AS median, count(*) AS n counted as one column and the Row with two fields was refused. GROUP and ORDER end the list only with their BY.
  • A database round trip over block_warning_ms no longer draws "handler held its thread": core.Limits gained waiting/waited, the Engine routes them to the watchdog, and the Postgres wire reports every statement through them, once from the exchange to the result's close. A handler that computes without parking is still reported (ADR 210).
  • app.tryStatic and app.tryStaticWith on a directory that is not there hand back error.StaticDirNotFound and log nothing; the error: line belonged to static, which stops the process on it. A problem inside a directory that is there is still said in one line, since the error cannot name the file (ADR 207).
  • A WebSocket client that leaves with a reset rather than a FIN, which is every load generator that keeps its ports out of TIME_WAIT and every tab that was killed rather than closed, ends receive with null the way a FIN does, instead of error.ReadFailed and a warning per connection. The warning was one lock every connection queued on to leave; at 70,000 connections a second it held reset sockets open to max_connections, and the server began refusing at accept. 512 WebSocket connections closed after ten frames each: 461K → 1.68M frames a second, descriptors mid-run 10,024 → 560 (ADR 202).
  • A server that is not busy spends a third less CPU per request. zio's scheduler dozes for 100 µs before each park so that work stealing does not churn, and on a thread with nothing coming that is a second context switch per request: 100 µs of CPU a request at 500 req/s on two threads, 70 with stealing off, and +3% at saturation. Stealing is now off, and a handler runs on one OS thread from its first line to its last, across every wait in it. bench/paced.py is the instrument (ADR 199).
  • c.clientIp() reads every X-Forwarded-For field, as one list in wire order, rather than the first. HAProxy's option forwardfor adds a field of its own instead of appending to the client's, so a forged header arrived as two fields with the forgery first — and with .trusted_proxies set, the walk started from the forgery and returned it. nginx appends, which is why the tests passed. Both the rules and .trusted_hops now walk proxies.Forwarded, from the last field's right end; more than eight fields is answered with the socket's address. No allocation (ADR 102, the closing section).
  • A client that connects and gives up before the server reaches it in the backlog no longer stops the server. zio v0.17.0 surfaced that as error.ConnectionAborted from accept, and the accept loop returned on anything but a timeout — so one aborted connection ended listen() with a clean "nilo stopping" in the log. The pin is v0.18.0, whose accept retries it on the same deadline. Rare on Linux, which usually hands the socket over and fails the read instead; the ordinary path on the BSDs. The same bump takes the BroadcastChannel fix the roadmap was waiting on, and lets the Engine hand a fired completion straight back to submit instead of rebuilding it first (zio#673, fixed by zio#674).
  • A chunked request body nobody read no longer panics on a chunk size that overflows a u64. It was added to the running total before the total was checked, so ffffffffffffffff after any earlier chunk overflowed — a crash in a safe build, a wrapped limit in a fast one. It is refused on the announced size now, before a read, the way a buffered chunked body already was.
  • A chunk size is read as strict 1*HEXDIG rather than through a lenient integer parse. +5, 1_0 (which read as 16), and a size with leading or trailing whitespace were accepted, each a length a front end could frame differently — the request-smuggling shape a duplicated Content-Length is.
  • Whitespace between a header field name and its colon — Content-Length : — is a 400 rather than a line that is silently dropped, which RFC 9112 §5.1 requires and which closes the same framing disagreement.
  • A Content-Length body being discarded to reuse a keep-alive connection is bounded by max_body, and a body over it closes the connection instead of being read in full. The drain path ignored the limit the handler path enforces, so a body larger than the server would ever accept was read only to be thrown away.

Docs

  • The guide is a website: nevindra.github.io/nilo, one copy per minor release, published when the release is tagged (ADR 219).
  • Getting started says what a save has to touch for zig build dev to restart the server: a .zig the binary is built from, or a file it @embedFiles, and nothing else in the checkout, so a front end kept beside the server keeps its own dev server. The static-files guide and decided.md said the loop could watch a bundler's directory; it cannot, and both now say so. bench/devloop.py is the check: a save the build never reads must leave the server up, a save it reads must restart it. The dev step on that page also gained the line that forwards b.args, without which the zig build dev -- --incremental beside it never reached nilo-dev.
  • Past one table says what a raw parameter may be (an optional binds NULL, and ($1 IS NULL OR …) is the sql.given of raw SQL), has a section on reporting statements (rawExactlyOne, raw per group, rawPage for a paged join, dates per dialect) and one on what SQLite does differently. SQLite has the strftime(col / 1000000, 'unixepoch') recipe a microsecond Timestamp needs.
  • Handlers has the table of legal return shapes and the refused ones beside it. The App reference says which calls return an error and which return nothing.
  • Getting started and the README say what to pass when the native link fails at crt1.o:.sframe on a glibc built by GCC 16: -Dtarget=x86_64-linux-gnu or -Dllvm.
  • Static files says why .reload does not pick up a bundler's hashed filenames, and the two ways round it. Writing says a Str column takes a literal, a []const u8, a []u8 or a Str on insert.
  • Ctx.body() says that a gzipped body comes back inflated while header("Content-Encoding") and header("Content-Length") still describe the wire, because the head is read in place and nothing rewrites it — and what a handler forwarding the body should send instead. ADR 089 carries the same note.
  • Deploying has one table for every bound listen() takes: what a client sees past it, what the log says, and what has to happen before the server takes that work again. The prose under it was already there; the lookup was not.

Where to read next

The guide is a website now, one copy a minor release: nevindra.github.io/nilo. Every module has a page under docs/guide/, and the whole public API is one page a module under docs/reference/. What is still open, the defects the line-by-line read found and has not fixed among it, is docs/roadmap.md.

The full list, every entry with its ADR, is in CHANGELOG.md at v0.6.0. What was measured and what turned out false on the way is in docs/history.md.