Skip to content

Coroutines as State Machines

Brian Szmyd edited this page Jun 9, 2026 · 3 revisions

Coroutines as State Machines (and how Fibers differ)

A companion to Folly to Coroutine Migration. That page tells you what changed; this one gives you the model underneath it — what the compiler actually does to a co_await function — so the two things that page cares about, the heap frame (allocation) and the resume (threading), stop being magic. If you've used folly::fibers, the last section maps it onto the same vocabulary so you can see exactly where coroutines and fibers diverge.

The one rule that explains everything

A C++ coroutine is stackless. The compiler cuts the function apart at every co_await and rewrites it as a state machine: a struct holding your live locals plus a "where was I" index, and a resume() that jumps back to the right spot. When a coroutine suspends, its stack frame is gone — the call stack unwinds back to whoever resumed it. The only thing that survives is one heap allocation: the coroutine frame.

That single fact is the source of both pillars:

  • the frame is what gets heap-allocated (Pillar 1), and
  • resume() is just a function call, so it runs on whatever thread calls it (Pillar 2).

A two-state example

One suspension point splits a coroutine into two states: before the await and after it.

async_result<size_t> write_one(volume_handle vol, uint64_t addr, sg_list sgs) {
    auto r = co_await async_write(vol, addr, std::move(sgs));   // the only suspension point
    co_return r ? r.value() : -EIO;
}

What the compiler lowers it to

A faithful sketch — not literal output (real frames also carry initial/final_suspend, a destroy path, and HALO may elide the new) — but every moving part is here:

// (1) THE FRAME — heap-allocated; outlives every suspension.
//     Every local live across a co_await becomes a MEMBER, not a stack variable.
struct write_one$frame {
    int                     state = 0;       // "where was I" — the resume index
    promise_type            promise;         // owns the eventual result<size_t>
    std::coroutine_handle<> continuation{};  // who resumes when WE finish (set by our awaiter)

    volume_handle  vol;                       // params + live locals, spilled onto the frame
    uint64_t       addr;
    sg_list        sgs;
    write_awaiter  awaiter;                   // the sub-op we're suspended on
    result<size_t> r;
};

// (2) THE RAMP — the function you actually call. Allocates the frame, copies args in,
//     hands back the lazy task. It does NOT run the body.
async_result<size_t> write_one(volume_handle vol, uint64_t addr, sg_list sgs) {
    auto* f   = new write_one$frame{.vol = vol, .addr = addr, .sgs = std::move(sgs)};
    auto  ret = f->promise.get_return_object();   // the [[nodiscard]] task handed to the caller
    // sisl::async tasks are LAZY: initial_suspend() == suspend_always, so we return here without
    // executing. state 0 runs on the FIRST resume() — i.e. when the caller co_awaits `ret`.
    return ret;                                   // (an *eager* coroutine would resume(f) right now)
}

// (3) THE BODY — turned inside-out into a resumable switch. resume() re-enters HERE.
void write_one$resume(write_one$frame* f) {
    auto self = std::coroutine_handle<promise_type>::from_promise(f->promise);
    switch (f->state) {
    case 0:
        f->awaiter = async_write(f->vol, f->addr, std::move(f->sgs));
        f->state = 1;                             // ← set the REENTRY POINT *before* suspending
        if (!f->awaiter.await_ready()) {
            f->awaiter.await_suspend(self);       // give the sub-op OUR handle as ITS continuation
            return;                               // ★ SUSPEND: unwind to the caller/reactor.
                                                  //   the C stack is gone; only *f survives.
        }
        [[fallthrough]];                          // ready-already fast path: never left the thread
    case 1:                                       // ★ the COMPLETING THREAD re-enters here via resume()
        f->r = f->awaiter.await_resume();         // ← FRAME RELOAD: read locals back off *f
        f->promise.return_value(f->r ? f->r.value() : -EIO);
        auto k = f->continuation;                 // whoever awaited us
        delete f;                                 // frame + its locals freed
        if (k) k.resume();                        // ← CONTINUATION runs — ON THIS THREAD
        return;
    }
}

The three things to walk away with (marked ★ / ← above):

  1. The state-machine jump to the reentry point. f->state = 1 is set before the suspend, so the next resume() falls through the switch straight to the line after the co_await. That's "jump back to where I was."
  2. The frame reload. Locals don't survive on the stack across a suspend — the stack is gone. They live on *f, and resume reads them back (f->awaiter.await_resume(), f->r). This is why the migration page keeps saying "anything referenced across a co_await must live on the frame or the heap." The compiler spills the locals it can see onto *f for you — but a reference into a dead temporary's scope is on you (that's the sgs_keepalive heap-own in the read fan-out).
  3. The continuation. await_suspend(self) hands our handle to the sub-op as its continuation; when the sub-op finishes it calls self.resume() → us at state 1 → we finish and call our continuation k.resume(). A chain of "when you're done, call me." And because k.resume() is an ordinary call, the continuation runs on the thread that completed the sub-op — no executor, no hand-off. That is sisl::async's sticky affinity, and it's why a write that completes on the commit reactor finishes on the commit reactor.

What's an "awaitable"? — sisl::async::disk_task decomposed

The lowering above called three methods on f->awaiter (await_ready / await_suspend / await_resume) and returned suspend_always from initial_suspend. Those aren't sisl inventions — they're the standard awaiter protocol every co_await speaks. sisl::async::disk_task<T> (sisl/include/sisl/async/disk_task.hpp) is the smallest real thing in our tree that shows the whole vocabulary at once, because it is both a coroutine (it carries a promise_type) and an awaitable (you can co_await one). Decomposing it covers every "funny type" you'll see.

It implements two separate protocols

Role Methods Who calls them
It is a coroutine promise_type::{get_return_object, initial_suspend, final_suspend, return_value, unhandled_exception} the compiler, to run this task's body
It is an awaitable disk_task::{await_ready, await_suspend, await_resume} the compiler, when another coroutine does co_await someDiskTask

A type needs only the first set to be a task, only the second to be co_await-able. disk_task has both — which is exactly why one disk_task can co_await another.

The awaiter triple — what runs at a co_await

When some coroutine writes co_await child and child is a disk_task<T>, the compiler emits exactly these three calls on it:

bool await_ready() const noexcept { return false; }            // (1) skip the suspend? no — the I/O hasn't run yet
std::coroutine_handle<> await_suspend(std::coroutine_handle<> cont) noexcept {
    _coro.promise()._continuation = cont;                      // (2) remember who resumes us when WE finish
    return _coro;                                              //     ...and symmetric-transfer INTO the child now
}
T await_resume() noexcept { return _coro.promise()._value; }   // (3) resume lands here; hand back the result
Method Returns Job in disk_task
await_ready() bool "Can we skip suspending?" false ⇒ always park (I/O still pending). hot_task instead returns _coro.done() — the fast path for I/O that already completed (the [[fallthrough]] in the lowering).
await_suspend(cont) void / bool / handle Runs after the caller's frame is saved. Stores cont (the caller's handle) as the child's _continuation, then returns a handle — see symmetric transfer below.
await_resume() T Runs on resume (or the ready fast-path). Its return value is the value of co_await child — here _coro.promise()._value.

suspend_always / suspend_never — and where they plug in

These are the standard library's two no-op awaiters: they carry no value and just answer await_ready one way. You rarely co_await them directly — you return them from the promise's initial_suspend() / final_suspend(), the two awaits the compiler injects at the top and bottom of every body.

await_ready() meaning disk_task uses it for
std::suspend_always false always suspend initial_suspend()lazy
std::suspend_never true never suspend (not used here — that would be an eager task)
  • initial_suspend() → std::suspend_alwayslazy. The body doesn't run on construction; it runs on the first resume(). That's what start() triggers (_coro.resume() advances it to the first SQE submission), and it's what enables fan-out: start() every child to submit all the SQEs, then co_await the results.
  • final_suspend() returns a custom final_awaiter, not suspend_always — because at co_return the frame must not self-destruct yet; it has to hand control back to the awaiter. That's the second symmetric transfer.

Symmetric transfer — the coroutine_handle-returning await_suspend

The powerful variant of await_suspend is the one that returns a coroutine_handle: it means "instead of returning to whoever resumed me, immediately resume this handle" — a guaranteed tail call, no stack growth, no scheduler hop. disk_task uses it at both ends of a call:

// (A) ENTERING a child — co_await child: start the child right now, no scheduler bounce
std::coroutine_handle<> await_suspend(std::coroutine_handle<> cont) noexcept {
    _coro.promise()._continuation = cont;
    return _coro;                                         // ← resume the CHILD
}

// (B) LEAVING a child — at co_return, hand control straight back to the caller
struct final_awaiter {
    bool await_ready() noexcept { return false; }
    std::coroutine_handle<> await_suspend(std::coroutine_handle< promise_type > h) noexcept {
        auto cont = h.promise()._continuation;
        return cont ? cont : std::noop_coroutine();       // ← resume the CALLER (or stop the chain)
    }
    void await_resume() noexcept {}
};

So a chain A → co_await B → co_await C threads continuations down (each await_suspend returns the child) and tail-resumes them up (each final_awaiter returns the stored _continuation), entirely without a scheduler — the "resumes the caller via final_suspend symmetric transfer without any scheduler round-trip" the file's header comment promises. std::noop_coroutine() is the standard "resume nothing, the chain ends here" handle.

Lifetime — the disk_task object is the frame's owner

disk_task is move-only and owns its coroutine handle: the move ctor std::exchangees it away, and the destructor calls _coro.destroy(). So the lowering's rule — the frame outlives every suspension and is freed exactly once — is enforced here by the disk_task's own RAII, not by hand. hot_task<T> is the already-started sibling start() returns: same three await_*, but its await_ready() is _coro.done(), so an I/O that completed synchronously never suspends at all.

Many awaits, one frame — fewer allocations than .thenValue

A coroutine with N sequential co_awaits is still one frame. It's sized once, at compile time, to the peak set of locals simultaneously live across a suspend — non-overlapping locals share storage, and a local not live across any suspend never lands on the frame at all. More await points just add case labels to the switch; they add no runtime allocations. The frame does not grow.

That is where the migration cut allocation calls. In Folly the continuation chain is heap objects — each .thenValue is a new Future backed by its own heap-allocated, atomically-refcounted shared state (Core), plus storage for the callback:

return write_data(vol, addr, sgs)                       // Future  <- Core #1
    .thenValue([=](auto){ return write_index(...); })   // Future  <- Core #2  (+ callback)
    .thenValue([=](auto){ return append_wal(...);  })   // Future  <- Core #3
    .thenValue([=](auto){ return ok();             });  // Future  <- Core #4
// 4 stages ~ 4 Core allocations, all refcounted, + the request object.

In coroutine-land the continuation chain is state indices on the one frame — zero allocation per step; the co_awaits are just more states in the switch:

async_result<size_t> write(vol, addr, sgs) {   // <- ONE frame, allocated once
    co_await write_data(...);                  //   state 0 -> 1   (a jump, no alloc)
    co_await write_index(...);                 //   state 1 -> 2   (a jump, no alloc)
    co_await append_wal(...);                  //   state 2 -> 3   (a jump, no alloc)
    co_return ok();
}
// 4 steps -> 1 frame. The glue between steps is a switch jump, not a heap callback.

So: Folly allocation calls scale with the number of async steps (a Core per .thenValue); coroutine allocation calls scale with the number of coroutines (one frame, however many co_awaits inside). Fold more steps into a coroutine and you strictly cut allocation calls.

A chain of coroutines is still a frame per level — and HALO can't fuse them

The caveat: that win is within a coroutine. A chain of separate coroutine functions — async_write awaits volume::write awaits the data service — is one frame per coroutine, linked by continuation handles. That's frame-per-coroutine vs Folly's Core-per-stage: closer to even on the chain itself, so the win is concentrated in how many steps you fold into each level.

You might hope the compiler fuses the chained frames (HALO — coroutine frame elision). It can't here, and the reason is worth stating precisely because it's neither a defect nor a tuning knob: HALO requires proving a coroutine starts and finishes inside its caller's scope without the handle escaping. The HomeBlocks I/O coroutines do the opposite — they suspend on one thread and resume on another. The evidence is in sisl::async::value_awaitable (what co_await req->promise_ lands on):

// Producer side (ANY thread). Resumes the consumer iff it already suspended.
void complete(T value) noexcept {
    _result.emplace(std::move(value));
    if (_state.exchange(k_done, std::memory_order_acq_rel) == k_waiting) { _waiter.resume(); }  // foreign thread
}

co_await req->promise_ stores the frame's handle in _waiter and suspends; the caller's stack unwinds and returns; later on_write, on the commit thread, calls complete()_waiter.resume(). The frame outlives its caller's activation and resumes on a foreign thread — the literal definition of the handle escaping, and an escaped frame cannot be elided into its caller. Heap allocation here is mandatory, not a missed optimization. (sisl::async::task is stdexec's exec::task; there's no custom operator new and no opt-out trickery — the frames simply outlive their callers.)

If frame-allocation pressure ever shows up in a profile (it didn't in our A/B — the I/O path is throughput-bound, and tcmalloc makes small allocations cheap), the levers are not HALO:

  • Cut the frame count — flatten pure forwarders. A coroutine that only does co_return co_await inner(...) is a wasted frame; make it a plain function that returns inner(...). Only coroutines that do real work before a suspend (setup, the sgs_keepalive heap-own) need to be coroutines at all.
  • Make the unavoidable frames cheap — a pooled operator new. A promise can allocate frames from a per-thread pool instead of malloc; lowers per-alloc cost without changing HALO's reach. sisl uses default new today.

The same idea as a Folly fiber

A folly::fibers task reaches the same place — linear code that suspends mid-function and resumes later with locals intact — by the opposite mechanism: it is stackful.

folly::fibers::FiberManager& fm = ...;             // runs on an EventBase
fm.addTask([vol, addr, sgs] {                      // runs on a fiber with ITS OWN stack (~256 KiB)
    folly::fibers::Baton baton;
    size_t r{};
    start_async_write(vol, addr, sgs, [&](size_t n){ r = n; baton.post(); });
    baton.wait();      // ★ SUSPEND: jump_fcontext saves SP + callee regs and switches back to the
                       //   FiberManager loop. The fiber's WHOLE stack is left frozen in place.
    use(r);            // ★ RESUME: baton.post() marks the fiber ready; the loop jump_fcontext's back —
                       //   restoring SP + regs — and execution continues here, the entire stack intact.
});

There is no compiler transform and no per-suspend frame struct. The "state" is the entire native stack — all locals and the whole chain of nested calls beneath you — preserved byte-for-byte. Suspend/resume is a register + stack-pointer swap (boost::context::jump_fcontext); a FiberManager is the runtime that owns the ready/blocked fibers and their stacks.

The consequence that matters: because the whole stack is preserved, a fiber can suspend from any call depth — even inside a plain function, or third-party code, that has no idea fibers exist. A coroutine can only suspend at its own co_awaits; a normal function it calls cannot yield. Fibers buy that flexibility with a full stack per in-flight task — reserved up front, fixed-size: memory-heavy at high concurrency, and a stack-overflow risk if you under-size it.

Side by side

C++ coroutine (sisl::async) Folly fiber
Stack model Stackless — one heap frame of live locals Stackful — a full native stack per task
Preserved across suspend Just the spilled locals on *frame The entire stack (all locals + nested calls)
Suspend transform Compiler-generated state machine (switch) None — runtime register/SP swap (jump_fcontext)
Where can you suspend? Only at this coroutine's own co_await Anywhere, at any call depth
Resume is… A call that jumps to the reentry state Restore registers + switch stacks
Cost per in-flight op One right-sized frame (tens–hundreds of bytes) One reserved stack (~256 KiB default)
Scheduler None — the awaiter chain is the schedule A FiberManager runtime
Thread on resume Whatever thread calls resume() (sticky affinity) Whatever thread runs the FiberManager loop

That bottom-left-vs-right is the allocation story (Pillar 1) in miniature: the migration swapped "a full fiber stack, or a heap Future Core per .thenValue" for "a right-sized coroutine frame per coroutine in the chain."

Back to HomeBlocks

Re-read the Folly to Coroutine Migration write path with this in hand and it decodes cleanly:

  • co_await rd()->async_write(...) — a frame is allocated; this is one suspension index; it usually resumes on the issuing reactor (the io_uring CQE lands there).
  • co_await req->promise_ — the frame parks; on_write on the commit thread calls the continuation's resume(); per rule (3), the rest of volume::write runs on the commit thread.
  • the per-sub-read sg_list kept alive in sgs_keepalive — a local the compiler can't keep on the frame for you (it's captured by-reference into a lazily-run task), so you heap-own it explicitly.