core-cpp 0.2.0
Patch-level fixes and the API fastcached needed to move onto core-cpp, found by migrating it: the MSVC ARM64 exception-table defect (fastcached#1546), port sharing, socket options on accepted sockets, adoptSocket, static-CRT variants, a closed-waiter idle bug, and per-request parity with fastcached's own reactor. Minor version because of one Breaking entry and new API.
Breaking
-
The
await_readyof seven public awaiters answers aconstexprconstantfalse:
core::net::DelayAwaiter, the awaiterwhenAllandwhenAnyreturn,AsyncQueue<T>::PopAwaiter,
and the fourTuiRuntimeawaiters (which also becamenoexcept). It is the fix for
fastcached#1546 under Fixed: the decision it made isawait_suspend's now. (Task's two
awaiters answer as before -- true for a task owning no frame or already finished -- but from a
member their constructor sets, and are not on this list.) Aco_awaitbehaves as before; code that calledawait_ready()
directly to learn whether an await would park getsfalsewhere it gottrue-- for instance
sleepUntil(nullptr, t).await_ready()-- and core-cpp's ownSleepUntil_test.cppand
AsyncQueue_test.cppasserted exactly that. It stays aconstmember rather than astatic
one, which clang-tidy'sreadability-static-accessed-through-instancewould report at every
co_awaitin a caller's code.ResultAwaitable::await_readystill answers whether the operation
settled inline.Migration:
co_awaitthe awaiter, and never callawait_ready()on it directly; it is the
compiler's half of the protocol, not a question a caller can ask. A test that asserted an await
resolves without parking asserts it through the flow instead: that it has finished after one
turn of the loop, or that its continuation ran exactly once. fastcached'sSleepUntil_test.cpp
assertstrueon its own copy, and changes with the migration.
Fixed
-
EventLoop::runUntilIdleandtesting::TestLoop::drainno longer return while the waiters of
a closed handle are still queued.notifyHandleClosingrecords the parks on a closing handle
and the next turn queues their waiters after its wait; that turn still reported itself idle,
because none of its counters counted them, so a drain returned with a closed listener's pending
accept-- or any flow parked throughwaitReadable/waitWritable-- woken and not yet run. A
caller that tore its objects down next left~EventLoopto resume or free those frames after
their owners were gone: fastcached's server teardown reported it as a heap-use-after-free under
ASan and TSan. A turn is now idle only if it also leaves the ready queue empty, which covers
every step that queues work rather than the four that had a counter.ClosedParkIdle_test.cpp
closes a listener under a parked accept and asks onerunUntilIdleto finish it. -
An
OperationCancelledthrown out of aco_awaitondelayorsleepUntil, and on twelve
other awaiters, is caught again under MSVC 19.44 on ARM64
(fastcached#1546). That compiler's
ARM64 code generator drops the enclosingtryof aco_awaiton a temporary awaiter whose
await_readymakes a call, so the exception passed a typedcatchandcatch (...)alike and
reached whatever awaited the flow. The newwindows (cl-release-arm64)leg then found the same
loss with no call at all: anawait_suspendthat returned the awaiting coroutine's own handle,
followed by anawait_resumethat threw --Task's awaiter over a task owning no frame, whose
std::logic_errorpassed the awaiting coroutine'scatch. SoTask's two awaiters decide in
their constructor, where the awaiter is made by theco_await, andawait_readyreads the
answer.ResultAwaitable, every socket operation's awaitable, transfers back into an
OperationCancelledfor a flow already stopped too; on the same leg it keeps its handler, which
CancelRead_test.cppnow holds.core::net::DelayAwaiter::await_readyread the clock through
the virtualIClock::now(). The 0.1.0 notes say this awaiter carried fastcached's fix when
TuiRuntimemoved onto it; it did not, and everydelayandsleepUntilhad the shape. Every
await_readyinsrc/corenow reads a member or answers a constant, and the decision it made is
the constructor's orawait_suspend's, asked before the flow's stop token is read, so what aco_awaitobserves is
unchanged (a direct call ofawait_ready()is not; see Breaking): an elapsed deadline, a null loop, a finished task, an emptywhenAll, a queued item,
a free TLS gate, a finished lookup and a buffered input event or agent message all resume without
parking, and a flow that is already stopped still resumes normally on each of them. The awaiters:
DelayAwaiter,interruptibleSleepUntil's,Task<T>::AwaiterandTask<void>::Awaiter,
whenAll's andwhenAny's,AsyncQueue::PopAwaiter,ResultAwaitable, the threaded resolver's
and the TLS serial gate's, andTuiRuntime'sNextInputEventAwaiter,NextEventForAwaiter,
NextActivityAwaiterandNextAgentReadyAwaiter. -
A socket a listener accepts carries
TCP_NODELAY, as a dialled one always did. Only the
dial paths set it (posix/DialPrimitives.cpp,windows/DialPrimitives.cpp, and through them the
completion-port dial);accept4/accepton POSIX, the WFMO listener'sacceptand the IOCP
listener'sAcceptExhanded out sockets with close-on-exec and nothing else, so a server's small
replies waited on Nagle for the client's delayed ACK. Every dial and every accept on every
platform now goes through one helper,detail::applyStreamSocketOptions, which sets close-on-exec
(a non-inheritable handle on Windows),TCP_NODELAY, and keepalive when a dial asks for it; the
buffer sizes below aredetail::applySocketBufferSizes's, before the connection exists.WindowsSocket::native()joinsPosixSocket::native()and
IocpSocket::native(), for diagnostics and tests. -
The WFMO backend's TCP listener claims its port exclusively. It bound with
SO_REUSEADDR,
which on Windows lets a second socket bind a port a live listener serves and take its
connections, where the IOCP listener has always usedSO_EXCLUSIVEADDRUSE. It uses that too
now, and fails the bind if the option is refused.BackendKind::Wfmois not the Windows
default, so only a caller that asked for it by name was exposed. -
testing::TestLoop::pendingSubmissions()andpendingTimers()count what was handed over
between turns. Asubmitorschedulefrom a thread that is not the loop's worker -- the
case's own thread outside a turn included -- goes to the inbound queue until turn step 1, and the
two counters read only the ready queue and the park table, so a case that submitted and then
counted read 0 (at least 12 of fastcached's cases, on migration). They now add the inbound
submissions and scheduled deadlines, read under the inbound lock through two new
EventLoopaccessors,inboundSubmissionCount()andinboundScheduledCount()--
cancelPendingalready searched that queue, so the answer no longer depends on which thread
submitted. Both counters lostnoexcept: taking the lock can throw.
Changed
- A
PosixSocketkeeps one backend registration for its life, not one per parked operation.
Every read or write that parked attached a registration, armed it and detached it again -- on
epoll anEPOLL_CTL_ADDand anEPOLL_CTL_DELper request -- and fastcached's GET benchmark ran
6.8% slower onEventLoopthan on its own epoll reactor (geomean of 24 scenarios, -22% at the
worst). The socket's parks now ask forRegistrationLifetime::UntilClosed: the loop registers the
descriptor the first time it is parked on, a park takes a slot on that registration and changes
what it is armed for only when it must, andclose()ends it by announcing the close, as it
already did. Readability stays armed after a read completes, so the steady state of a
request/response connection costs noepoll_ctlat all; writability is dropped as soon as the
write is taken, and a readiness report nobody is parked to take narrows the registration after
that wait. On a loopback echo over epoll (WSL2, clang-22 Release, median of 5) the server's CPU
per request fell from 9.3-10.5 us to 5.1 us, and its throughput rose from 96-106k to 193-195k
requests a second at 16, 64 and 256 connections; a raw epoll echo, the floor, is 4.2 us. The
turn is unchanged.SocketRegistration_test.cppcounts the backend calls, and fails with the
per-park registration. - An operation parked on a socket allocates nothing in
EventLoop, and a turn nobody handed
work to takes no lock. Each park was a fresh allocation, filed in astd::unordered_mapand,
whatever its registration, in a second map by handle: three allocations and three frees per
parked operation. Parks are now recycled and kept in an open-addressing table, and a park on a
socket's lifetime registration is found through that registration rather than the handle map. A
turn skips the inbound mutex when nothing was posted and the timer heap when no deadline is armed,
stop()sets an atomic,ResultAwaitableregisters no stop callback on a token that can never
be stopped, a queued entry no longer moves an empty work item through a temporary, and
EpollBackendno longer zeroes a 768-byte event array on every wait. On fastcached's GET at 16
connections (WSL2, clang-22 Release, median of 5, the daemon's CPU per request), memcached text
went from 24.1 to 23.4 us and RESP from 26.9 to 25.0 us, against 23.7 and 25.7 us on fastcached's
own epoll reactor, and allocations per request from 10.1 to 7.1 and from 26.1 to 23.1.
Added
CORE_CPP_MSVC_STATIC_RUNTIME_VARIANTS(OFF): with an MSVC-ABI compiler, every compiled module
butcore::testing_main(whose Catch2 and dialog suppression are built/MD, and whose target
is edited after it is declared) also gets a static-CRT twin,core::<name>_mt(targetcore-cpp-<name>-mt), built/MTor
/MTdand linking the twins of the modules it links (header-only modules are shared as they
are). The MSVC linker refuses to mix C runtimes (LNK2038 ... 'RuntimeLibrary', or lld-link's
/failifmismatch), so one build of core-cpp could not serve both fastcached's/MDdaemon and
its/MTlauncher, fastcache-cc; with the option on it can. The twins are declared by
core_cpp_add_modulefrom the module table, so they follow it; they joinCORE_CPP_TARGETS,
and areEXCLUDE_FROM_ALL, so a parent builds only those it links. Other compilers ignore the
option with one status line. Two tests pin it on Windows: a/MTprogram linkingcore::net
andcore::logfails to link, namingRuntimeLibrary, and the same program linking
core::net_mtandcore::log_mtlinks and runs a loopback echo.RegistrationLifetime(<core/net/detail/ParkTable.hpp>, through<core/net/EventLoop.hpp>),
and alifetimefield of it onParkEntry, whichParkEntry::onReadyCallbacktakes as a new
trailing parameter with a default.PerPark, the default, is the registration every park had.
UntilClosedshares one registration per handle among every park on it that asks, kept until
EventLoop::notifyHandleClosingnames the handle -- so a caller that asks for it promises to
announce every close, which is whatPosixSocketdoes and what it now asks for.testing::ScriptedBackend::pushReadiness(HandlerId, Readiness): a scripted wait that reports
several conditions at once, as one kernel answer does (EPOLLIN|EPOLLOUT), so a case can reach
the choice a backend makes between two watched directions reported in the same wait.adoptSocket(EventLoop&, platform::NativeHandle, std::string peerAddress), beside
adoptFdandadoptListenerin<core/net/Sockets.hpp>: a connected socket accepted or
dialled outside core-cpp, driven by the loop the caller chooses. It is what a Windows server
needs to spread connections over several loops -- one thread accepts and deals each handle out,
since Windows has noSO_REUSEPORTand a completion-port association is one socket to one port
-- andadoptFdanswersUnsupportedthere. It builds the socket the loop's backend drives
(IocpSocketwhere the loop lends a completion port,WindowsSocketunder WFMO,PosixSocket
on POSIX), takes ownership of the handle on every path and closes it when adoption fails --
allocating the wrapper throwing included (the opposite ofadoptListener, which leaves a refused
handle with its caller), changes no socket
option beyond what the transport needs to run (non-blocking mode on POSIX), and asserts it is
called on the loop's thread.adoptFdis unchanged. Under WFMO, a readiness event that cannot
be created or associated is an error value, through the newWindowsSocket::adopt; the
WindowsSocketconstructor, whichconnectand the WFMO listener still use, can only carry on
with a socket that never becomes ready.SocketBufferSizes(<core/net/SocketBuffers.hpp>), and abuffersfield of it on both
ListenOptionsandDialOptions: the kernel send and receive buffers (SO_SNDBUF,
SO_RCVBUF) of every socket a listener accepts, and of one dialled socket. Each size is a
std::optional<std::size_t>, and an unset one leaves the kernel's value untouched, which is the
default. fastcached sizes both to 1 MiB so that a large reply leaves in onesendmsg. The sizes
are asked for before the connection exists -- of the listening socket beforelisten, which its
accepted sockets inherit (anAcceptExsocket included), and of a dialled socket before
connect-- because the TCP window scale is announced in the handshake and tcp(7) asks for them
first. A size is a request: Linux reports twice what was set and caps an
unprivileged request atnet.core.wmem_max/rmem_max.PosixListener::bindtakes two new
trailing parameters with defaults,PortSharing sharing(below) and then the sizes;
WindowsListener::bindandIocpListener::bindtake the sizes; andPosixListener,
WindowsListenerandIocpListenergainnative(), as the sockets have. The internal
detail::DialStep,detail::dialReadinessanddetail::dialCompletiontake a
detail::StreamSocketOptionswhere they took aKeepAlive.ListenOptions::sharing, aPortSharingdefaulting toPortSharing::Exclusive. With
PortSharing::Shared, several listeners may bind one port; fastcached's daemon binds one per loop
by default, and without it the second loop's bind failed withAddressInUse. Whether the
connections are spread across them is the platform's: Linux spreads them (SO_REUSEPORT), and
FreeBSD does withSO_REUSEPORT_LB, whichlisten()uses wherever the constant is defined. On
macOS and the other BSDs the binds coexist and the newest listener gets every connection
(SO_REUSEPORT), so one listener per loop leaves all but one loop idle there; accepting on one
and handing sockets out withadoptSocketis what spreads them. On Windows a shared listener is refused withNetErrorCode::Unsupportedrather
than mapped toSO_REUSEADDR, which there lets a later socket take a held port over. It is the
UDP sockets' existing enum, so<core/net/Sockets.hpp>now includes<core/net/UdpSocket.hpp>.
tools/migrate/renames.json'score::net::ReusePortrow points at it.core-cpp.await-ready, atree-levelcheck with a self-test (scripts/check-await-ready.py,
run by thestylejob), refusing anawait_readybody undersrc/ortests/that calls a
function or constructs an object, and anawait_suspendreturning a coroutine handle that
returns the handle it was given, unless its row names the test that runs it on the ARM64 leg. It cannot see an overloaded operator; the fixed public awaiters
also carry astatic_assert(core::async::awaitReadyIsConstantFalse<A>()), new in
<core/async/Awaitable.hpp>, which asks the question at compile time without constructing an
awaiter (P2280), so a call there fails to compile on GCC 14, Clang 20 and MSVC 19.51 or newer;
on older compilers it asserts nothing.- The
cl-release-arm64preset and thewindows (cl-release-arm64)CI leg, on
windows-11-arm: MSVC's ARM64 code generator is the only one that miscompiles the shape above,
and no other leg can observe it. The preset expects an arm64 developer shell.
Assets. core-cpp-v0.2.0-vendor.tar.gz is the vendoring file set; SHA256SUMS covers it.