fastcached v0.1.0
This is the first release of fastcached, so rather than a list of changes, here is what the project is and what it can do today.
What it is
Two things ship from this repository:
fastcached — an in-memory cache daemon that speaks the memcached text, memcached binary, memcached meta and Redis RESP2 protocols on a single port. It works out which one a client is using from the first bytes of the connection, so the port selects no protocol; every client reaches the same daemon on the same one.
fastcache-cc — a compiler launcher in the style of ccache and sccache, backed by that daemon. Its distinguishing property is that cache entries are portable across checkout paths: paths under your source root and build tree are rewritten to tokens before hashing, so the same source compiled at /home/alice/proj and at /ci/runner/w/1/s/proj produces the same cache key. CI runners and developer machines can therefore share one cache even when their trees live at different depths — which is the reason the launcher exists, since sccache keys on absolute paths and cannot.
They are useful together as a shared compile cache, and fastcached is useful on its own as a memcached/Redis-compatible cache, including as a plain sccache backend.
What is in this release
The daemon
- Four wire protocols on one port, detected per connection — memcached text, binary and meta (
mg/ms/md/ma/me/mn), and Redis RESP2. The default port is 6674, unassigned by IANA, unprivileged, and below the ephemeral range. Additional listeners are repeatable via--listen, so a client you cannot re-point keeps working on its own port, and that port speaks every protocol too. - Optional persistence with
--storage: a copy-on-write B+tree where every commit is crash-consistent. The file always matches either the previous transaction or the new one, so akill -9at any instant leaves no half-written state, and a restart picks the cache back up with no warm-up. Storage is composed as an in-memory LRU over the on-disk tree, sharded by key hash so writes to different shards do not block each other. --threads=Nruns N single-threaded reactors (epoll, kqueue or IOCP), each pinned to a core, with every connection pinned to one reactor for its lifetime. Concurrent clients are bounded by memory rather than by a worker count.- Authentication with
--requirepass(RedisAUTH, memcached SASL PLAIN), TLS on an OpenSSL build, and Prometheus/metricsplus/healthzon a separate admin port with--metrics. - YAML configuration with CLI flags taking precedence, re-read on
SIGHUPor the Windows service manager'sPARAMCHANGE. Started without--config, the daemon finds its own file from a per-platform list of locations — and which locations apply depends on whether the process could actually be the machine-wide service. - Packaged for Linux (
.deb,.rpm), macOS (.pkg,.dmg) and Windows (.msi), each installing both executables and registering the daemon to start automatically, plus aDockerfilefor containers.
The launcher
- Drop-in via
CMAKE_<LANG>_COMPILER_LAUNCHER, or as a plain prefix for a single compile. Supportsgcc/g++,clang/clang++including versioned names, and MSVCclandclang-cl. - If anything goes wrong — daemon unreachable, network gone, cache corrupt — it silently runs the real compiler. A broken cache can slow a build down; it should not be able to break one.
--show-statsreports hit rate, latency distributions and, importantly, a separateunavailablecount, so "the cache did not help" and "the cache was never reached" do not look alike.
Performance
On in-memory GET throughput, measured on one AMD Ryzen 9 9950X3D against redis 8.10.0 and memcached 1.6.40 built and run as native binaries on the same host:
| Concurrency | fastcached | vs redis | vs memcached |
|---|---|---|---|
| 1 | ~125k ops/s | ~1.05× (tie) | ~1.0× (tie) |
| 16 | ~1.06M ops/s | 3.0× | ~1.07× (tie) |
| 64 | ~1.45M ops/s | 4.5× | 1.6× |
| 256 | ~1.32M ops/s | 4.4× | 1.6× |
These are honest but narrow numbers, and worth reading with their limits attached: one fast desktop CPU, a single machine, and a Python load generator sharing that CPU with the server. At one connection there is no parallelism to exploit and all three tie. The redis baseline runs at its default io-threads 1, so a redis tuned for several I/O threads would narrow the network gap — what stands is the multi-core architecture, not a claim about redis at its best. The suite is in bench/ and reproduces with python bench/fastcached_bench.py --vs redis,memcached.
Scope, honestly
fastcached is not a general-purpose replacement for memcached or Redis. It implements only the slice of each protocol a cache backend actually uses, and the coverage matrix in the documentation says exactly which commands that is. If you need the rest of Redis, you need Redis.
It is, however, in production use as a shared compile cache for a large C++ codebase, backing both CI runners and developer machines.
Getting started
fastcached --storage=$HOME/.cache/fastcached/cache.cowThen point clients at it, or wire the launcher into a build — the README walks through both, and the full documentation lives at https://lastrada-software.github.io/fastcached/.
This being a first release, there will be rough edges we have not found yet. Issue reports and feedback are genuinely welcome.
Licensed under the Apache License, Version 2.0. Memcached and Redis are registered trademarks of their respective owners.