Repository navigation
Releases: kemo159/multicyclone
Release list
MultiCyclone v3.1 - autosave timer
Adds --autosavetimer SECONDS — periodic checkpoints while the search runs.
A checkpoint was only written on a clean exit: Ctrl+C, --seconds, or finishing
the range. A crash, a power cut or kill -9 left nothing behind, which is
exactly the case a checkpoint exists to cover. On a multi-day search that is the
whole run gone.
CUDACyclone --range AAAA:BBBB --address 1Abc... --autosavetimer 300
[autosave] checkpoint written to cyclone_checkpoint.txt at 12.04% (48318382080 keys checked)
Recovery is the ordinary --resume path — an autosaved file is the same format
and the same encryption the exit path already writes.
Notes
- It does not pause the search. The save reads the per-thread counters
without draining the GPU pipeline. An earlier version synchronised the stream
and stalled the search ~6 s at every save; that is gone. - It cannot skip keys. The counters are read while kernels still run. They
only ever count down and the code takes the max across threads, so a
mid-flight read understates progress — a resume repeats a little work rather
than missing any. - A crash during a save cannot corrupt the previous one. The file is written
to<checkpoint>.tmpand renamed over the target. - Ignored in random mode, which has no linear progress to record.
--autosaveworks as an alias.
Pick the interval for what you are protecting against; 300 (5 min) is sensible
for an overnight run.
Assets
| asset | architectures | built with |
|---|---|---|
CUDACyclone-v3.1-windows-x64-sm120.exe |
sm_120 only | CUDA 13.1 + MSVC 14.44 |
CUDACyclone-v3.1-linux-x64-sm120 |
sm_120 only | CUDA 12.8 + g++, Ubuntu 24.04 |
These are RTX 50-series only. On any other card take
v3.1-multiarch,
which covers Turing through Blackwell, or build from source — see README.md.
The Linux build is deliberately on CUDA 12.8 rather than 13.x: it asks for an
older minimum driver, so it runs on machines that have not updated. chmod +x
it after downloading.
MultiCyclone v3.1 - multi-architecture (Turing and up)
Same code as v3.1
— only the compiled GPU architectures differ. Grab these if you are not on an
RTX 50-series card.
| architectures | GPUs | size | |
|---|---|---|---|
| this release | sm_75, 80, 86, 89, 90, 120 | Turing → Blackwell | 41–54 MB |
| v3.1 | sm_120 only | RTX 50-series only | 9–15 MB |
That covers RTX 20 / GTX 16 (Turing), A100, RTX 30 (Ampere), RTX 40 (Ada), H100
(Hopper) and RTX 50 (Blackwell).
| asset | built with | needs |
|---|---|---|
CUDACyclone-v3.1-windows-x64-multiarch.exe |
CUDA 13.1 + MSVC 14.44 | NVIDIA driver for CUDA 13.x |
CUDACyclone-v3.1-linux-x64-multiarch |
CUDA 12.8 + g++, Ubuntu 24.04 | NVIDIA driver for CUDA 12.8+, chmod +x |
The Linux build is deliberately on CUDA 12.8 rather than 13.x: it asks for an
older minimum driver, so it runs on machines that have not updated.
There is no speed penalty. A fatbinary carries native SASS per architecture,
so the card runs its own code either way. The only cost is file size.
What is new in v3.1
--autosavetimer SECONDS rewrites the checkpoint every N seconds while the
search keeps running, so a crash, a power cut or kill -9 costs at most one
interval instead of the whole run. Ctrl+C and --seconds already saved on the
way out; this covers the stops that never get the chance.
CUDACyclone --range AAAA:BBBB --address 1Abc... --autosavetimer 300
[autosave] checkpoint written to cyclone_checkpoint.txt at 12.04% (48318382080 keys checked)
Recovery is the ordinary --resume path. The save does not pause the search, it
cannot skip keys, and a crash mid-save cannot truncate the previous good
checkpoint — see v3.1
for the details.
Anything older than Turing
Not covered. CUDA 13.x removed Maxwell, Pascal and Volta entirely, and while
CUDA 12.8 can still target them, none of it is tested here. Build from source if
you need one.
MultiCyclone v3.0
Not on an RTX 50-series card? The binaries here are
sm_120only and will
not run on anything else. Use
v3.0-multiarch
instead — same code, built for Turing through Blackwell (sm_75/80/86/89/90/120),
at identical speed.
Read this first
Both binaries are built for sm_120 only — RTX 50-series. They will not run on
any other GPU. For anything else, build from source; it takes a few minutes:
make -j$(nproc) CUDA_ARCHS=120 # just your own GPU
make -j$(nproc) # every architecture your nvcc supports
| asset | built with | needs |
|---|---|---|
CUDACyclone-v3.0-windows-x64-sm120.exe |
CUDA 13.1 + MSVC 14.44 | NVIDIA driver for CUDA 13.x |
CUDACyclone-v3.0-linux-x64-sm120 |
CUDA 12.8 + g++, Ubuntu 24.04 | NVIDIA driver for CUDA 12.8+, chmod +x |
Stop and resume
Ctrl+C or --seconds now writes a checkpoint; the same command plus --resume
picks up where it left off, as many times as you like.
^C
======== INTERRUPTED (Ctrl+C) ==========================
Checkpoint saved to cyclone_checkpoint.txt at 18.32% (73282879488 keys checked, encrypted)
Resume with the same command plus --resume
--resume refuses to run unless the target, --range, --grid, --slices,
--tpb and the GPU set all match what the checkpoint was written under — each
of those re-tiles the range across threads, so a saved offset would no longer
mean what it did. Mismatches are reported field by field rather than silently
searching the wrong keys.
Checkpoints are encrypted. In the clear the file would name the address you
are hunting, the range, and how far you have got. The key is derived from the
search identity (target hash160 + range), so resuming needs no extra secret —
--resume already requires the same --address and --range. Add
--checkpoint-pass PASS (or set CUDACYCLONE_CHECKPOINT_PASS) to mix in a
passphrase, which also seals it against someone who does know the target;
lose that passphrase and the progress is unrecoverable. SHA-256 counter mode,
encrypt-then-MAC with HMAC-SHA256, fresh random nonce per write, MAC verified
before anything is parsed.
Idle GPUs take over the CPU tail
--cpu-threads splits a --cpu-percent slice off the end of the range for the
CPU. That split is fixed up front, so if it over-allocates the CPU the GPUs
finish and sit idle. They no longer wait:
GPUs finished their share; taking over the CPU sidecar's remaining
FFCAD3992A - 10000000001 (0.89B keys)
Each CPU thread reports its position, and the takeover starts at the lowest one
— a superset of the outstanding work, so nothing can be missed. Measured 8.4 s
against ~169 s on a 2-billion-key CPU tail. --cpu-auto still benchmarks both
sides and splits so they finish together; the takeover is the safety net for
when that drifts.
Fixes
--secondsno longer claims the range was exhausted. It printed
KEY NOT FOUND (exhaustive)and exited 0 after stopping at any coverage at
all. Now printsSTOPPED (time limit)and exits 2.- Reported hash rate was swinging ±10% while real throughput was steady. The
per-thread counter flushed at 65536, but a thread only doesslices * Bkeys
per launch (32768 at the default64 x 512), so it never fired and the
counter only moved at kernel end — aliasing against the 1 s sampling. Now
3.7% solo and 1.4% with the CPU active, at unchanged throughput. makecould not build on Windows at all (23 errors inInt.h).cpu_avx2
gates its MSVC vs GCC intrinsic paths onWIN64, not the_WIN64MSVC
predefines, and the Makefile never defined it.CMakeLists.txtalready did.host_sha256::sha256overruns its stack buffer above 55 bytes. It handles
exactly one 64-byte block. Safe for its existing 25- and 32-byte callers, but
a landmine for reuse; the checkpoint crypto uses a proper multi-block
implementation, verified against FIPS 180-4 and RFC 4231 vectors.
Exit codes
| code | meaning |
|---|---|
| 0 | key found, or range searched exhaustively |
| 1 | bad arguments, or a --resume checkpoint that does not match |
| 2 | stopped by --seconds before the range was exhausted |
| 130 | interrupted with Ctrl+C |
2 and 130 mean the range was not fully searched — both leave a checkpoint.
Earlier releases and the pre-v3.0 history remain available under the v2.5,
v2.1 and Release tags.
MultiCyclone v3.0 - multi-architecture (Turing and up)
Same code as v3.0
— only the compiled GPU architectures differ. Grab these if you are not on an
RTX 50-series card.
| architectures | GPUs | size | |
|---|---|---|---|
| this release | sm_75, 80, 86, 89, 90, 120 | Turing → Blackwell | 41–54 MB |
| v3.0 | sm_120 only | RTX 50-series only | 10–15 MB |
That covers RTX 20 / GTX 16 (Turing), A100, RTX 30 (Ampere), RTX 40 (Ada), H100
(Hopper) and RTX 50 (Blackwell).
| asset | built with | needs |
|---|---|---|
CUDACyclone-v3.0-windows-x64-multiarch.exe |
CUDA 13.1 + MSVC 14.44 | NVIDIA driver for CUDA 13.x |
CUDACyclone-v3.0-linux-x64-multiarch |
CUDA 12.8 + g++, Ubuntu 24.04 | NVIDIA driver for CUDA 12.8+, chmod +x |
The Linux build is deliberately on CUDA 12.8 rather than 13.x: it asks for an
older minimum driver, so it runs on machines that have not updated.
There is no speed penalty. A fatbinary carries native SASS per architecture,
so the card runs its own code either way — measured on an RTX 5080, 30 s runs:
sm_120-only build 4848 Mkeys/s
multi-arch build 4864 Mkeys/s
The only cost is file size. If you are on an RTX 50-series card and care about
the download, take the smaller v3.0 build; otherwise this one is fine everywhere.
Anything older than Turing
Not covered. CUDA 13.x removed Maxwell, Pascal and Volta entirely, and while
CUDA 12.8 can still target them, none of it is tested here. If you want
sm_50–sm_70, build from source with a CUDA 12.x toolkit and add the
architectures to KNOWN_SM_ARCHS in the Makefile:
make -j$(nproc) CUDA_ARCHS=61 # Pascal, for example
Verified
Both binaries were checked before upload — they report MULTICYCLONE v3.0,
recover the known key 0x100001, and cuobjdump confirms all six architectures
are embedded. Only sm_120 was exercised on real hardware; the rest are compiled
but untested, so please open an issue if an older card misbehaves.
Note for anyone inspecting the Linux binary: cuobjdump from CUDA 12.8 misreads
the sm_120 section and aborts with Support for 'sm_2' has been removed. That is
a bug in the 12.8 tool, not a broken binary — CUDA 13.1's cuobjdump lists all
six correctly.
Full feature notes for v3.0 — encrypted checkpoint/resume, GPU takeover of the
CPU tail, and the fixes — are on the
v3.0 release.
MultiCyclone v2.5
MultiCyclone v2.5
Highlights:
- Embedded AVX2 CPU worker in the main executable.
- Hybrid CPU/GPU search with --cpu-auto benchmark-based range split.
- Windows and Linux/WSL sidecar process support.
- Launch guard for very large grids, preserving thread count while shortening kernel launches.
- Default builds now target every supported CUDA architecture from GTX 1660 / Turing sm_75 upward.
Windows artifact:
- MultiCyclone-v2.5-windows-sm75plus.zip
- Built for sm_75, sm_80, sm_86, sm_87, sm_88, sm_89, sm_90, sm_100, sm_103, sm_110, sm_120, sm_121 plus compute_121 PTX.
Smoke tested:
- Known-key GPU test passed.
- CPU sidecar known-key test passed.
MultiCyclone v2.1
Windows x64 rebuild with CUDA 13.1 for sm_75, sm_80, sm_86, sm_87, sm_88, sm_89, sm_90, sm_100, sm_103, sm_110, sm_120, and sm_121. Requires a compatible NVIDIA driver; CUDA Toolkit is not required to run.
Windows Mulitcyclone 2.0
Windows and Linux Version CUDA 13