Releases: bitcoin-pow/cuda_btcw_miner
Release list
CUDA Miner new hard fork
BTCW CUDA Miner — RTX 4080/4070 SUPER new Hard Fork
CUDA miner for the BTCW NO_EXT_WORK/NEWFORK signing-work algorithm. The
RTX 4080 SUPER build uses the W26 lookup-table profile and BTCW's original
224-byte shared-memory payload.
Algorithm
For each block template, the BTCW node supplies a 32-byte secp256k1 secret
key, 160 bytes of context, and a fixed 32-byte hash_no_sig message.
The GPU searches a 32-bit test_case. Each candidate is evaluated as follows:
- Form RFC6979 extra entropy as
LE32(test_case) || 28 zero bytes. - Generate deterministic ECDSA nonce
kfrom the secret key, reduced
message, and extra entropy. Test case zero is skipped because zero is
reserved by the result mailbox. - Compute
R = kGon secp256k1 andr = R.x mod n. - Compute
s = k^-1 * (hash_no_sig + r * secret_key) mod n. - Normalize
sto low-S form and DER-encode(r,s). - Accept only 70- or 71-byte DER signatures.
- Compute
SHA256(SHA256(DER_signature)). - Compare the result with the fixed Stage-2 target
2^228 - 1, equivalent
to 28 leading zero bits in displayed hash order. - Return a successful
test_casethrough the existing 64-bit POSIX shared
memory nonce mailbox.
The BTCW node remains the final authority. It regenerates the signature using
CKey::Sign(hash_no_sig, ..., false, test_case) and validates it before block
submission.
CUDA design
Important optimizations include:
- fixed-layout, word-native RFC6979/HMAC-SHA256;
- reuse of secret-key and reduced-message SHA prefix state;
- secp256k1 GLV scalar splitting;
- a five-group signed W26 fixed-base lookup table of approximately 9 GiB;
- XYZZ mixed point addition;
- 128-candidate batched scalar inversion;
- two 64-candidate field-inversion sub-batches;
- PTX 32-bit Comba field and scalar multiplication;
- specialized DER SHA256d;
- direct comparison of the final SHA word with the fixed target.
The miner runs an RFC6979 GPU self-test at startup. It compares the optimized
extra-entropy implementation with the generic reference implementation for
known test cases and aborts if they differ.
Requirements
- Linux x86-64
- NVIDIA driver
- CUDA Toolkit with
nvcc - NVIDIA RTX 4080 SUPER
- substantially more than 12 GiB of free GPU memory for the W26 configuration
- a compatible BTCW NEWFORK node exposing
/shared_mem