-
Notifications
You must be signed in to change notification settings - Fork 7
NKDS Storage Model
How NKDS physically stores disc images, how writes are made safe, and exactly what happens during compaction. This page is for users who want to understand the internals before trusting the store with their collections.
For day-to-day usage see NKDS and NKDS CLI.
/data/nkit/
wii.nkds ← index file (256-byte header + block map + image metadata)
wii_0000.nkds ← shard 0 (compressed block data)
wii_0001.nkds ← shard 1 (created automatically when shard 0 fills up)
gamecube.nkds
gamecube_0000.nkds
A set is one index file plus zero or more shard files. Everything NKDS needs to reconstruct an image lives in these files.
Embedded mode (--shard-size 0) packs the index and all block data into a single .nkds file — useful for per-game OGMR sets you want to move around as a unit.
Every piece of disc data is stored as a block (default 64 KiB). Each block has a unique identity derived entirely from its content:
BlockKey = (XxHash64, CRC32) → 96-bit collision resistance
This is computed over the raw, uncompressed block bytes. The same content — in any game, in any format — always produces the same key. Different content never produces the same key (with astronomically high probability).
Why this matters:
- Two images sharing any 64 KiB of identical data automatically share that physical block — zero extra storage
- Reads are verified: if the decompressed block doesn't match its stored key, the read fails rather than silently returning wrong data
- Deduplication is exact and lossless — NKDS never approximates or discards anything
Deduplication happens at the filesystem level inside the disc image, not at raw sector level. NKit parses the disc's internal filesystem first, then deduplicates the actual files within it. This is why dedup rates are so high: Wii games share SDK libraries, system update data, and common media files — all identified and shared at the file level, even if they're at different disc offsets.
When you run nkds add or nkds ogmr, NKDS processes the image through the full NKit pipeline, then stores it:
- NKit reads and normalises the source image (ISO, RVZ, WUX, CHD, archive — any supported format)
- Filesystem parse — the disc's internal file table is read up front
- Block chunking — each internal file is divided into 64 KiB blocks
- Dedup check — each block key is looked up in the index; existing blocks are referenced, not stored again
- Compression — new blocks are compressed with Zstandard before writing to the shard
- Append to shard — compressed blocks are appended at end-of-file; nothing is overwritten
- Index update — block locations and image metadata are written to the index
- Atomic commit — the header is updated in a specific safe order (see below)
Steps 1–7 leave the existing committed state completely intact if interrupted. Only step 8 makes the new data visible.
NKDS never overwrites existing block data. Every write is an append.
The only bytes ever written in place are the two 256-byte headers at the start of the index file. All block data, all block maps, and all image metadata go to the end of the current file. Old sections become dead space that compaction later reclaims.
This means:
- A power cut mid-add leaves the last committed state fully intact
- There is no window where a partial write could corrupt previously stored images
- Multiple images can be added in sequence without risk to already-stored data
NKDS uses a dual-header commit protocol to make every write atomic:
Index file layout:
Offset 0x000: Primary header (256 bytes)
Offset 0x100: Secondary header (256 bytes)
... data ...
When committing a new image, writes happen in this exact order:
- Append all new data → flush to disk
- Write the secondary header (backup) → flush
- Write the primary header (commit point) → flush
The primary header is the last durable write. It is the commit point.
On recovery:
- NKDS reads the primary header first
- If the primary fails validation (torn write, partial flush), it falls back to the secondary header as authoritative
- The secondary was written before the primary, so it always describes the last fully consistent state
If a crash happens between steps 2 and 3: the primary is stale, the secondary is the new committed state — recovery uses the secondary. If a crash happens before step 2: both headers describe the previous committed state — no data is lost and no partial write is visible.
Compaction reclaims the space used by removed images. It is the only operation that permanently deletes data.
- Identifies all blocks referenced only by removed images (blocks still needed by live images are kept)
- Builds a new shard set containing only the live blocks, written to temporary files (
.nkds.tmp) - Validates the new temporary files before touching the originals
- Commits the new block index (atomic, using the dual-header protocol above)
- Only after the index is committed — atomically renames the temp shards over the originals
- It does not touch images that have not been removed — their data is preserved exactly
- It does not run automatically — you must explicitly call
nkds compact - It does not proceed if the temp file validation fails — if something went wrong building the new shard, the originals are left untouched
If a crash occurs during compaction, NKDS uses a promote-or-delete recovery on the next open:
- Temp shard files exist and the committed index expects the compacted sizes → compaction reached its commit point → the temp files are promoted (renamed over originals)
- Temp shard files exist but the index still describes the pre-compaction sizes → compaction never committed → the temp files are deleted; the originals are used as-is
The committed index is the single source of truth. One of those two outcomes always applies — there is no ambiguous state.
Earlier versions of NKDS had a compaction bug that could lose data under specific conditions. This was fixed in v3.0.0. The fix introduced the validate-before-replace step (the temp shard is opened and validated before it replaces the original) and the promote-or-delete recovery protocol. Both are now covered by automated tests.
If you used NKDS before v3.0.0 and ran compaction, it is worth re-verifying your sets with nkds verify. For collections that were only added to (no compaction), the data is intact — the compaction bug only affected the space-reclaim path.
The block index is a sorted binary array of (XxHash64, CRC32, FileId, Offset, Size) entries — approximately 28 bytes per unique block. For a 3 TB Wii collection with high dedup this is typically a few hundred MB.
Lookups are O(log n) binary search over the sorted keys. The index is memory-mapped for VFS reads, making random access fast even for very large collections.
Each shard is a flat file of compressed blocks. Blocks are never split across shards — if the next block would exceed the shard size limit, a new shard is started first. Once a block's position (shard file, offset, size) is committed, it is immutable for the life of that shard. Only compaction moves blocks, and it updates the index in lockstep.
Shards are opened with FileShare.ReadWrite | FileShare.Delete. This means:
- Multiple concurrent VFS readers can read the same shard simultaneously
- A writer adding new images can write to the same shard readers are reading (appending past their read position)
- Compaction can atomically rename a shard while readers still hold handles — the stale handle is detected and refreshed on the next read
| Scenario | Safe? |
|---|---|
| Multiple concurrent VFS readers | ✅ Lock-free read cache, independent handles |
| Reader + writer simultaneously | ✅ Writer appends; readers read committed data |
| Two writers to the same set | ❌ Single-writer enforced |
| Operations on different sets | ✅ Per-set isolation |
| Mount active while adding images | ✅ Mount sees the previous committed state until the new commit lands |
| Closing a set while mounted | ✅ Close always unmounts first and waits for the mount host thread to exit |
nkds verify re-reads each stored block, decompresses it, and checks the content against the stored (XxHash64, CRC32) key. Any block that fails is reported. No "trust me" — every byte is checked.
You can also verify against a dat file to confirm names and checksums match the expected dump:
nkds verify --datastore /data/nkit/wii.nkds --mask "*.iso" --config nkit.yaml
nkds verify --datastore D:\NKitData\wii.nkds --mask "*.iso" --config nkit.yaml- NKDS — Core concepts and quick start
- NKDS Data Safety — Plain-language summary of what is safe and what to watch for
- NKDS CLI — Full command reference
- Architecture — How NKit and NKDS fit together