Skip to content

NKDS Storage Model

Nanook edited this page Sep 24, 2026 · 1 revision

How NKDS physically stores disc images, how writes are made safe, and exactly what happens during compaction. This page is for users who want to understand the internals before trusting the store with their collections.

For day-to-day usage see NKDS and NKDS CLI.


Files on Disk

/data/nkit/
  wii.nkds            ← index file (256-byte header + block map + image metadata)
  wii_0000.nkds       ← shard 0 (compressed block data)
  wii_0001.nkds       ← shard 1 (created automatically when shard 0 fills up)
  gamecube.nkds
  gamecube_0000.nkds

A set is one index file plus zero or more shard files. Everything NKDS needs to reconstruct an image lives in these files.

Embedded mode (--shard-size 0) packs the index and all block data into a single .nkds file — useful for per-game OGMR sets you want to move around as a unit.


Content-Addressed Blocks

Every piece of disc data is stored as a block (default 64 KiB). Each block has a unique identity derived entirely from its content:

BlockKey = (XxHash64, CRC32)  →  96-bit collision resistance

This is computed over the raw, uncompressed block bytes. The same content — in any game, in any format — always produces the same key. Different content never produces the same key (with astronomically high probability).

Why this matters:

  • Two images sharing any 64 KiB of identical data automatically share that physical block — zero extra storage
  • Reads are verified: if the decompressed block doesn't match its stored key, the read fails rather than silently returning wrong data
  • Deduplication is exact and lossless — NKDS never approximates or discards anything

Deduplication happens at the filesystem level inside the disc image, not at raw sector level. NKit parses the disc's internal filesystem first, then deduplicates the actual files within it. This is why dedup rates are so high: Wii games share SDK libraries, system update data, and common media files — all identified and shared at the file level, even if they're at different disc offsets.


How a Write Works

When you run nkds add or nkds ogmr, NKDS processes the image through the full NKit pipeline, then stores it:

  1. NKit reads and normalises the source image (ISO, RVZ, WUX, CHD, archive — any supported format)
  2. Filesystem parse — the disc's internal file table is read up front
  3. Block chunking — each internal file is divided into 64 KiB blocks
  4. Dedup check — each block key is looked up in the index; existing blocks are referenced, not stored again
  5. Compression — new blocks are compressed with Zstandard before writing to the shard
  6. Append to shard — compressed blocks are appended at end-of-file; nothing is overwritten
  7. Index update — block locations and image metadata are written to the index
  8. Atomic commit — the header is updated in a specific safe order (see below)

Steps 1–7 leave the existing committed state completely intact if interrupted. Only step 8 makes the new data visible.


Append-Only Writes

NKDS never overwrites existing block data. Every write is an append.

The only bytes ever written in place are the two 256-byte headers at the start of the index file. All block data, all block maps, and all image metadata go to the end of the current file. Old sections become dead space that compaction later reclaims.

This means:

  • A power cut mid-add leaves the last committed state fully intact
  • There is no window where a partial write could corrupt previously stored images
  • Multiple images can be added in sequence without risk to already-stored data

Dual-Header Atomic Commit

NKDS uses a dual-header commit protocol to make every write atomic:

Index file layout:
  Offset 0x000: Primary header   (256 bytes)
  Offset 0x100: Secondary header (256 bytes)
  ... data ...

When committing a new image, writes happen in this exact order:

  1. Append all new data → flush to disk
  2. Write the secondary header (backup) → flush
  3. Write the primary header (commit point) → flush

The primary header is the last durable write. It is the commit point.

On recovery:

  • NKDS reads the primary header first
  • If the primary fails validation (torn write, partial flush), it falls back to the secondary header as authoritative
  • The secondary was written before the primary, so it always describes the last fully consistent state

If a crash happens between steps 2 and 3: the primary is stale, the secondary is the new committed state — recovery uses the secondary. If a crash happens before step 2: both headers describe the previous committed state — no data is lost and no partial write is visible.


Compaction

Compaction reclaims the space used by removed images. It is the only operation that permanently deletes data.

What compaction does

  1. Identifies all blocks referenced only by removed images (blocks still needed by live images are kept)
  2. Builds a new shard set containing only the live blocks, written to temporary files (.nkds.tmp)
  3. Validates the new temporary files before touching the originals
  4. Commits the new block index (atomic, using the dual-header protocol above)
  5. Only after the index is committed — atomically renames the temp shards over the originals

What compaction does NOT do

  • It does not touch images that have not been removed — their data is preserved exactly
  • It does not run automatically — you must explicitly call nkds compact
  • It does not proceed if the temp file validation fails — if something went wrong building the new shard, the originals are left untouched

Crash safety during compaction

If a crash occurs during compaction, NKDS uses a promote-or-delete recovery on the next open:

  • Temp shard files exist and the committed index expects the compacted sizes → compaction reached its commit point → the temp files are promoted (renamed over originals)
  • Temp shard files exist but the index still describes the pre-compaction sizes → compaction never committed → the temp files are deleted; the originals are used as-is

The committed index is the single source of truth. One of those two outcomes always applies — there is no ambiguous state.

History and current status

Earlier versions of NKDS had a compaction bug that could lose data under specific conditions. This was fixed in v3.0.0. The fix introduced the validate-before-replace step (the temp shard is opened and validated before it replaces the original) and the promote-or-delete recovery protocol. Both are now covered by automated tests.

If you used NKDS before v3.0.0 and ran compaction, it is worth re-verifying your sets with nkds verify. For collections that were only added to (no compaction), the data is intact — the compaction bug only affected the space-reclaim path.


The Block Index

The block index is a sorted binary array of (XxHash64, CRC32, FileId, Offset, Size) entries — approximately 28 bytes per unique block. For a 3 TB Wii collection with high dedup this is typically a few hundred MB.

Lookups are O(log n) binary search over the sorted keys. The index is memory-mapped for VFS reads, making random access fast even for very large collections.


Shard Files

Each shard is a flat file of compressed blocks. Blocks are never split across shards — if the next block would exceed the shard size limit, a new shard is started first. Once a block's position (shard file, offset, size) is committed, it is immutable for the life of that shard. Only compaction moves blocks, and it updates the index in lockstep.

Shards are opened with FileShare.ReadWrite | FileShare.Delete. This means:

  • Multiple concurrent VFS readers can read the same shard simultaneously
  • A writer adding new images can write to the same shard readers are reading (appending past their read position)
  • Compaction can atomically rename a shard while readers still hold handles — the stale handle is detected and refreshed on the next read

Thread Safety Summary

Scenario Safe?
Multiple concurrent VFS readers ✅ Lock-free read cache, independent handles
Reader + writer simultaneously ✅ Writer appends; readers read committed data
Two writers to the same set ❌ Single-writer enforced
Operations on different sets ✅ Per-set isolation
Mount active while adding images ✅ Mount sees the previous committed state until the new commit lands
Closing a set while mounted ✅ Close always unmounts first and waits for the mount host thread to exit

Verification

nkds verify re-reads each stored block, decompresses it, and checks the content against the stored (XxHash64, CRC32) key. Any block that fails is reported. No "trust me" — every byte is checked.

You can also verify against a dat file to confirm names and checksums match the expected dump:

nkds verify --datastore /data/nkit/wii.nkds --mask "*.iso" --config nkit.yaml

nkds verify --datastore D:\NKitData\wii.nkds --mask "*.iso" --config nkit.yaml

Further Reading

  • NKDS — Core concepts and quick start
  • NKDS Data Safety — Plain-language summary of what is safe and what to watch for
  • NKDS CLI — Full command reference
  • Architecture — How NKit and NKDS fit together

Clone this wiki locally