Skip to content

CASSANDRA-21632: Read the BTI partition index preload in chunks - #5089

Open
aviau wants to merge 2 commits into
apache:cassandra-5.0from
aviau:bti-preload-chunked
Open

CASSANDRA-21632: Read the BTI partition index preload in chunks#5089
aviau wants to merge 2 commits into
apache:cassandra-5.0from
aviau:bti-preload-chunked

Conversation

@aviau

@aviau aviau commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Meta

Description

Bounding the preload makes startup survivable, but touching one byte per page is still a slow way to warm a range: one pread or one page fault per page, 81.1M reads on my 309 GiB index. Reading in 1 MiB positioned reads on the channel gets the same pages resident at device bandwidth instead. Positioned reads are unaffected by disk_access_mode, and page cache is shared between the buffered and mapped views of a file, so a mapped lookup still finds the data resident.

This matters most when the warm is large, which is to say at the default (whole index) or at a large bound. With a small bound the walk is already cheap and this buys little.

Trade-offs:

  • The whole range is copied into userspace rather than one byte per page. The kernel reads the same pages either way, so this costs memory bandwidth, not I/O.
  • Under a mapped access mode the walk also populated the process page table entries. Channel reads leave a minor fault per page to the first lookup that touches it. No I/O is involved, but it is not nothing.
  • It allocates a 1 MiB direct buffer per opening thread, freed explicitly with FileUtils.clean rather than at GC.
  • Under disk_access_mode: standard the warm no longer populates the chunk cache, nor evicts it. Lookups fill it on first touch instead. Under the default mmap_index_only the partition index was never in the chunk cache to begin with.

Assisted-by: Claude Code:claude-opus-5

patch by Alexandre Viau; reviewed by TBD for CASSANDRA-21632

Opening a BTI SSTable whose bloom filter is uninformative warms the whole
Partitions.db, one byte per page through a reader that follows disk_access_mode. That
costs one pread or one readahead-bounded page fault per page: on a 309 GiB index,
81.1M reads at 50 MiB/s regardless of access mode. It also warms far more than the
available page cache can hold.

Add cassandra.bti.partition_index_preload_size to bound how much is warmed. The tail
is warmed because the trie is written bottom-up, so the upper levels traversed by
every lookup are at the end of the file, while the bulk at the front is leaf pages
that a lookup reads one of.

The property takes a human-readable size, e.g. 512MiB. 0B skips warming entirely,
which was not previously possible: preload is enabled whenever the bloom filter is
uninformative, so the only way to avoid it was to lower bloom_filter_fp_chance below
1.0. A negative value warms the whole index and is the default, so behaviour is
unchanged unless the property is set.

Assisted-by: Claude Code:claude-opus-5

patch by Alexandre Viau; reviewed by TBD for CASSANDRA-21632
@aviau
aviau force-pushed the bti-preload-chunked branch from f9968f7 to 389b5ae Compare September 1, 2026 20:53
@aviau
aviau changed the base branch from trunk to cassandra-5.0 September 1, 2026 21:11

// ChannelProxy.read may return a short read; advance by what was actually read.
int read = fh.channel.read(buffer, pos);
if (read <= 0)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should treat read < 0 as an EOFException, as FileDataInput.readByte() would have previously.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair request. Done!

Touching one byte per page costs a pread or a page fault per page: 81.1M reads on a
309 GiB index. Read it in 1 MiB positioned reads on the channel instead, which are
unaffected by disk_access_mode; page cache is shared between the buffered and mapped
views of a file, so a mapped lookup still finds the data resident.

The trade-offs: the whole range is copied into userspace rather than one byte per
page, and under a mapped mode the first lookup takes a minor fault the walk had
already resolved. Under disk_access_mode: standard the warm no longer populates the
chunk cache, nor evicts it; lookups fill it on first touch instead.

Assisted-by: Claude Code:claude-opus-5

patch by Alexandre Viau; reviewed by TBD for CASSANDRA-21632
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants