Minimize memory internal fragmentation for Bloom filters #6427

pdillinger · 2020-02-18T18:56:12Z

Summary: New experimental option BBTO::optimize_filters_for_memory builds
filters that maximize their use of "usable size" from malloc_usable_size,
which is also used to compute block cache charges.

Rather than always "rounding up," we track state in the
BloomFilterPolicy object to mix essentially "rounding down" and
"rounding up" so that the average FP rate of all generated filters is
the same as without the option. (YMMV as heavily accessed filters might
be unluckily lower accuracy.)

Thus, the option near-minimizes what the block cache considers as
"memory used" for a given target Bloom filter false positive rate and
Bloom filter implementation. There are no forward or backward
compatibility issues with this change, though it only works on the
format_version=5 Bloom filter.

With Jemalloc, we see about 10% reduction in memory footprint (and block
cache charge) for Bloom filters, but 1-2% increase in storage footprint,
due to encoding efficiency losses (FP rate is non-linear with bits/key).

Why not weighted random round up/down rather than state tracking? By
only requiring malloc_usable_size, we don't actually know what the next
larger and next smaller usable sizes for the allocator are. We pick a
requested size, accept and use whatever usable size it has, and use the
difference to inform our next choice. This allows us to narrow in on the
right balance without tracking/predicting usable sizes.

Why not weight history of generated filter false positive rates by
number of keys? This could lead to excess skew in small filters after
generating a large filter.

Results from filter_bench with jemalloc (irrelevant details omitted):

(normal keys/filter, but high variance)
$ ./filter_bench -quick -impl=2 -average_keys_per_filter=30000 -vary_key_count_ratio=0.9
Build avg ns/key: 29.6278
Number of filters: 5516
Total size (MB): 200.046
Reported total allocated memory (MB): 220.597
Reported internal fragmentation: 10.2732%
Bits/key stored: 10.0097
Average FP rate %: 0.965228
$ ./filter_bench -quick -impl=2 -average_keys_per_filter=30000 -vary_key_count_ratio=0.9 -optimize_filters_for_memory
Build avg ns/key: 30.5104
Number of filters: 5464
Total size (MB): 200.015
Reported total allocated memory (MB): 200.322
Reported internal fragmentation: 0.153709%
Bits/key stored: 10.1011
Average FP rate %: 0.966313

(very few keys / filter, optimization not as effective due to ~59 byte
 internal fragmentation in blocked Bloom filter representation)
$ ./filter_bench -quick -impl=2 -average_keys_per_filter=1000 -vary_key_count_ratio=0.9
Build avg ns/key: 29.5649
Number of filters: 162950
Total size (MB): 200.001
Reported total allocated memory (MB): 224.624
Reported internal fragmentation: 12.3117%
Bits/key stored: 10.2951
Average FP rate %: 0.821534
$ ./filter_bench -quick -impl=2 -average_keys_per_filter=1000 -vary_key_count_ratio=0.9 -optimize_filters_for_memory
Build avg ns/key: 31.8057
Number of filters: 159849
Total size (MB): 200
Reported total allocated memory (MB): 208.846
Reported internal fragmentation: 4.42297%
Bits/key stored: 10.4948
Average FP rate %: 0.811006

(high keys/filter)
$ ./filter_bench -quick -impl=2 -average_keys_per_filter=1000000 -vary_key_count_ratio=0.9
Build avg ns/key: 29.7017
Number of filters: 164
Total size (MB): 200.352
Reported total allocated memory (MB): 221.5
Reported internal fragmentation: 10.5552%
Bits/key stored: 10.0003
Average FP rate %: 0.969358
$ ./filter_bench -quick -impl=2 -average_keys_per_filter=1000000 -vary_key_count_ratio=0.9 -optimize_filters_for_memory
Build avg ns/key: 30.7131
Number of filters: 160
Total size (MB): 200.928
Reported total allocated memory (MB): 200.938
Reported internal fragmentation: 0.00448054%
Bits/key stored: 10.1852
Average FP rate %: 0.963387

And from db_bench (block cache) with jemalloc:

$ ./db_bench -db=/dev/shm/dbbench.no_optimize -benchmarks=fillrandom -format_version=5 -value_size=90 -bloom_bits=10 -num=2000000 -threads=8 -compaction_style=2 -fifo_compaction_max_table_files_size_mb=10000 -fifo_compaction_allow_compaction=false
$ ./db_bench -db=/dev/shm/dbbench -benchmarks=fillrandom -format_version=5 -value_size=90 -bloom_bits=10 -num=2000000 -threads=8 -optimize_filters_for_memory -compaction_style=2 -fifo_compaction_max_table_files_size_mb=10000 -fifo_compaction_allow_compaction=false
$ (for FILE in /dev/shm/dbbench.no_optimize/*.sst; do ./sst_dump --file=$FILE --show_properties | grep 'filter block' ; done) | awk '{ t += $4; } END { print t; }'
17063835
$ (for FILE in /dev/shm/dbbench/*.sst; do ./sst_dump --file=$FILE --show_properties | grep 'filter block' ; done) | awk '{ t += $4; } END { print t; }'
17430747
$ #^ 2.1% additional filter storage
$ ./db_bench -db=/dev/shm/dbbench.no_optimize -use_existing_db -benchmarks=readrandom,stats -statistics -bloom_bits=10 -num=2000000 -compaction_style=2 -fifo_compaction_max_table_files_size_mb=10000 -fifo_compaction_allow_compaction=false -duration=10 -cache_index_and_filter_blocks -cache_size=1000000000
rocksdb.block.cache.index.add COUNT : 33
rocksdb.block.cache.index.bytes.insert COUNT : 8440400
rocksdb.block.cache.filter.add COUNT : 33
rocksdb.block.cache.filter.bytes.insert COUNT : 21087528
rocksdb.bloom.filter.useful COUNT : 4963889
rocksdb.bloom.filter.full.positive COUNT : 1214081
rocksdb.bloom.filter.full.true.positive COUNT : 1161999
$ #^ 1.04 % observed FP rate
$ ./db_bench -db=/dev/shm/dbbench -use_existing_db -benchmarks=readrandom,stats -statistics -bloom_bits=10 -num=2000000 -compaction_style=2 -fifo_compaction_max_table_files_size_mb=10000 -fifo_compaction_allow_compaction=false -optimize_filters_for_memory -duration=10 -cache_index_and_filter_blocks -cache_size=1000000000
rocksdb.block.cache.index.add COUNT : 33
rocksdb.block.cache.index.bytes.insert COUNT : 8448592
rocksdb.block.cache.filter.add COUNT : 33
rocksdb.block.cache.filter.bytes.insert COUNT : 18220328
rocksdb.bloom.filter.useful COUNT : 5360933
rocksdb.bloom.filter.full.positive COUNT : 1321315
rocksdb.bloom.filter.full.true.positive COUNT : 1262999
$ #^ 1.08 % observed FP rate, 13.6% less memory usage for filters

(Due to specific key density, this example tends to generate filters that are "worse than average" for internal fragmentation. "Better than average" cases can show little or no improvement.)

Test Plan: unit test added, 'make check' with gcc, clang and valgrind

facebook-github-bot

@pdillinger has imported this pull request. If you are a Facebook employee, you can view this diff on Phabricator.

facebook-github-bot · 2020-02-18T21:15:28Z

@pdillinger has updated the pull request. Re-import the pull request

facebook-github-bot

@pdillinger has imported this pull request. If you are a Facebook employee, you can view this diff on Phabricator.

facebook-github-bot · 2020-02-19T06:15:28Z

@pdillinger has updated the pull request. Re-import the pull request

facebook-github-bot

@pdillinger has imported this pull request. If you are a Facebook employee, you can view this diff on Phabricator.

facebook-github-bot · 2020-02-19T17:35:27Z

@pdillinger has updated the pull request. Re-import the pull request

facebook-github-bot

@pdillinger has imported this pull request. If you are a Facebook employee, you can view this diff on Phabricator.

facebook-github-bot · 2020-02-19T18:17:32Z

@pdillinger has updated the pull request. Re-import the pull request

facebook-github-bot · 2020-02-19T23:00:09Z

@pdillinger has updated the pull request. Re-import the pull request

facebook-github-bot

@pdillinger has imported this pull request. If you are a Facebook employee, you can view this diff on Phabricator.

siying

It looks good! I am surprised how sophisticated it has become. Is there a way to simplify a little bit, e.g. randomized the choice, rather than looking at balance?

siying · 2020-02-20T23:23:34Z

util/math.h

+  } else {
+    int lz = __builtin_clzll(static_cast<unsigned long long>(v));
+    return int{sizeof(unsigned long long)} * 8 - 1 - lz;
+  }


Wow, this is so sophisticated.

Summary: Useful in validating/testing internal fragmentation changes (#6427) Pull Request resolved: #6444 Test Plan: manual (no changes to production code) Differential Revision: D20040076 Pulled By: pdillinger fbshipit-source-id: 32d26f363d2a9ab9f5bebd281dcebd9915ae340e

Summary: Make kPageSize extern const size_t (used in draft #6427) Make kLitteEndian constexpr bool Clarify a couple of comments Pull Request resolved: #6443 Test Plan: make check, CI Differential Revision: D20044558 Pulled By: pdillinger fbshipit-source-id: e0c5cc13229c82726280dc0ddcba4078346b8418

pdillinger · 2020-03-04T17:14:57Z

It looks good! I am surprised how sophisticated it has become. Is there a way to simplify a little bit, e.g. randomized the choice, rather than looking at balance?

Hmm. I don't know why I didn't give that much thought before. It should give similar aggregate behavior without the extra state tracking.

pdillinger · 2020-03-04T17:31:59Z

It looks good! I am surprised how sophisticated it has become. Is there a way to simplify a little bit, e.g. randomized the choice, rather than looking at balance?

Hmm. I don't know why I didn't give that much thought before. It should give similar aggregate behavior without the extra state tracking.

Note to self: in order for CalculateSpace to predict the size returned by Finish (for partitioned filter), save a random threshold value in the builder object on construction and regenerate after each Finish.

facebook-github-bot · 2020-03-04T17:47:44Z

@pdillinger has updated the pull request. Re-import the pull request

facebook-github-bot

@pdillinger has imported this pull request. If you are a Facebook employee, you can view this diff on Phabricator.

pdillinger · 2020-04-15T05:28:46Z

This change is probably of limited value unless we also change the block cache charge policy to charge for internal fragmentation, so that the block cache can more reliably use only its configured limit. (If people are already choosing setting expecting internal fragmentation overhead, they will probably need to re-tweak settings after such a change.)

Summary: New experimental option BBTO::optimize_filters_for_memory builds filters that maximize their use of "usable size" from malloc_usable_size, which is also used to compute block cache charges. Rather than always "rounding up," we track state in the BloomFilterPolicy object to mix essentially "rounding down" and "rounding up" so that the average FP rate of all generated filters is the same as without the option. (YMMV as heavily accessed filters might be unluckily lower accuracy.) Thus, the option near-minimizes what the block cache considers as "memory used" for a given target Bloom filter false positive rate and Bloom filter implementation. There are no forward or backward compatibility issues with this change, though it only works on the format_version=5 Bloom filter. With Jemalloc, we see about 10% reduction in memory footprint (and block cache charge) for Bloom filters, but 1-2% increase in storage footprint, due to encoding efficiency losses (FP rate is non-linear with bits/key). Why not weighted random round up/down rather than state tracking? By only requiring malloc_usable_size, we don't actually know what the next larger and next smaller usable sizes for the allocator are. We pick a requested size, accept and use whatever usable size it has, and use the difference to inform our next choice. This allows us to narrow in on the right balance without tracking/predicting usable sizes. Why not weight history of generated filter false positive rates by number of keys? This could lead to excess skew in small filters after generating a large filter. Results with jemalloc (irrelevant details omitted): (normal keys/filter, but high variance) $ ./filter_bench -quick -impl=2 -average_keys_per_filter=30000 -vary_key_count_ratio=0.9 Build avg ns/key: 29.6278 Number of filters: 5516 Total size (MB): 200.046 Reported total allocated memory (MB): 220.597 Reported internal fragmentation: 10.2732% Bits/key stored: 10.0097 Average FP rate %: 0.965228 $ ./filter_bench -quick -impl=2 -average_keys_per_filter=30000 -vary_key_count_ratio=0.9 -optimize_filters_for_memory Build avg ns/key: 30.5104 Number of filters: 5464 Total size (MB): 200.015 Reported total allocated memory (MB): 200.322 Reported internal fragmentation: 0.153709% Bits/key stored: 10.1011 Average FP rate %: 0.966313 (very few keys / filter, optimization not as effective due to ~59 byte internal fragmentation in blocked Bloom filter representation) $ ./filter_bench -quick -impl=2 -average_keys_per_filter=1000 -vary_key_count_ratio=0.9 Build avg ns/key: 29.5649 Number of filters: 162950 Total size (MB): 200.001 Reported total allocated memory (MB): 224.624 Reported internal fragmentation: 12.3117% Bits/key stored: 10.2951 Average FP rate %: 0.821534 $ ./filter_bench -quick -impl=2 -average_keys_per_filter=1000 -vary_key_count_ratio=0.9 -optimize_filters_for_memory Build avg ns/key: 31.8057 Number of filters: 159849 Total size (MB): 200 Reported total allocated memory (MB): 208.846 Reported internal fragmentation: 4.42297% Bits/key stored: 10.4948 Average FP rate %: 0.811006 (high keys/filter) $ ./filter_bench -quick -impl=2 -average_keys_per_filter=1000000 -vary_key_count_ratio=0.9 Build avg ns/key: 29.7017 Number of filters: 164 Total size (MB): 200.352 Reported total allocated memory (MB): 221.5 Reported internal fragmentation: 10.5552% Bits/key stored: 10.0003 Average FP rate %: 0.969358 $ ./filter_bench -quick -impl=2 -average_keys_per_filter=1000000 -vary_key_count_ratio=0.9 -optimize_filters_for_memory Build avg ns/key: 30.7131 Number of filters: 160 Total size (MB): 200.928 Reported total allocated memory (MB): 200.938 Reported internal fragmentation: 0.00448054% Bits/key stored: 10.1852 Average FP rate %: 0.963387 Test Plan: unit test added, 'make check' with gcc, clang and valgrind

facebook-github-bot · 2020-06-18T20:37:32Z

@pdillinger has updated the pull request. Re-import the pull request

pdillinger · 2020-06-18T20:41:23Z

Simpler implementation that makes no assumptions other than using malloc_usable_size, as block cache does (previous comment inaccurate)

facebook-github-bot · 2020-06-18T21:04:23Z

@pdillinger has updated the pull request. Re-import the pull request

facebook-github-bot

@pdillinger has imported this pull request. If you are a Facebook employee, you can view this diff on Phabricator.

facebook-github-bot · 2020-06-19T03:32:32Z

@pdillinger has updated the pull request. Re-import the pull request

facebook-github-bot

@pdillinger has imported this pull request. If you are a Facebook employee, you can view this diff on Phabricator.

siying · 2020-06-19T18:38:22Z

table/block_based/filter_policy.cc

+          usable_len = size_t{0xffffffc0};
+        }
+        rv = (static_cast<uint32_t>(usable_len) & ~uint32_t{63}) +
+             /* metadata */ 5;


Here is what man page https://man7.org/linux/man-pages/man3/malloc_usable_size.3.html says:

The value returned by malloc_usable_size() may be greater than the requested size of the allocation because of alignment and minimum size constraints. Although the excess bytes can be overwritten by the application without ill effects, this is not good programming practice: the number of excess bytes in an allocation depends on the underlying implementation. The main use of this function is for debugging and introspection.

Are you sure we want to use those bytes? At least consulting jemalloc friends before doing that.

facebook-github-bot · 2020-06-19T22:05:54Z

@pdillinger has updated the pull request. Re-import the pull request

facebook-github-bot

@pdillinger has imported this pull request. If you are a Facebook employee, you can view this diff on Phabricator.

facebook-github-bot · 2020-06-19T22:10:39Z

@pdillinger has updated the pull request. Re-import the pull request

facebook-github-bot

@pdillinger has imported this pull request. If you are a Facebook employee, you can view this diff on Phabricator.

siying · 2020-06-20T00:05:15Z

include/rocksdb/table.h

+  // Because some memory counted by block cache might be unmapped pages within
+  // internal fragmentation, this option can increase observed RSS memory
+  // usage. With cache_index_and_filter_blocks=true, this option makes the
+  // block cache better at using space it is allowed.


Can we mention in the comment that the implementation is against the best-practice suggestion in malloc_usable_size()'s man page?

facebook-github-bot · 2020-06-22T18:26:26Z

@pdillinger has updated the pull request. Re-import the pull request

facebook-github-bot

@pdillinger has imported this pull request. If you are a Facebook employee, you can view this diff on Phabricator.

facebook-github-bot · 2020-06-22T20:42:18Z

@pdillinger merged this pull request in 5b2bbac.

pdillinger requested a review from siying February 18, 2020 18:56

facebook-github-bot added the CLA Signed label Feb 18, 2020

facebook-github-bot reviewed Feb 18, 2020

View reviewed changes

pdillinger removed the request for review from siying February 18, 2020 22:04

facebook-github-bot reviewed Feb 19, 2020

View reviewed changes

facebook-github-bot reviewed Feb 20, 2020

View reviewed changes

siying reviewed Feb 20, 2020

View reviewed changes

This was referenced Feb 21, 2020

Share kPageSize (and other small tweaks) #6443

Closed

Misc filter_bench improvements #6444

Closed

pdillinger force-pushed the filter-frag branch from e88237e to 6428c60 Compare March 4, 2020 17:47

facebook-github-bot reviewed Mar 4, 2020

View reviewed changes

pdillinger force-pushed the filter-frag branch from 6428c60 to 15de8d4 Compare June 18, 2020 20:37

pdillinger changed the title ~~WIP/RFC: Reclaim filter memory wasted to internal fragmentation~~ Minimize memory internal fragmentation for Bloom filters Jun 18, 2020

Add to HISTORY.md

22b3030

Fix unused variable on !ROCKSDB_MALLOC_USABLE_SIZE

64af4de

facebook-github-bot reviewed Jun 18, 2020

View reviewed changes

pdillinger added 3 commits June 18, 2020 20:20

Add optimize_filters_for_memory to db_bench

015c130

Fix test for 128-byte cache line

1d8d607

Add optimize_filters_for_memory to db_stress/db_crashtest

bca3c37

facebook-github-bot reviewed Jun 19, 2020

View reviewed changes

pdillinger requested a review from siying June 19, 2020 03:38

siying reviewed Jun 19, 2020

View reviewed changes

pdillinger added 2 commits June 19, 2020 13:43

Merge remote-tracking branch 'origin/master' into filter-frag

0cab66a

Some updates based on internal feedback

a1a4b1a

facebook-github-bot reviewed Jun 19, 2020

View reviewed changes

make format

afc21c9

facebook-github-bot reviewed Jun 19, 2020

View reviewed changes

siying approved these changes Jun 20, 2020

View reviewed changes

pdillinger added 2 commits June 22, 2020 11:19

More comments about malloc_usable_size

530a547

Merge remote-tracking branch 'origin/master' into filter-frag

d1294d8

facebook-github-bot reviewed Jun 22, 2020

View reviewed changes

facebook-github-bot closed this in 5b2bbac Jun 22, 2020

facebook-github-bot added the Merged label Jun 22, 2020

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Minimize memory internal fragmentation for Bloom filters #6427

Minimize memory internal fragmentation for Bloom filters #6427

pdillinger commented Feb 18, 2020 •

edited

facebook-github-bot left a comment

facebook-github-bot commented Feb 18, 2020

facebook-github-bot left a comment

facebook-github-bot commented Feb 19, 2020

facebook-github-bot left a comment

facebook-github-bot commented Feb 19, 2020

facebook-github-bot left a comment

facebook-github-bot commented Feb 19, 2020

facebook-github-bot commented Feb 19, 2020

facebook-github-bot left a comment

siying left a comment

siying Feb 20, 2020

pdillinger commented Mar 4, 2020

pdillinger commented Mar 4, 2020

facebook-github-bot commented Mar 4, 2020

facebook-github-bot left a comment

pdillinger commented Apr 15, 2020

facebook-github-bot commented Jun 18, 2020

pdillinger commented Jun 18, 2020

facebook-github-bot commented Jun 18, 2020

facebook-github-bot left a comment

facebook-github-bot commented Jun 19, 2020

facebook-github-bot left a comment

siying Jun 19, 2020

facebook-github-bot commented Jun 19, 2020

facebook-github-bot left a comment

facebook-github-bot commented Jun 19, 2020

facebook-github-bot left a comment

siying Jun 20, 2020

facebook-github-bot commented Jun 22, 2020

facebook-github-bot left a comment

facebook-github-bot commented Jun 22, 2020

Minimize memory internal fragmentation for Bloom filters #6427

Minimize memory internal fragmentation for Bloom filters #6427

Conversation

pdillinger commented Feb 18, 2020 • edited

facebook-github-bot left a comment

Choose a reason for hiding this comment

facebook-github-bot commented Feb 18, 2020

facebook-github-bot left a comment

Choose a reason for hiding this comment

facebook-github-bot commented Feb 19, 2020

facebook-github-bot left a comment

Choose a reason for hiding this comment

facebook-github-bot commented Feb 19, 2020

facebook-github-bot left a comment

Choose a reason for hiding this comment

facebook-github-bot commented Feb 19, 2020

facebook-github-bot commented Feb 19, 2020

facebook-github-bot left a comment

Choose a reason for hiding this comment

siying left a comment

Choose a reason for hiding this comment

siying Feb 20, 2020

Choose a reason for hiding this comment

pdillinger commented Mar 4, 2020

pdillinger commented Mar 4, 2020

facebook-github-bot commented Mar 4, 2020

facebook-github-bot left a comment

Choose a reason for hiding this comment

pdillinger commented Apr 15, 2020

facebook-github-bot commented Jun 18, 2020

pdillinger commented Jun 18, 2020

facebook-github-bot commented Jun 18, 2020

facebook-github-bot left a comment

Choose a reason for hiding this comment

facebook-github-bot commented Jun 19, 2020

facebook-github-bot left a comment

Choose a reason for hiding this comment

siying Jun 19, 2020

Choose a reason for hiding this comment

facebook-github-bot commented Jun 19, 2020

facebook-github-bot left a comment

Choose a reason for hiding this comment

facebook-github-bot commented Jun 19, 2020

facebook-github-bot left a comment

Choose a reason for hiding this comment

siying Jun 20, 2020

Choose a reason for hiding this comment

facebook-github-bot commented Jun 22, 2020

facebook-github-bot left a comment

Choose a reason for hiding this comment

facebook-github-bot commented Jun 22, 2020

pdillinger commented Feb 18, 2020 •

edited