Overloading scylla with materialized view writes can lead to deadlock #15844

nuivall · 2023-10-25T18:22:15Z

Observed effect is that view update backlog grows and doesn't clear even after write load stops. We encountered it when table had 5 materialized views, potentially with hot partitions.

…vice group When base write triggers mv write and it needs to be send to another shard it used the same service group and we could end up with a deadlock. This fix affects also alternator's secondary indexes. Testing was done using (yet) uncommited framework for easy alternator performance testing: scylladb#13121. I've changed hardcoded max_nonlocal_requests config in scylla from 5000 to 500 and then ran: ./build/release/scylla perf-alternator-workloads --workdir /tmp/scylla-workdir/ --smp 2 \ --developer-mode 1 --alternator-port 8000 --alternator-write-isolation forbid --workload write_gsi \ --duration 60 --ring-delay-ms 0 --skip-wait-for-gossip-to-settle 0 --continue-after-error true --concurrency 2000 Without the patch when scylla is overloaded (i.e. number of scheduled futures being close to max_nonlocal_requests) after couple seconds scylla hangs, cpu usage drops to zero, no progress is made. We can confirm we're hitting this issue by seeing under gdb: p seastar::get_smp_service_groups_semaphore(2,0)._count $1 = 0 With the patch I wasn't able to observe the problem, even with 2x concurrency. I was able to make the process hang with 10x concurrency but I think it's hitting different limit as there wasn't any depleted smp service group semaphore and it was happening also on non mv loads. Fixes scylladb#15844

…vice group When base write triggers mv write and it needs to be send to another shard it used the same service group and we could end up with a deadlock. This fix affects also alternator's secondary indexes. Testing was done using (yet) not committed framework for easy alternator performance testing: scylladb#13121. I've changed hardcoded max_nonlocal_requests config in scylla from 5000 to 500 and then ran: ./build/release/scylla perf-alternator-workloads --workdir /tmp/scylla-workdir/ --smp 2 \ --developer-mode 1 --alternator-port 8000 --alternator-write-isolation forbid --workload write_gsi \ --duration 60 --ring-delay-ms 0 --skip-wait-for-gossip-to-settle 0 --continue-after-error true --concurrency 2000 Without the patch when scylla is overloaded (i.e. number of scheduled futures being close to max_nonlocal_requests) after couple seconds scylla hangs, cpu usage drops to zero, no progress is made. We can confirm we're hitting this issue by seeing under gdb: p seastar::get_smp_service_groups_semaphore(2,0)._count $1 = 0 With the patch I wasn't able to observe the problem, even with 2x concurrency. I was able to make the process hang with 10x concurrency but I think it's hitting different limit as there wasn't any depleted smp service group semaphore and it was happening also on non mv loads. Fixes scylladb#15844

…vice group When base write triggers mv write and it needs to be send to another shard it used the same service group and we could end up with a deadlock. This fix affects also alternator's secondary indexes. Testing was done using (yet) not committed framework for easy alternator performance testing: #13121. I've changed hardcoded max_nonlocal_requests config in scylla from 5000 to 500 and then ran: ./build/release/scylla perf-alternator-workloads --workdir /tmp/scylla-workdir/ --smp 2 \ --developer-mode 1 --alternator-port 8000 --alternator-write-isolation forbid --workload write_gsi \ --duration 60 --ring-delay-ms 0 --skip-wait-for-gossip-to-settle 0 --continue-after-error true --concurrency 2000 Without the patch when scylla is overloaded (i.e. number of scheduled futures being close to max_nonlocal_requests) after couple seconds scylla hangs, cpu usage drops to zero, no progress is made. We can confirm we're hitting this issue by seeing under gdb: p seastar::get_smp_service_groups_semaphore(2,0)._count $1 = 0 With the patch I wasn't able to observe the problem, even with 2x concurrency. I was able to make the process hang with 10x concurrency but I think it's hitting different limit as there wasn't any depleted smp service group semaphore and it was happening also on non mv loads. Fixes #15844 Closes #15845

…vice group When base write triggers mv write and it needs to be send to another shard it used the same service group and we could end up with a deadlock. This fix affects also alternator's secondary indexes. Deadlock scenario: 1. Let's assume that all smp service group semaphores have current value 1 2. Write arrives to shard A and gets forwarded to owning shard B 3. At this point A->B semaphore decreases to 0 4. On shard B it generates mv write with a different key 5. Meanwhile another (unrelated) write arrives to shard B 6. Similarly let's assume it gets forwarded to owning shard A 7. At this point both A->B and B->A semaphores are set to 0 8. Similarly second write, now executing on shard A generated mv write with a different key 9. Now we have 2 mv writes which need to be forwarded to owning shards, but they are waiting on two semaphores, blocking 2 base writes which own resources on those semaphores - deadlock. Testing was done using (yet) not committed framework for easy alternator performance testing: scylladb#13121. I've changed hardcoded max_nonlocal_requests config in scylla from 5000 to 500 and then ran: ./build/release/scylla perf-alternator-workloads --workdir /tmp/scylla-workdir/ --smp 2 \ --developer-mode 1 --alternator-port 8000 --alternator-write-isolation forbid --workload write_gsi \ --duration 60 --ring-delay-ms 0 --skip-wait-for-gossip-to-settle 0 --continue-after-error true --concurrency 2000 Without the patch when scylla is overloaded (i.e. number of scheduled futures being close to max_nonlocal_requests) after couple seconds scylla hangs, cpu usage drops to zero, no progress is made. We can confirm we're hitting this issue by seeing under gdb: p seastar::get_smp_service_groups_semaphore(2,0)._count $1 = 0 With the patch I wasn't able to observe the problem, even with 2x concurrency. I was able to make the process hang with 10x concurrency but I think it's hitting different limit as there wasn't any depleted smp service group semaphore and it was happening also on non mv loads. Fixes scylladb#15844

mykaul · 2023-11-19T12:45:31Z

@avikivity - can you prioritize backporting this? I think it's an important fix.

…vice group When base write triggers mv write and it needs to be send to another shard it used the same service group and we could end up with a deadlock. This fix affects also alternator's secondary indexes. Testing was done using (yet) not committed framework for easy alternator performance testing: #13121. I've changed hardcoded max_nonlocal_requests config in scylla from 5000 to 500 and then ran: ./build/release/scylla perf-alternator-workloads --workdir /tmp/scylla-workdir/ --smp 2 \ --developer-mode 1 --alternator-port 8000 --alternator-write-isolation forbid --workload write_gsi \ --duration 60 --ring-delay-ms 0 --skip-wait-for-gossip-to-settle 0 --continue-after-error true --concurrency 2000 Without the patch when scylla is overloaded (i.e. number of scheduled futures being close to max_nonlocal_requests) after couple seconds scylla hangs, cpu usage drops to zero, no progress is made. We can confirm we're hitting this issue by seeing under gdb: p seastar::get_smp_service_groups_semaphore(2,0)._count $1 = 0 With the patch I wasn't able to observe the problem, even with 2x concurrency. I was able to make the process hang with 10x concurrency but I think it's hitting different limit as there wasn't any depleted smp service group semaphore and it was happening also on non mv loads. Fixes #15844 Closes #15845 (cherry picked from commit 020a9c9)

avikivity · 2023-11-19T17:00:53Z

Backported to 5.1, 5.2, 5.4.

nuivall self-assigned this Oct 25, 2023

nuivall mentioned this issue Oct 25, 2023

db: view: run local materialized view mutations on a separate smp service group #15845

Closed

nuivall added type/bug Backport candidate labels Oct 25, 2023

mykaul added this to the 5.4 milestone Oct 26, 2023

mykaul added the area/materialized views label Oct 26, 2023

mykaul added the P1 Urgent label Oct 26, 2023

scylladb-promoter closed this as completed in 020a9c9 Oct 30, 2023

avikivity removed the Backport candidate label Nov 19, 2023

mykaul mentioned this issue Dec 26, 2023

Shard going 100% CPU #16564

Closed

1 task

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Overloading scylla with materialized view writes can lead to deadlock #15844

Overloading scylla with materialized view writes can lead to deadlock #15844

nuivall commented Oct 25, 2023

mykaul commented Nov 19, 2023

avikivity commented Nov 19, 2023

Overloading scylla with materialized view writes can lead to deadlock #15844

Overloading scylla with materialized view writes can lead to deadlock #15844

Comments

nuivall commented Oct 25, 2023

mykaul commented Nov 19, 2023

avikivity commented Nov 19, 2023