TiFlash segment removed after a TiKV node disconnected (after upgrade to v7.5.0) #66487
Replies: 3 comments
|
Original author: TiDB-AI-Assistant I can see that you are facing an issue with TiFlash after upgrading your TiDB cluster to version 7.5.0. It seems that the data between TiKV and TiFlash is out of sync, and you are experiencing empty data when querying via TiFlash. Additionally, the progress of your tables in the To address this issue, you have tried setting the replica count of all tables to 0, scaling in both of your TiFlash nodes, waiting for the Pending Offline status to end, pruning them, and then scaling them back into the cluster. After setting the replica count back to 2 for your tables, the replication runs normally. However, after about a day, the issue reoccurs, even after multiple attempts of scaling in and scaling out the TiFlash nodes . The scenario you described involves one of the TiKV nodes evicting all leader regions and then rebalancing, which seems to trigger the issue. The evicted TiKV node logs errors related to sending extra messages, starting a new election, and receiving a MsgTimeoutNow message. Subsequently, the TiFlash nodes exhibit strange logs, and the data on TiFlash becomes empty . The logs you provided indicate that the TiDBSchemaSyncer is unable to find related database IDs and logical table IDs, leading to a failure in syncing table schemas. This results in data loss on the TiFlash nodes, even though the region size does not decrease. Querying TiFlash at this point returns inaccurate results, while TiKV queries remain normal. Setting the replica count to 0 and then back to 2 causes the TiFlash nodes to become disconnected and then down, with some out-of-sync errors . Given the complexity of the issue and the detailed information you provided, it seems like a deeper investigation into the synchronization and communication between TiKV and TiFlash is required. It may be beneficial to review the cluster configuration, network settings, and resource allocation to ensure optimal performance and data consistency between the components . If you need further assistance or guidance on troubleshooting this issue, feel free to provide more details or logs for a more in-depth analysis. |
|
Original author: chidat.lu update: A little minutes before the segment drop at TiFlash, i found this log at TiKV |
|
Original author: TiDB-AI-Assistant It seems that you have identified a log in TiKV just before the segment drop in TiFlash. The log indicates an issue with getting a snapshot, specifically mentioning an error related to the epoch not matching for a particular region. The error message states that the current epoch of the region does not match the epoch sent in the request, leading to a failure in getting the snapshot . This discrepancy in epoch versions could potentially be a contributing factor to the synchronization issues between TiKV and TiFlash. It may be worth investigating further into the epoch mismatch and ensuring that the regions are properly synchronized and communicating with each other to prevent data inconsistencies and segment drops in TiFlash . If you need assistance in resolving this specific error or further guidance on troubleshooting the synchronization issue, feel free to provide additional details or logs for a more detailed analysis. |
Uh oh!
There was an error while loading. Please reload this page.
Original author: chidat.lu
Original time: 2024-02-21 03:56:56.782344
Original ID/Slug: 791 / tiflash-segment-removed-after-a-tikv-node-disconnected-after-upgrade-to-v7-5-0
Hello everyone,
I’m facing an issue with TiFlash after upgrade my TiDB cluster from v5.4 to v7.5.
A days after the upgrade, I realized that the data became out of sync between TiKV and TiFlash. The query via TiFlash returned almost empty data. When show tiflash_replica, all of my table progress is back to 0 or something not 1 but avaiable still 1.
After that, I decided to set all replica to 0, scale-in all 2 of my TiFlash node, waiting end of Pending Offline, prune its and scale-out them into cluster again. After set replica to 2 for my tables again, the replication is run normally. After replication of all tables completed, my query on TiFlash is on-sync again. But after about 1 days, the issue occurred again I did more 2 times scale-in scale-out, the issue still hasn’t been resolved.
Application environment:
Production
TiDB version:
v7.5.0
Reproduction method:
Set TiFlash replica of all table to 0, scale-in, scale-out and set to 2 again.
Problem:
After several attempts, I realized they all shared the same scenario. After all of tables had completed the replication about 1 days. 1 of TiKV node evicted all leader region and then rebalance again (1)
The evicted TiKV node log raised some ERROR like this
and re-join with this log
I don’t think it have OOM kill or restart occurred here.
But after that, two TiFlash node have some strange log and data on TiFlash become empty.
and then
As (2) image below, data on 2 TiFlash (10.0.0.4-5) node is gone but region size still not decrease. All TiFlash query at this time is return an not accurate result but TiKV query is normal.
If set replica to 0 and re-set them to 2 again, to TiFlash node become Disconnect and then Down with some not-sync error.
Retry scale-in scale-out at 3rd attemp still have same issue.
Resource allocation:
3 TiDB node: 32 CPU, 128 GB RAM, 500 GB SSD disk, 10GBitS NIC
3 PD node: deploy on same TiDB node
3 TiKV node: 40 CPU, 128 GB RAM, 3.5 TB SSD disk, 10GBitS NIC (10.0.0.1-3)
2 TiFlash node: 48 CPU, 128 GB RAM, 3.5TB SSD disk, 10GBitS NIC (10.0.0.4-5)
Attachment:
(1)
(2)
All reactions