DynamoDB: composite primary keys can collide because the internal key encoding is ambiguous #3755
Replies: 3 comments 1 reply
|
Thanks for the careful write-up, @chris-enroly. I reproduced it on current A few things from digging into it:
A bug issue would be useful for tracking. If you'd like to take the PR, it's yours; otherwise we'll pick it up. |
|
Filed as #3800 for tracking, with the reproduction, the TransactWriteItems symptom, and the option 2 direction we agreed here. |
|
Thanks for digging into it, and agreed on option 2. The escaping approach is better than what I sketched: keeping every key without No need to assign it to me. #3804 landed a few hours ago and already covers the ground, including the Happy to review #3804 if that is useful. One thing I would want to check there: escaping changes the byte ordering of encoded keys, and the Scan cursor advances through the map by that ordering, so it is worth confirming pagination order still matches DynamoDB's. |
Uh oh!
There was an error while loading. Please reload this page.
The problem
A composite primary key is stored under a single map key built by joining the partition and sort values with a literal
#:with the same join at line 3392 for the cursor/scan path.
#is a legal character inside a DynamoDB key attribute value, so two different primary keys can encode to the same address:A#BA##BA#BA##BThe second
PutItemreplaces the first. Both keys then read back as item 2, and item 1 is gone. Real DynamoDB stores both, since they are distinct primary keys.This is silent data loss rather than an error, which is what makes it awkward to notice: nothing fails at write time.
Reproduction
Table with a composite key:
Expected: the item with
marker=FIRST.Observed against
main(a2fddaf): the item withmarker=SECOND. Both key combinations return the second item, andScanshows one item where there should be two.I confirmed this at the service layer on current
main: after both writes,getItemfor("A#", "B")returns{"customerId":{"S":"A"},"orderId":{"S":"#B"},"marker":{"S":"SECOND"}}.Keys that embed
#as a separator are a common single-table design convention, so the colliding shapes are not especially exotic.Two ways to fix it
These have quite different trade-offs and the choice affects stored state, so it seemed better to ask before anyone writes code.
Option 1: detect the collision and reject the write
Keep the current encoding. Compare the structured key identity against the item already stored at that encoded address, and if they are different logical keys, fail the write with a
ValidationExceptioninstead of overwriting. The same check runs on theBatchWriteItemandTransactWriteItemspaths before any write is applied, so a batch cannot half-apply.Option 2: make the encoding unambiguous
Change the encoding so the delimiter cannot be confused with data, for example length-prefixed segments (
2#A#1#B) or escaping the delimiter inside each segment. Collisions become impossible by construction.One thing that does not work
Making the delimiter configurable. Whatever character is chosen is still legal inside a key value, so collisions remain possible and become rarer and harder to diagnose rather than being fixed.
Question
Which direction would you prefer, and is AWS fidelity the deciding factor here, or does the migration cost of changing the stored key format outweigh it? Happy to file this as a formal bug report as well if that is more useful for tracking.
All reactions