Repository navigation
[Feature] partial hashing for faster and less problematic uploads of large files #32170
LucaTheHacker
started this conversation in
Feature Request
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
I have searched the existing feature requests, both open and closed, to make sure this is not a duplicate request.
The feature
Somewhat inspired by #18991 and my attempt to upload a 35 GiB video from Safari.
Currently, Immich performs a full hash on the client and another on the server to guarantee that no duplicate files are present. Client-provided hashes are not trusted, which is a good choice, but this may cause performance issues or even make uploads impossible on some clients and browsers.
What I believe would be an improvement is the client-side duplicate check:
Checking file size is faster than computing a hash.
A partial hash is usually fine for a high-entropy file such as a video.
For this reason, I believe a cut-off size should be decided and, only for files above it, the duplicate checker should base itself on the file size plus the hash of the first N MB of the file.
If the file is "the same at the start" because it is a different variation of the same video, the partial hash may match, but the file size will not. If the size is the same, the chance that the partial hash also matches is very small.
In case nothing matches, which is very likely, the file will be uploaded directly, without wasting battery, CPU and, potentially, RAM on computing a hash that will probably NEVER match.
If something matches, the user could either be prompted to skip hashing (if the file is over the cut-off size) and trust the server, or proceed with the client-side check. Alternatively, the user could be shown which file matches as a duplicate and choose whether to skip the upload.
Avoiding local hashing may also benefit the mobile apps. Calculating a few hashes is not a big deal in the scheme of the whole backup operation, but if you have recorded a 5-minute ProRes video with your iPhone, I'm quite sure you'll appreciate it.
Since the server always makes the final call, we can be sure nothing will ever be duplicated.
Since collisions are very unlikely with file size plus tens of MB of data, we can almost certainly prevent wasted hashing.
Since a match does not automatically skip the upload, and we 1) ask the user to confirm their intentions, 2) bypass hashing, or 3) calculate the full hash, the user will never lose any file.
Obviously, this will require the server to store the hash of only the first N MB of data, but it costs nothing more than a column in the database, as we can just record the calculated hash at the N MB mark and then continue to the full hash.
Eventually, given the entropy of just the file size, we could event implement a "pre-hash" check even for files below the cut-off, it's an extra HTTP request, but may save hashing larger pictures on phones.
For this last one, I'm not sure if making an extra HTTP request is better than calculating a 20-30MB hash, but I believe this had to be considered anyway.
Platform
All reactions