Skip to content

Fix locking bug

Pre-release
Pre-release

Choose a tag to compare

@Charl-AI Charl-AI released this 27 Oct 14:44
· 20 commits to main since this release
3873851

Implementing mutex locks involved having one lock per slot in the cache. Somewhere between 50k and 100k dataset samples, this fails because Python itself couldn't create a list with that many lock objects.

Solved by removing all mutex locking. This is generally fine -- in each epoch, the same datapoint will never be accessed twice anyway.

If reintroducing in future, there are basically three options:

  • lock for each slot (like before), but all locking gets automatically disabled for large datasets.
  • one lock for the whole object, but it protects writes only. This would slow down the first epoch, but then be fine.
  • Implement both of the above, with a flag for advanced users to toggle the options