Releases: xqms/gpu_claim
Releases · xqms/gpu_claim
v0.10.4
Compare
Sorry, something went wrong.
No results found
xqms
released this
01 Apr 07:32
Changes:
try again on GPU release error, hopefully fixes "Card N is still in use" message in most cases.
Full Changelog : v0.10.3...v0.10.4
v0.10.3
Compare
Sorry, something went wrong.
No results found
xqms
released this
06 Mar 13:07
Changes:
Relax depencies for the debian package to be compatible with older NVIDIA drivers
Full Changelog : v0.10.2...v0.10.3
v0.10.2
Compare
Sorry, something went wrong.
No results found
xqms
released this
06 Mar 09:09
Changes:
Rename debian package to gpu-claim
Full Changelog : v0.10.1...v0.10.2
v0.10.1
Compare
Sorry, something went wrong.
No results found
xqms
released this
10 Jan 16:35
Changes:
Correctly pass through environment variables like LD_LIBRARY_PATH or LD_PRELOAD (see commit e4a1f93 for details).
v0.10.0
Compare
Sorry, something went wrong.
No results found
xqms
released this
19 Sep 10:19
Changes:
Support the --card option, which lets you run a job in parallel on the GPUs which you already own.
v0.9.0
Compare
Sorry, something went wrong.
No results found
xqms
released this
25 Aug 13:07
Changes:
Major code refactoring
Nicer handling of client <-> server connection losses
v0.8.0
Compare
Sorry, something went wrong.
No results found
xqms
released this
24 Aug 15:45
Changes:
Make sure GPUs are freed directly if the gpu process is killed (e.g. by VSCode)
Remove the claim command, since nobody uses it anyway and it interferes with the above
Set CUDA_VISIBLE_DEVICES with numeric IDs, anything else confuses older pytorch versions (torch.cuda.device_count() returns 0)
Fix NVML struct versioning
Print idle time in seconds rather than in minutes
Decrease idle timeout to 1 min
v0.7.1
Compare
Sorry, something went wrong.
No results found
xqms
released this
24 Aug 10:17
Changes:
Use versioned NVML API - fixes ABI breakage on newest NVIDIA drivers (535.86.10)
v0.7.0
Compare
Sorry, something went wrong.
No results found
xqms
released this
21 Aug 14:42
Changes:
Isolate /dev/shm for each GPU container. This ensures correct cleanup (torch.multiprocessing, I'm looking at you!).
Ensure that the initramfs is rebuilt after installation of this package, so that the correct nvidia flags are set.
v0.6.0
Compare
Sorry, something went wrong.
No results found
xqms
released this
04 Aug 13:44
Changes:
Force exit of all descendant user processes. Sometimes people use complicated python multiprocessing pipelines which leave stuck
child processes behind. Using Linux PID namespaces, we can make sure that all processes are killed as soon as the main process exits.