Faster CUDA kernel and dispatch to other dtypes
- Add a script to benchmark kernels between torchdtw version:
benchmark/src/dtw_benchmark/compare_release.py
- Dispatch to all kinds of dtypes for distances and sx in CPU and CUDA kernels.
- Faster CUDA kernel:
[---------------- dtw / cpu -----------------]
| released 0.2.0 | current
1 threads: -----------------------------------
64x64 | 9.2 | 9.4
128x128 | 23.9 | 24.4
256x256 | 85.6 | 86.1
512x512 | 327.1 | 327.8
1024x1024 | 1291.0 | 1295.6
Times are in microseconds (us).
[-------------------- dtw_batch / cpu --------------------]
| released 0.2.0 | current
1 threads: ------------------------------------------------
n=4x4, s=128x128, sym | 72.8 | 72.8
n=4x4, s=128x128, asym | 155.6 | 154.9
n=8x8, s=256x256, sym | 665.9 | 663.5
n=4x8, s=256x512, asym | 1899.2 | 1901.0
Times are in microseconds (us).
[---------------- dtw / cuda ----------------]
| released 0.2.0 | current
1 threads: -----------------------------------
64x64 | 96.7 | 81.3
128x128 | 149.4 | 123.6
256x256 | 288.0 | 228.2
512x512 | 745.2 | 560.8
1024x1024 | 2581.7 | 1159.4
Times are in microseconds (us).
[-------------------- dtw_batch / cuda -------------------]
| released 0.2.0 | current
1 threads: ------------------------------------------------
n=4x4, s=128x128, sym | 100.0 | 81.3
n=4x4, s=128x128, asym | 104.2 | 84.1
n=8x8, s=256x256, sym | 305.3 | 153.3
n=4x8, s=256x512, asym | 514.8 | 353.0