Build 1 - sentence-transformers/all-MiniLM-L6-v2
·
9 commits
to main
since this release
Benchmark sentence-transformers/all-MiniLM-L6-v2
- Taille (quant8) : 21.66 Mo
- Latence mediane : 44.636 ms (runs=30)
- Cosinus torch vs fp32 : 0.999984
- Cosinus fp32 vs quant8 : 0.999775
- compute_unit demande : ComputeUnit.CPU_AND_NE
- compute_unit_plan (par operation) : {'inconnu (non estime)': 243, 'MLCPUComputeDevice': 180}