Skip to content

Releases: Rezarys/ane-embed-pipeline

Build 2 - sentence-transformers/all-MiniLM-L6-v2

Choose a tag to compare

@github-actions github-actions released this 03 Sep 01:33

Benchmark sentence-transformers/all-MiniLM-L6-v2

  • Taille (quant8) : 21.66 Mo
  • Latence mediane : 25.398 ms (runs=30)
  • Cosinus torch vs fp32 : 0.999984
  • Cosinus fp32 vs quant8 : 0.999775
  • compute_unit demande : ComputeUnit.CPU_AND_NE
  • compute_unit_plan (par operation) : {'inconnu (non estime)': 243, 'MLCPUComputeDevice': 180}

v1.0.0

Choose a tag to compare

@Rezarys Rezarys released this 02 Sep 22:41

Première version publique : paquet Swift ANEEmbed (all-MiniLM-L6-v2 quantifié 8 bits pour l'Apple Neural Engine), pipeline CI de conversion/benchmark. Voir README pour les résultats du benchmark et les limites connues.

Build 1 - sentence-transformers/all-MiniLM-L6-v2

Choose a tag to compare

@github-actions github-actions released this 02 Sep 22:43

Benchmark sentence-transformers/all-MiniLM-L6-v2

  • Taille (quant8) : 21.66 Mo
  • Latence mediane : 44.636 ms (runs=30)
  • Cosinus torch vs fp32 : 0.999984
  • Cosinus fp32 vs quant8 : 0.999775
  • compute_unit demande : ComputeUnit.CPU_AND_NE
  • compute_unit_plan (par operation) : {'inconnu (non estime)': 243, 'MLCPUComputeDevice': 180}