GitHub - piskvorky/ann-benchmarks: Benchmarks of approximate nearest neighbor libraries in Python

Benchmarking nearest neighbors

This project contains some tools to benchmark various implementations of approximate nearest neighbor (ANN) search.

Evaluated

Annoy
FLANN
scikit-learn: LSHForest, KDTree, BallTree
PANNS
NearPy
KGraph

Data sets

GloVe
SIFT

Motivation

Doing fast searching of nearest neighbors in high dimensional spaces is an increasingly important problem, but with little attempt at objectively comparing methods.

Install

Clone the repo and run bash install.sh. This will install all libraries as well as downloading and preprocessing all data sets. It could take a while. It has been tested in Ubuntu 14.04.

There is also a Docker image available under erikbern/ann containing all libraries and data sets.

Principles

Everyone is welcome to submit pull requests with tweaks and changes to how each library is being used.
In particular: if you are the author of any of these libraries, and you think the benchmark can be improved, consider making the improvement and submitting a pull request.
This is meant to be an ongoing project and represent the current state.
Make everything easy to replicate, including installing and preparing the datasets.
To make it simpler, look only at the precision-performance tradeoff.
Try many different values of parameters for each library and ignore the points that are not on the precision-performance frontier.
High-dimensional datasets with approximately 100-1000 dimensions. This is challenging but also realistic. Not more than 1000 dimensions because those problems should probably be solved by doing dimensionality reduction separately.
Use single core benchmarks. I believe most real world scenarios could be parallelized in other ways (eg. do multiple queries in parallel).
Avoid extremely costly index building (more than several hours).
Focus on datasets that fit in RAM. Out of core ANN could be the topic of a later comparison.
Do proper train/test set of index data and query points.

Results

This is very much a work in progress... more results coming later!

1.19M vectors from GloVe (100 dimensions, trained from tweets), cosine similarity, run on a c4.4xlarge instance on EC2.

https://raw.github.com/erikbern/ann-benchmarks/master/results/glove.png

1M SIFT features (128 dimensions), Euclidean distance, also run on a c4.4xlarge:

https://raw.github.com/erikbern/ann-benchmarks/master/results/sift.png

References

sim-shootout by Radim Řehůřek
NonMetricSpaceLib
This blog post

Name		Name	Last commit message	Last commit date
Latest commit History 63 Commits
install		install
results		results
Dockerfile		Dockerfile
README.rst		README.rst
ann_benchmarks.py		ann_benchmarks.py
install.sh		install.sh
plot.py		plot.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

install

install

results

results

Dockerfile

Dockerfile

README.rst

README.rst

ann_benchmarks.py

ann_benchmarks.py

install.sh

install.sh

plot.py

plot.py

Repository files navigation

Benchmarking nearest neighbors

Evaluated

Data sets

Motivation

Install

Principles

Results

References

About

Releases

Packages

Languages

piskvorky/ann-benchmarks

Folders and files

Latest commit

History

Repository files navigation

Benchmarking nearest neighbors

Evaluated

Data sets

Motivation

Install

Principles

Results

References

About

Resources

Stars

Watchers

Forks

Languages