No description, website, or topics provided.
Python C++ C
Latest commit 0153ef0 Apr 28, 2016 @larsmans larsmans Merge pull request #3 from larsmans/master
simple benchmark script
Permalink
Failed to load latest commit information.
leven ship headers in tarball Nov 19, 2013
tests fix test setup Nov 20, 2013
.gitignore fix test setup Nov 20, 2013
LICENSE transfer copyright to my employer; relicensed (Apache2) Nov 18, 2013
MANIFEST.in ship headers in tarball Nov 19, 2013
README.rst
benchmark.py simple benchmark script Jan 16, 2014
setup.py fix test setup Nov 20, 2013

README.rst

Leven

Levenshtein edit distance library for Python, Apache-licensed. Written by Lars Buitinck, Netherlands eScience Center, with contributions from Isaac Sijaranamual, University of Amsterdam.

Performs distance computations on either byte strings or Unicode codepoints.

Installation

Make sure you have Cython and a C++ compiler installed:

pip install cython

Installing a C++ compiler is so platform-dependent that I won't show instructions. Consult your package manager.

Then:

python setup.py install

To run the tests, but not to actually use leven, you need six and Nose.

Usage

>>> from leven import levenshtein
>>> levenshtein("hello, world!", "goodbye, cruel world!")
13

About the implementation

The core algorithms have been implemented in C++. I used this instead of C to get templates, easier memory management and a better standard library, so the C++ code probably looks C-ish.

Todo

  • Implement Ukkonen's algorithm for bounded Levenshtein distance
  • Implement Levenshtein automata for fast neighbor search in string spaces
  • Implement weighted Levenshtein distance