This repository builds a Python package that installs a pii-extract-base
plugin to performs PII detection for text data based on regular expressions
(with optional context). The name of the plugin entry point is
piisa-detectors-regex.
The PII Tasks in the package are structured by language & country, since many of the PII elements are language- and/or -country dependent.
The package
- needs at least Python 3.8
- needs the pii-data and the pii-extract-base base packages
- uses the regex package (instead of the standard
repackage in the core Python library) - uses the python-stdnum package to validate many identifiers (and the python-phonenumbers to validate phone numbers)
The package does not have any user-facing entry points, and it is [used automatically] by the PIISA framework.
The provided Makefile can be used to process the package:
make pkgwill build the Python package, creating a file that can be installed withpipmake unitwill launch all unit tests (using pytest, so pytest must be available)make installwill install the package in a Python virtualenv. The virtualenv will be chosen as, in this order:- the one defined in the
VENVenvironment variable, if it is defined - if there is a virtualenv activated in the shell, it will be used
- otherwise, a default is chosen as
/opt/venv/bigscience(it will be created if it does not exist)
- the one defined in the
To add a new PII processing task, please see the contributing instructions.