Skip to content

Release v0.1.0

Latest

Choose a tag to compare

@gowthamrao gowthamrao released this 03 Sep 22:32
0005025
Documentation, test coverage

* Add documentation

* Improve test coverage for several modules (#34)

This commit improves the test coverage for the following modules:
- `extract_regex_paragraphs_udf.py` (from 0% to 45%)
- `medical_code_extractor.py` (from 67% to 96%)
- `apply_regex_functions.py` (from 72% to 100%)

New test files were added for `extract_regex_paragraphs_udf.py` and `apply_regex_functions.py`.
New test cases were added to `test_medical_code_extractor.py` to cover more functionality.

Also, `pyarrow` is added as a dependency in `pyproject.toml` as it is required for the PySpark pandas UDF tests.


* docs: Create detailed documentation for the package

This commit adds a comprehensive `README.md` file that serves as the main documentation for the `pyRegularExpression` package.

The new `README.md` file includes:

- An overview of the package and installation instructions.
- A master catalog of all the "finder" modules, including a brief description and a list of all the functions within each module.
- Usage examples for finder functions and helper functions.

* Make package ready for PyPI submission

This change addresses an issue where the package would fail to import if the optional `pyspark` dependency was not installed.

The `src/pyregularexpression/__init__.py` file has been modified to wrap the dynamic import of modules in a `try...except ModuleNotFoundError` block. This allows the package to be imported even if `pyspark` is not present.

The tests for the spark functionality in `tests/test_extract_regex_paragraphs_udf.py` have been updated to use `pytest.importorskip("pyspark")`, so they are skipped when `pyspark` is not installed.