Releases: MALIBA-AI/bambara-text-normalization
Release list
v1.0.4
bambara-text-normalizer v1.0.4
What's Changed
CLI Improvements
- Improve the command-line interface using Typer and Rich.
- Add formatted help output, tables for evaluation metrics, and framed terminal output.
- Preserve plain-text output when piping or redirecting.
- Add comprehensive CLI test coverage.
- Expand CLI documentation with usage examples and available options.
Bug Fixes
- Fix number parsing where trailing punctuation (
.or,) was incorrectly interpreted as the beginning of a decimal value. - Add regression tests covering trailing punctuation and numeric separator handling.
Documentation
- Add a PyPI version badge to the README.
- Update CLI documentation.
Packaging
- Bump package version to 1.0.4.
- Add
typerandrichas package dependencies. - Synchronize package version metadata.
Full Changelog: v1.0.2...v1.0.4
v1.0.2
Release Notes v1.0.2
Full Changelog: v1.0.1...v1.0.2
Highlights
Order-independent numeric expansion (fixes pipeline order bug)
Date, time, measurement, and number normalizers previously relied on rewriting raw text in a fixed call order, causing incorrect output when called out of sequence. This release refactors all four into span-reporting functions (they identify spans instead of mutating text directly), consolidated into a single-pass expansion step. Numeric expansion is now order-independent by construction, with tests verifying this and documentation covering expansion precedence rules.
- Added span primitives for numeric expansion
- Date, time, measurement, and number normalizers now report spans instead of rewriting text
- Added single-pass expansion combining all span types
- Exported the numeric span API
- Added tests for order independence
- Documented expansion precedence
Features (Jan 8–24)
- Contraction expansion mode (expand/contract)
- Number normalization
- Date normalization
- Time/duration normalization
- Measurement normalization
- Docs and automation
- Fixed an invalid classifier issue
Notes
This release directly addresses the pipeline order-dependency bug reported for normalize() the fix removes the dependency rather than just documenting a required call order.
v1.0.1
Bambara Text Normalizer v1.0.1
Release Date: 24th January 2025
License: MIT
What's New
This release adds Measurement Normalization to the TN/ITN toolkit, enabling bidirectional conversion between standard measurement units and Bambara expressions.
New Features
Measurement Normalization (TN/ITN)
from bambara_normalizer import measurement_to_bambara, bambara_to_measurement, normalize
# Weight
measurement_to_bambara(5, "kg")
# → "kilogaramu duuru"
# Length
measurement_to_bambara(100, "m")
# → "mɛtɛrɛ kɛmɛ"
# Volume
measurement_to_bambara(2.5, "L")
# → "litiri fila tomi duuru"
# Area
measurement_to_bambara(3, "ha")
# → "ɛkitari saba"
# Inverse conversion
bambara_to_measurement("kilogaramu duuru")
# → (5, 'kg')
# In text
normalize("A ye 5 kg san", expand_measurements=True)
# → "a ye kilogaramu duuru san"
normalize("So in bɛ 100 m", preset="wer")
# → "so in bɛ mɛtɛrɛ kɛmɛ"Supported Measurement Units
| Category | Units | Bambara |
|---|---|---|
| Weight | kg, g, mg, t | kilogaramu, garamu, miligaramu, tɔni |
| Length | km, m, cm, mm | kilomɛtɛrɛ, mɛtɛrɛ, santimɛtɛrɛ, milimɛtɛrɛ |
| Volume | L, mL | litiri, mililitiri |
| Area | ha, m² | ɛkitari, mɛtɛrɛ kare |
Configuration
from bambara_normalizer import normalize, BambaraNormalizerConfig
# Individual flag
normalize("A ye 5 kg san", expand_measurements=True)
# → "a ye kilogaramu duuru san"
# WER/CER presets enable all expansions by default
config = BambaraNormalizerConfig.for_wer_evaluation()
# expand_numbers=True, expand_dates=True, expand_times=True, expand_measurements=TrueInstallation
pip install bambara-text-normalizeror from github
pip install git+https://github.com/sudoping01/bambara-text-normalization.git
Changelog
Added
measurements.pymodule with weight, length, volume, area TN/ITNexpand_measurementsconfiguration optionmeasurement_to_bambara()functionbambara_to_measurement()functionnormalize_measurements_in_text()functiondenormalize_measurements_in_text()functionis_measurement_word()utility functionget_unit_category()utility function- 30 measurement normalization tests
Changed
- WER/CER presets now enable
expand_measurements=True - Updated README with Measurement Normalization section
Full Documentation
See README.md for complete documentation.
Citation
@software{bambara_tn_itn,
title = {Bambara Text Normalization and Inverse Text Normalization (TN/ITN) Toolkit},
author = {MALIBA-AI},
year = {2025},
version = {2.4.0},
url = {https://github.com/sudoping01/bambara-text-normalization},
note = {A comprehensive TN/ITN framework for Bambara (Bamanankan) with support for numbers, dates, times, measurements, and ASR evaluation}
}
MALIBA-AI 🇲🇱
# Bambara Text Normalizer v1.0.1v1.0.0
Bambara Text Normalizer v1.0.0
License: MIT
Bambara Text Normalizer is a comprehensive text normalization and ASR evaluation framework for Bambara (Bamanankan). It provides bidirectional text normalization (TN/ITN) for numbers, dates, and times, along with intelligent contraction handling and evaluation metrics for speech recognition systems.
Overview
Bambara Text Normalizer is designed to handle text normalization tasks for Bambara language processing. Its modular architecture makes it suitable for integration into ASR pipelines, TTS systems, machine translation, and other NLP applications. The normalizer accounts for valid orthographic variations in Bambara, ensuring fair evaluation metrics that measure actual recognition errors rather than writing convention differences.
Features
| Feature | Description |
|---|---|
| Contraction Modes | Three modes: expand, contract, preserve |
| Number TN/ITN | Bidirectional conversion for integers, decimals, up to millions |
| Date TN/ITN | French (DD-MM-YYYY) and ISO (YYYY-MM-DD) format support |
| Time TN/ITN | Clock time (HH:MM) and duration (XhYmZs) conversion |
| ASR Evaluation | WER, CER, MER, WIL, WIP, DER metrics |
| Disambiguation | Context-aware k'/n' contraction expansion |
Installation
pip install git+https://github.com/sudoping01/bambara-text-normalization.gitUsage
Text Normalization
from bambara_normalizer import normalize
# Contraction expansion (default)
normalize("B'a fɔ k'a ta")
# → "bɛ a fɔ ka a ta"
# Contraction mode
normalize("bɛ a fɔ ka a ta", mode="contract")
# → "b'a fɔ k'a ta"
# Preserve mode
normalize("B'a fɔ", mode="preserve")
# → "b'a fɔ"Number Normalization
from bambara_normalizer import number_to_bambara, bambara_to_number, normalize
# Digits to words
number_to_bambara(123)
# → "kɛmɛ ni mugan ni saba"
# Words to digits
bambara_to_number("kɛmɛ ni mugan ni saba")
# → 123
# In text
normalize("A ye 100 sɔrɔ", expand_numbers=True)
# → "a ye kɛmɛ sɔrɔ"Date Normalization
from bambara_normalizer import date_to_bambara, bambara_to_date, normalize
# Date to Bambara
date_to_bambara(2024, 10, 13)
# → "Oktɔburu tile tan ni saba san baa fila ni mugan ni naani"
# Bambara to date
bambara_to_date("Oktɔburu tile tan ni saba san baa fila ni mugan ni naani")
# → datetime.date(2024, 10, 13)
# In text
normalize("A bɛ na 13-10-2024 la", expand_dates=True)
# → "a bɛ na oktɔburu tile tan ni saba san baa fila ni mugan ni naani la"Time Normalization
from bambara_normalizer import time_to_bambara, duration_to_bambara, normalize
# Clock time
time_to_bambara(7, 30)
# → "Nɛgɛ kaɲɛ wolonwula ni sanga bi saba"
# Duration
duration_to_bambara(hours=1, minutes=30)
# → "lɛrɛ kelen ni miniti bi saba"
# In text
normalize("A nana 7:30 la", expand_times=True)
# → "a nana nɛgɛ kaɲɛ wolonwula ni sanga bi saba la"ASR Evaluation
from bambara_normalizer import evaluate, BambaraEvaluator
# Single evaluation
result = evaluate("B'a fɔ", "bɛ a fɔ")
print(f"WER: {result.wer:.2%}") # → 0.00%
# Batch evaluation
evaluator = BambaraEvaluator(mode="expand")
references = ["k'a ta", "b'a fɔ"]
hypotheses = ["ka a ta", "bɛ a fɔ"]
aggregate, individual = evaluator.evaluate_batch(references, hypotheses)
print(f"Overall WER: {aggregate.wer:.2%}")Configuration Presets
from bambara_normalizer import BambaraNormalizer, BambaraNormalizerConfig
# WER evaluation preset
normalizer = BambaraNormalizer(BambaraNormalizerConfig.for_wer_evaluation())
# CER evaluation preset
normalizer = BambaraNormalizer(BambaraNormalizerConfig.for_cer_evaluation())
# Preserve tones
normalizer = BambaraNormalizer(BambaraNormalizerConfig.preserving_tones())
# Minimal normalization
normalizer = BambaraNormalizer(BambaraNormalizerConfig.minimal())Supported Languages
| Language | Code | Status |
|---|---|---|
| Bambara (Bamanankan) | bm |
✅ Supported |
Technical Specifications
| Specification | Details |
|---|---|
| Python Version | 3.9+ |
| Dependencies | jiwer |
| Test Coverage | 157 tests |
| License | MIT |
References
- [Bambara Reference Corpus](http://cormand.huma-num.fr/)
- [Daba Morphological Analyzer](https://github.com/maslinych/daba)
- [Bamadaba Dictionary](http://cormand.huma-num.fr/bamadaba.html)
- DNAFLA / AMALAN — Bambara standardization body
Citation
@software{bambara_tn_itn,
title = {Bambara Text Normalization and Inverse Text Normalization (TN/ITN) Toolkit},
author = {MALIBA-AI},
year = {2025},
url = {https://github.com/sudoping01/bambara-text-normalization},
note = {A comprehensive TN/ITN framework for Bambara (Bamanankan) with support for numbers, dates, times, and ASR evaluation}
}MALIBA-AI