Skip to content

Releases: MALIBA-AI/bambara-text-normalization

v1.0.4

Choose a tag to compare

@sudoping01 sudoping01 released this 02 Aug 14:41
6c978ff

bambara-text-normalizer v1.0.4

What's Changed

CLI Improvements

  • Improve the command-line interface using Typer and Rich.
  • Add formatted help output, tables for evaluation metrics, and framed terminal output.
  • Preserve plain-text output when piping or redirecting.
  • Add comprehensive CLI test coverage.
  • Expand CLI documentation with usage examples and available options.

Bug Fixes

  • Fix number parsing where trailing punctuation (. or ,) was incorrectly interpreted as the beginning of a decimal value.
  • Add regression tests covering trailing punctuation and numeric separator handling.

Documentation

  • Add a PyPI version badge to the README.
  • Update CLI documentation.

Packaging

  • Bump package version to 1.0.4.
  • Add typer and rich as package dependencies.
  • Synchronize package version metadata.

Full Changelog: v1.0.2...v1.0.4

v1.0.2

Choose a tag to compare

@sudoping01 sudoping01 released this 26 Jul 15:54
76d3b4d

Release Notes v1.0.2

Full Changelog: v1.0.1...v1.0.2

Highlights

Order-independent numeric expansion (fixes pipeline order bug)
Date, time, measurement, and number normalizers previously relied on rewriting raw text in a fixed call order, causing incorrect output when called out of sequence. This release refactors all four into span-reporting functions (they identify spans instead of mutating text directly), consolidated into a single-pass expansion step. Numeric expansion is now order-independent by construction, with tests verifying this and documentation covering expansion precedence rules.

  • Added span primitives for numeric expansion
  • Date, time, measurement, and number normalizers now report spans instead of rewriting text
  • Added single-pass expansion combining all span types
  • Exported the numeric span API
  • Added tests for order independence
  • Documented expansion precedence

Features (Jan 8–24)

  • Contraction expansion mode (expand/contract)
  • Number normalization
  • Date normalization
  • Time/duration normalization
  • Measurement normalization
  • Docs and automation
  • Fixed an invalid classifier issue

Notes

This release directly addresses the pipeline order-dependency bug reported for normalize() the fix removes the dependency rather than just documenting a required call order.

v1.0.1

Choose a tag to compare

@sudoping01 sudoping01 released this 24 Jan 21:38

Bambara Text Normalizer v1.0.1

Release Date: 24th January 2025

License: MIT


What's New

This release adds Measurement Normalization to the TN/ITN toolkit, enabling bidirectional conversion between standard measurement units and Bambara expressions.


New Features

Measurement Normalization (TN/ITN)

from bambara_normalizer import measurement_to_bambara, bambara_to_measurement, normalize

# Weight
measurement_to_bambara(5, "kg")
# → "kilogaramu duuru"

# Length
measurement_to_bambara(100, "m")
# → "mɛtɛrɛ kɛmɛ"

# Volume
measurement_to_bambara(2.5, "L")
# → "litiri fila tomi duuru"

# Area
measurement_to_bambara(3, "ha")
# → "ɛkitari saba"

# Inverse conversion
bambara_to_measurement("kilogaramu duuru")
# → (5, 'kg')

# In text
normalize("A ye 5 kg san", expand_measurements=True)
# → "a ye kilogaramu duuru san"

normalize("So in bɛ 100 m", preset="wer")
# → "so in bɛ mɛtɛrɛ kɛmɛ"

Supported Measurement Units

Category Units Bambara
Weight kg, g, mg, t kilogaramu, garamu, miligaramu, tɔni
Length km, m, cm, mm kilomɛtɛrɛ, mɛtɛrɛ, santimɛtɛrɛ, milimɛtɛrɛ
Volume L, mL litiri, mililitiri
Area ha, m² ɛkitari, mɛtɛrɛ kare

Configuration

from bambara_normalizer import normalize, BambaraNormalizerConfig
# Individual flag
normalize("A ye 5 kg san", expand_measurements=True)
# → "a ye kilogaramu duuru san"

# WER/CER presets enable all expansions by default
config = BambaraNormalizerConfig.for_wer_evaluation()
# expand_numbers=True, expand_dates=True, expand_times=True, expand_measurements=True

Installation

pip install bambara-text-normalizer

or from github

pip install git+https://github.com/sudoping01/bambara-text-normalization.git

Changelog

Added

  • measurements.py module with weight, length, volume, area TN/ITN
  • expand_measurements configuration option
  • measurement_to_bambara() function
  • bambara_to_measurement() function
  • normalize_measurements_in_text() function
  • denormalize_measurements_in_text() function
  • is_measurement_word() utility function
  • get_unit_category() utility function
  • 30 measurement normalization tests

Changed

  • WER/CER presets now enable expand_measurements=True
  • Updated README with Measurement Normalization section

Full Documentation

See README.md for complete documentation.


Citation

@software{bambara_tn_itn,
  title = {Bambara Text Normalization and Inverse Text Normalization (TN/ITN) Toolkit},
  author = {MALIBA-AI},
  year = {2025},
  version = {2.4.0},
  url = {https://github.com/sudoping01/bambara-text-normalization},
  note = {A comprehensive TN/ITN framework for Bambara (Bamanankan) with support for numbers, dates, times, measurements, and ASR evaluation}
}

MALIBA-AI 🇲🇱

# Bambara Text Normalizer v1.0.1

v1.0.0

Choose a tag to compare

@sudoping01 sudoping01 released this 20 Jan 00:13

Bambara Text Normalizer v1.0.0

License: MIT

Bambara Text Normalizer is a comprehensive text normalization and ASR evaluation framework for Bambara (Bamanankan). It provides bidirectional text normalization (TN/ITN) for numbers, dates, and times, along with intelligent contraction handling and evaluation metrics for speech recognition systems.


Overview

Bambara Text Normalizer is designed to handle text normalization tasks for Bambara language processing. Its modular architecture makes it suitable for integration into ASR pipelines, TTS systems, machine translation, and other NLP applications. The normalizer accounts for valid orthographic variations in Bambara, ensuring fair evaluation metrics that measure actual recognition errors rather than writing convention differences.


Features

Feature Description
Contraction Modes Three modes: expand, contract, preserve
Number TN/ITN Bidirectional conversion for integers, decimals, up to millions
Date TN/ITN French (DD-MM-YYYY) and ISO (YYYY-MM-DD) format support
Time TN/ITN Clock time (HH:MM) and duration (XhYmZs) conversion
ASR Evaluation WER, CER, MER, WIL, WIP, DER metrics
Disambiguation Context-aware k'/n' contraction expansion

Installation

pip install git+https://github.com/sudoping01/bambara-text-normalization.git

Usage

Text Normalization

from bambara_normalizer import normalize

# Contraction expansion (default)
normalize("B'a fɔ k'a ta")
# → "bɛ a fɔ ka a ta"

# Contraction mode
normalize("bɛ a fɔ ka a ta", mode="contract")
# → "b'a fɔ k'a ta"

# Preserve mode
normalize("B'a fɔ", mode="preserve")
# → "b'a fɔ"

Number Normalization

from bambara_normalizer import number_to_bambara, bambara_to_number, normalize

# Digits to words
number_to_bambara(123)
# → "kɛmɛ ni mugan ni saba"

# Words to digits
bambara_to_number("kɛmɛ ni mugan ni saba")
# → 123

# In text
normalize("A ye 100 sɔrɔ", expand_numbers=True)
# → "a ye kɛmɛ sɔrɔ"

Date Normalization

from bambara_normalizer import date_to_bambara, bambara_to_date, normalize

# Date to Bambara
date_to_bambara(2024, 10, 13)
# → "Oktɔburu tile tan ni saba san baa fila ni mugan ni naani"

# Bambara to date
bambara_to_date("Oktɔburu tile tan ni saba san baa fila ni mugan ni naani")
# → datetime.date(2024, 10, 13)

# In text
normalize("A bɛ na 13-10-2024 la", expand_dates=True)
# → "a bɛ na oktɔburu tile tan ni saba san baa fila ni mugan ni naani la"

Time Normalization

from bambara_normalizer import time_to_bambara, duration_to_bambara, normalize

# Clock time
time_to_bambara(7, 30)
# → "Nɛgɛ kaɲɛ wolonwula ni sanga bi saba"

# Duration
duration_to_bambara(hours=1, minutes=30)
# → "lɛrɛ kelen ni miniti bi saba"

# In text
normalize("A nana 7:30 la", expand_times=True)
# → "a nana nɛgɛ kaɲɛ wolonwula ni sanga bi saba la"

ASR Evaluation

from bambara_normalizer import evaluate, BambaraEvaluator

# Single evaluation
result = evaluate("B'a fɔ", "bɛ a fɔ")
print(f"WER: {result.wer:.2%}")  # → 0.00%

# Batch evaluation
evaluator = BambaraEvaluator(mode="expand")
references = ["k'a ta", "b'a fɔ"]
hypotheses = ["ka a ta", "bɛ a fɔ"]
aggregate, individual = evaluator.evaluate_batch(references, hypotheses)
print(f"Overall WER: {aggregate.wer:.2%}")

Configuration Presets

from bambara_normalizer import BambaraNormalizer, BambaraNormalizerConfig

# WER evaluation preset
normalizer = BambaraNormalizer(BambaraNormalizerConfig.for_wer_evaluation())

# CER evaluation preset
normalizer = BambaraNormalizer(BambaraNormalizerConfig.for_cer_evaluation())

# Preserve tones
normalizer = BambaraNormalizer(BambaraNormalizerConfig.preserving_tones())

# Minimal normalization
normalizer = BambaraNormalizer(BambaraNormalizerConfig.minimal())

Supported Languages

Language Code Status
Bambara (Bamanankan) bm ✅ Supported

Technical Specifications

Specification Details
Python Version 3.9+
Dependencies jiwer
Test Coverage 157 tests
License MIT

References


Citation

@software{bambara_tn_itn,
  title = {Bambara Text Normalization and Inverse Text Normalization (TN/ITN) Toolkit},
  author = {MALIBA-AI},
  year = {2025},
  url = {https://github.com/sudoping01/bambara-text-normalization},
  note = {A comprehensive TN/ITN framework for Bambara (Bamanankan) with support for numbers, dates, times, and ASR evaluation}
}

MALIBA-AI