Skip to content

v1.0.0

Choose a tag to compare

@sudoping01 sudoping01 released this 20 Jan 00:13
· 61 commits to main since this release

Bambara Text Normalizer v1.0.0

License: MIT

Bambara Text Normalizer is a comprehensive text normalization and ASR evaluation framework for Bambara (Bamanankan). It provides bidirectional text normalization (TN/ITN) for numbers, dates, and times, along with intelligent contraction handling and evaluation metrics for speech recognition systems.


Overview

Bambara Text Normalizer is designed to handle text normalization tasks for Bambara language processing. Its modular architecture makes it suitable for integration into ASR pipelines, TTS systems, machine translation, and other NLP applications. The normalizer accounts for valid orthographic variations in Bambara, ensuring fair evaluation metrics that measure actual recognition errors rather than writing convention differences.


Features

Feature Description
Contraction Modes Three modes: expand, contract, preserve
Number TN/ITN Bidirectional conversion for integers, decimals, up to millions
Date TN/ITN French (DD-MM-YYYY) and ISO (YYYY-MM-DD) format support
Time TN/ITN Clock time (HH:MM) and duration (XhYmZs) conversion
ASR Evaluation WER, CER, MER, WIL, WIP, DER metrics
Disambiguation Context-aware k'/n' contraction expansion

Installation

pip install git+https://github.com/sudoping01/bambara-text-normalization.git

Usage

Text Normalization

from bambara_normalizer import normalize

# Contraction expansion (default)
normalize("B'a fɔ k'a ta")
# → "bɛ a fɔ ka a ta"

# Contraction mode
normalize("bɛ a fɔ ka a ta", mode="contract")
# → "b'a fɔ k'a ta"

# Preserve mode
normalize("B'a fɔ", mode="preserve")
# → "b'a fɔ"

Number Normalization

from bambara_normalizer import number_to_bambara, bambara_to_number, normalize

# Digits to words
number_to_bambara(123)
# → "kɛmɛ ni mugan ni saba"

# Words to digits
bambara_to_number("kɛmɛ ni mugan ni saba")
# → 123

# In text
normalize("A ye 100 sɔrɔ", expand_numbers=True)
# → "a ye kɛmɛ sɔrɔ"

Date Normalization

from bambara_normalizer import date_to_bambara, bambara_to_date, normalize

# Date to Bambara
date_to_bambara(2024, 10, 13)
# → "Oktɔburu tile tan ni saba san baa fila ni mugan ni naani"

# Bambara to date
bambara_to_date("Oktɔburu tile tan ni saba san baa fila ni mugan ni naani")
# → datetime.date(2024, 10, 13)

# In text
normalize("A bɛ na 13-10-2024 la", expand_dates=True)
# → "a bɛ na oktɔburu tile tan ni saba san baa fila ni mugan ni naani la"

Time Normalization

from bambara_normalizer import time_to_bambara, duration_to_bambara, normalize

# Clock time
time_to_bambara(7, 30)
# → "Nɛgɛ kaɲɛ wolonwula ni sanga bi saba"

# Duration
duration_to_bambara(hours=1, minutes=30)
# → "lɛrɛ kelen ni miniti bi saba"

# In text
normalize("A nana 7:30 la", expand_times=True)
# → "a nana nɛgɛ kaɲɛ wolonwula ni sanga bi saba la"

ASR Evaluation

from bambara_normalizer import evaluate, BambaraEvaluator

# Single evaluation
result = evaluate("B'a fɔ", "bɛ a fɔ")
print(f"WER: {result.wer:.2%}")  # → 0.00%

# Batch evaluation
evaluator = BambaraEvaluator(mode="expand")
references = ["k'a ta", "b'a fɔ"]
hypotheses = ["ka a ta", "bɛ a fɔ"]
aggregate, individual = evaluator.evaluate_batch(references, hypotheses)
print(f"Overall WER: {aggregate.wer:.2%}")

Configuration Presets

from bambara_normalizer import BambaraNormalizer, BambaraNormalizerConfig

# WER evaluation preset
normalizer = BambaraNormalizer(BambaraNormalizerConfig.for_wer_evaluation())

# CER evaluation preset
normalizer = BambaraNormalizer(BambaraNormalizerConfig.for_cer_evaluation())

# Preserve tones
normalizer = BambaraNormalizer(BambaraNormalizerConfig.preserving_tones())

# Minimal normalization
normalizer = BambaraNormalizer(BambaraNormalizerConfig.minimal())

Supported Languages

Language Code Status
Bambara (Bamanankan) bm ✅ Supported

Technical Specifications

Specification Details
Python Version 3.9+
Dependencies jiwer
Test Coverage 157 tests
License MIT

References


Citation

@software{bambara_tn_itn,
  title = {Bambara Text Normalization and Inverse Text Normalization (TN/ITN) Toolkit},
  author = {MALIBA-AI},
  year = {2025},
  url = {https://github.com/sudoping01/bambara-text-normalization},
  note = {A comprehensive TN/ITN framework for Bambara (Bamanankan) with support for numbers, dates, times, and ASR evaluation}
}

MALIBA-AI