Full Changelog: v0.7.2...v0.8.0
Added
- Added spaced ellipsis normalization and token realignment for model punctuation
collapses
Changed
- Changed German area, volume, hectare, and speed quantities to use the shared
spokenform grammar - Changed the reviewed seven-language pipeline to use shared written-to-spoken
preparation - Changed migrated-language written-to-spoken semantics to use Spokenform 0.2.6 and
raised the abbr2words floor to 0.2.9 - Added Spokenform diagnostics and provenance tracking for upstream warnings
- Changed migrated-language processing order so Spokenform receives source syntax before
model-punctuation cleanup - Changed Spokenform source replacement handling to preserve exact spans and retain
provenance in token metadata - Changed dependency floors to abbr2words>=0.2.9 and spokenform>=0.2.6 with handoff
regressions - Changed English contextual N.0 release-label pronunciation through the spokenform
preparation layer - Changed German written-to-spoken semantic preparation to use spokenform while
preserving kokorog2p overrides - Changed French written-to-spoken semantic preparation to use spokenform and retained
number helpers as deprecated - Changed Spanish written-to-spoken semantic preparation to use spokenform and retained
dialect phoneme behavior - Changed Italian written-to-spoken semantic preparation to use spokenform and retained
Italian phoneme behavior - Changed Portuguese written-to-spoken semantic preparation to use spokenform and
retained Portuguese phoneme behavior - Changed Czech written-to-spoken semantic preparation to use spokenform and retained
Czech phonological rules - Changed English written-to-spoken semantic preparation to use spokenform and retained
English G2P behavior - Changed abbr2words to serve as shared lexical abbreviation source for all seven
languages
Fixed
- Fixed exact Spokenform replacement spans when sentence-final periods appear in
structured replacements - Fixed English abbreviation and core pipeline alignment compatibility with current
Spokenform semantics
Documentation
- Documented English semantic ownership and the Portuguese migration boundary
- Documented that accepted Spokenform semantics are preserved literally in the English
API documentation - Documented the abbr2words-to-spokenform-to-kokorog2p ownership boundary and
seven-language migration scope
Quality
- Added runtime and parity fixture assets to portable Codecrate packs
- Added compact Spokenform handoff regression tests covering special characters and
symbol preservation