Skip to content

Releases: KakaruHayate/tifa.cpp

v0.1.6 — Japanese kanji G2P (MeCab + UniDic) & one-click dictionary installer

Choose a tag to compare

@github-actions github-actions released this 04 Oct 11:40

v0.1.6 — Japanese kanji G2P (MeCab + UniDic) & one-click dictionary installer

日语汉字 G2P(MeCab + UniDic)与一键词典安装

Highlights / 重点

  • Japanese kanji lyrics now work. The japanese-mecab converter is a real
    MeCab + UniDic pipeline, exactly like upstream: MeCab segments the lyrics,
    UniDic supplies the whole-word kana readings (pron), and every reading
    candidate goes to the aligner. A kanji word without a reading fails per file
    instead of silently producing wrong phonemes.
    日语汉字歌词可用。 japanese-mecab 现在与上游一致:MeCab 分词,UniDic
    给出整词假名读音(pron),全部读音候选交给对齐器;无读音的汉字词按文件报
    错,不会静默给错误音素。

  • Kana lyrics keep working out of the box. Without the dictionary the same
    config id falls back to the kana converter (v0.1.5 behaviour, byte-identical
    output).
    假名歌词照旧开箱即用。 未装词典时同一配置 id 自动回退假名转换器(v0.1.5
    行为,输出逐字节一致)。

  • One-click dictionary install in the GUI. The environment panel has a new
    "日语词典" row: downloads with progress and installs to models/unidic/
    automatically.
    GUI 一键安装。 顶栏环境区新增"日语词典"行:带进度下载并自动装到
    models/unidic/。

  • unidic-lite-dicdir.zip ships as its own release asset (~49 MB packed,
    ~260 MB unpacked) — deliberately not inside the platform bundles. It is
    probed at models/unidic/ (the converter's unidic_dir kwarg overrides);
    bundle README.txt and docs/usage.md explain it. MeCab is BSD-3 and UniDic
    BSD/GPL/LGPL — see THIRD_PARTY_NOTICES.md.
    词典为独立 release 资产(压缩约 49 MB,解压约 260 MB),刻意不进平台包。
    CLI/GUI 在 models/unidic/ 自动探测(转换器 unidic_dir kwarg 可覆盖);
    详见 THIRD_PARTY_NOTICES.md 与包内 README.txt。

For OpenUtau users / OpenUtau 用户

The .oudep packages are unchanged: OpenUtau passes phones directly and never
runs G2P, so this release does not rebuild them.
.oudep 无需更新:OpenUtau 直接传入音素、不走 G2P,因此本次不重建 oudep。

Verified / 已验证

可愛いから好きになったなんて with -l ja on the stock model + dictionary:
1 ok, 25 phones, the texts tier keeps the kanji surfaces
(可愛い/から/好き/に/なっ/た/なんて), phones
k a w a i | k a r a | s u k i | n i | n a cl | t a | n a N t e
(促音→cl,ん→N,长音按上游约定丢弃). Stock model + dictionary, -l ja:
1 ok,25 phones,texts 层保留汉字词面;假名回退与 v0.1.5 完全一致。
The GUI one-click install was exercised against this very asset, and the
dictionary archive was downloaded and checked from this release.
GUI 一键安装即从本 release 下载该资产;词典压缩包亦已下载复核。

Thanks / 致谢

MeCab by Taku Kudo / NTT (BSD), UniDic by the UniDic Consortium (BSD),
unidic-lite packaging.

v0.1.5

Choose a tag to compare

@github-actions github-actions released this 04 Oct 06:18

What's Changed

  • fix(ui): resolve file and folder picker buttons not adding inputs by @Xiantaidu in #6
  • 2pass: keep the phrase line, and label the gaps by @KakaruHayate in #7
  • breath: port FoxBreatheLabeler as a second AP detector by @KakaruHayate in #8
  • breath: FoxBreatheLabeler as the default AP detector (re-opens #8) by @KakaruHayate in #9

New Contributors

Full Changelog: v0.1.4...v0.1.5

v0.1.4

Choose a tag to compare

@github-actions github-actions released this 03 Oct 15:36

Full Changelog: v0.1.3...v0.1.4

v0.1.3

Choose a tag to compare

@github-actions github-actions released this 02 Oct 04:18

tifa.cpp v0.1.3

两件事:PFML 1.0 一致性补齐,以及 word 层语言解析(三态)。命令行对齐结果不受影响。

新增

  • 语言三态。PFML 1.0 的语言有三种状态,此前被压成了一个字符串:

    PFML 写法 含义 v0.1.2 v0.1.3
    不写 language 继承外层 scope 报 "" 继承最近 scope 的语言
    language="" 无语言(清空) "" ""
    language-kind="any" 任意语言 接受后忽略 独立状态 language_is_any
    language="zh" 该标签 "zh" "zh"
  • 公开 API 相应扩展:G2PWord 与 G2PWordCandidates 新增
    bool language_is_any;G2PWord::language 现在返回解析后的值。

  • to_pfml 把 ANY 写回 language-kind="any"(且不写 language 属性),往返保持一致。

修正(PFML 1.0 一致性)

输入 规范 v0.1.2
<!-- 注释 --> 忽略 解析报错
<![CDATA[重]]> 普通字符数据 解析报错
未知属性(如 bogus="1") 报错 静默忽略

现在注释与 CDATA 可用,未知属性按规范报错(拼错属性名不会再被静默吞掉)。

兼容性

  • 命令行对齐结果零变化:8 个覆盖各语言状态的输入(普通、scope 内、异语言
    scope、多语言 -l zh,yue、显式多标签属性、裸音素、清空、ANY)与 v0.1.2
    产出逐字节相同的 TextGrid。
  • 显式多语言属性(language="zh,yue")仍按首个标签解析,与之前一致。
  • 其余修正只影响此前报错或标错的输入。

给 MaxLabel(PFML 编辑器)

  • G2PWord::language 的语义变了:对于省略 language 的 word,现在返回外层
    scope 的语言,以前是空串。如果编辑器把空串当作"未确定"来显示,需要改用
    新规则:空串 = 无语言(或已清空),language_is_any = ANY。
  • G2PWord::language_is_any 是新字段,G2PWordCandidates 上也有。
  • 依赖 pin 可以从 v0.1.2 移到 v0.1.3。
  • 验收清单:KakaruHayate/MaxLabel#1 正文里的表格;编辑器侧的
    tests/test_pfml_conformance.cpp 是对应落点。

下载

与 v0.1.2 相同:三平台 GUI 自包含包(tifa-label-<平台>)+ 三平台 CLI 包
(tifa-cli-<平台>-{full,q4})+ BreathLab 模型源文件。用法见包内 USAGE.md。
命令行务必带 -l(如 -l zh)。

CUDA 包由独立的 cuda workflow 编译,完成后自动补挂到本 release
(tifa-cuda-<平台>-{full,q4}.tar.gz)。


English

Two things: PFML 1.0 conformance, and word-level language resolution
(three states). Alignment output is unaffected.

New — a word's language now resolves against its enclosing scope: omitted
inherits, language="" clears, language-kind="any" is a distinct state.
G2PWord and G2PWordCandidates gain bool language_is_any, and
G2PWord::language now reports the resolved value. to_pfml writes ANY
back as language-kind="any".

Fixed — comments (<!-- -->) and CDATA are accepted instead of failing to
parse, and unknown attributes are errors rather than being silently ignored.

Compatible — 8 inputs covering every language state produce byte-identical
TextGrids against v0.1.2, and an explicit multi-language attribute still
resolves through its first tag.

For MaxLabel — G2PWord::language now reports the enclosing scope's
language for a word that omits it (it used to be empty), and
language_is_any is new. The pin can move to v0.1.3; the acceptance table is
in KakaruHayate/MaxLabel#1.

Same packaging as v0.1.2. Pass -l (e.g. -l zh) on the command line. CUDA
packages are appended to this release when the separate cuda build finishes.

v0.1.2

Choose a tag to compare

@github-actions github-actions released this 02 Oct 03:13

tifa.cpp v0.1.2

这一版把 G2P / PFML 抽成可以脱离 ggml 单独链接的库,并补全了 PFML 1.0 的完整发音树。

新增

  • 公开 G2P API(include/tifa_ggml/g2p.h):文本/PFML 转词(convert /
    convert_pfml)、写回 PFML(to_pfml,带 round-trip 保证)、纯语法校验
    (validate_pfml,不需要模型)、候选读音(candidates)、音素校验
    (resolve_phoneme)。
  • ggml-free 构建:-DTIFA_GGML_BUILD_MODEL=OFF 只构建 tifa_ggml_g2p,
    完全不拉取 ggml,适合"对齐器之前"的标注编辑器(如
    MaxLabel):
    set(TIFA_GGML_BUILD_MODEL OFF)
    add_subdirectory(tifa.cpp)
    target_link_libraries(your_editor PRIVATE tifa_ggml::g2p)
  • 完整 PFML 1.0 发音树:<reading> / <path> / <group> 嵌套,<phoneme>
    的 language / symbol 属性,以及省略容器只写音素的片段。

修正

  • 标签填充只在属性缺失时生效:text="" / script="" 是作者显式给的空值,
    往返后保持不变(此前会被音素"填"上,破坏 round-trip 保证)。
  • <word><phoneme>a</phoneme><phoneme>b</phoneme></word> 此前解析成空序列直接
    报错,现在正常。
  • <phoneme language="zh">ong</phoneme> 现在解析为 zh/ong(此前忽略 language)。
  • <word text="重" ...> 现在认 text 属性(此前忽略,用音素拼接当词形)。

兼容性

<word phonemes=...>、裸 <phoneme> 串、<scope> 这三种经典写法行为完全不变
(同一批输入对比,输出的 TextGrid 逐字节相同)。上面几条"修正"只影响此前报错或
标错的输入。

下载

与 v0.1.1 相同的布局:三平台 GUI 自包含包(tifa-label-<平台>)+ 三平台
CLI 包(tifa-cli-<平台>-{full,q4})+ BreathLab 模型源文件。用法见包内
USAGE.md(中英双语)。命令行务必带 -l(如 -l zh)。

CUDA 包由独立的 cuda workflow 编译,完成后自动补挂到本 release
(tifa-cuda-<平台>-{full,q4}.tar.gz)。


English

G2P and PFML are now a library you can link without ggml, and the PFML 1.0
pronunciation tree is complete.

New — a public API in include/tifa_ggml/g2p.h (convert, convert_pfml,
to_pfml with a round-trip guarantee, model-free validate_pfml, candidates,
resolve_phoneme); a ggml-free build (-DTIFA_GGML_BUILD_MODEL=OFF) that
fetches no ggml at all, for editors that sit before the aligner; and full
<reading>/<path>/<group> trees with language/symbol phoneme
attributes.

Fixed — label filling now applies only when an attribute is absent, so an
explicit text=""/script="" survives a round trip; <word> with <phoneme>
children no longer parses to an empty sequence; phoneme-level language and the
word-level text attribute are honoured.

Compatible — the three classic forms (<word phonemes=...>, bare
<phoneme> runs, <scope>) are byte-for-byte unchanged; the fixes only affect
inputs that previously errored or were mislabelled.

Same packaging as v0.1.1. Pass -l (e.g. -l zh) on the command line. CUDA
packages are appended to this release when the separate cuda build finishes.

tifa-ggml OpenUtau dependency packages (q4)

Choose a tag to compare

@KakaruHayate KakaruHayate released this 02 Oct 15:28

OpenUtau dependency packages (.oudep) for the tifa.cpp CLI, used to extract phoneme timing from a recording.

Three q4 packages, one per platform:

  • tifa-ggml-windows-x64-q4.oudep
  • tifa-ggml-linux-x64-q4.oudep
  • tifa-ggml-macos-arm64-q4.oudep

Built from the v0.1.5 CLI bundles with Misc/tifa-oudep/package.py (OpenUtau repo). Install by opening the .oudep file in OpenUtau (it extracts to Dependencies/tifa-ggml).

Each package carries the Q4_0 aligner, the syllable dictionaries and both AP detectors: models/breath-fbl-q4_0.gguf (FoxBreatheLabeler, the default) and models/breath-v5-24k-f16.gguf (BreathLab, the alternative).

v0.1.1

Choose a tag to compare

@github-actions github-actions released this 01 Oct 12:32

tifa.cpp v0.1.1

相对 v0.1.0 的修正与新增。

这一版改了什么

  • 呼吸模型的归属修正:这个阶段之前被标成 FBL,实际上用的是 BreathLab
    模型(作者 Xiantaidu)。全仓的文档、界面、
    命令行提示都已更正并注明出处。
  • 诊断 JSON 默认导出:<名称>.diagnosis.json 里是 agreement /
    confidence / determinacy / monotonicity。按它排序,只需要人工校对最差
    的 10%
    。不想要就传 --output-formats textgrid(界面里取消勾选)。
  • 图形工具支持三平台:Linux、macOS 也出 GUI 包了,不再只有 Windows。
  • 每个平台一个自包含压缩包:GUI + 引擎 + 运行库 + 模型 + 词典 + 说明全在
    一个包里,下载一个就能用,不需要自己拼装。

下载哪个

包 说明
tifa-label-windows-x64.zip Windows 图形标注工具,解压即用
tifa-label-linux-x64.tar.gz Linux 图形标注工具(需 libnss3 libatk-bridge2.0-0 libgtk-3-0 libgbm1 libasound2)
tifa-label-macos-arm64.tar.gz macOS 图形标注工具(未签名,首次右键 →「打开」)
tifa-cli-<平台>-full.tar.gz 命令行版,F16 全精度模型
tifa-cli-<平台>-q4.tar.gz 命令行版,Q4_0 最小模型

<平台> = windows-x64 / linux-x64 / macos-arm64。包内已含模型、词典与
运行库。命令行务必带 -l(如 -l zh)——模型是多语言的、音素表按语言限定。

breathLab_models_dml.zip 是呼吸模型的上游源文件,供自行转换
(scripts/convert_breath_to_gguf.py)。

CUDA 包由独立的 cuda workflow 编译,完成后会自动补挂到本 release 上
(tifa-cuda-<平台>-full/-q4.tar.gz)。


English

Fixes and additions since v0.1.0.

  • Breath model attribution corrected: the stage was labelled FBL but the
    model is BreathLab, by Xiantaidu. Docs,
    UI and CLI text all say so now.
  • The diagnosis JSON is exported by default (<name>.diagnosis.json with
    agreement / confidence / determinacy / monotonicity) — sort by it and
    proof-read only the worst ~10%. Opt out with --output-formats textgrid.
  • GUI builds for Linux and macOS, not just Windows.
  • One self-contained archive per platform: app + engine + runtime + models
    • dictionaries + guide. Nothing to assemble by hand.

tifa-label-<platform>.{zip,tar.gz} is the GUI; tifa-cli-<platform>-{full,q4}.tar.gz
is the CLI (F16 and Q4_0 weights). <platform> is windows-x64, linux-x64 or
macos-arm64. Pass -l (e.g. -l zh) on the command line — the model is
multilingual and its symbols are language-qualified. Linux GUI needs
libnss3 libatk-bridge2.0-0 libgtk-3-0 libgbm1 libasound2; the macOS build is
unsigned (right-click > Open).

CUDA packages are built by the separate cuda workflow and are attached to this
release once that build finishes.

v0.1.0

Choose a tag to compare

@github-actions github-actions released this 01 Oct 10:40

tifa.cpp v0.1.0

openvpi/TIFA 的 C++ 原生(ggml)推理实现——用于歌声数据集制作的多语言强制对齐器。运行时不需要 Python。

这个版本有什么

  • 原生推理:CPU / Vulkan / Metal / CUDA;边界与参考实现一致(浮点差 ≤ 1e-3)
  • TIFA Label:Windows 图形标注工具;模型与引擎按默认相对路径自动导入,无需配置
  • 完整数据集流程:TIFA 对齐 → FBL 呼吸检测(AP/SP)合并 → 2PASS 重对齐
  • 多语言 G2P:中文 / 粤语 / 日语 / 英文(含 LSTM OOV 推理),支持 PFML 输入
  • 两种精度:F16 全量与 Q4_0 最小,每个平台各一份
  • 对齐速度(RTX 2070 / Vulkan / F16):约 0.15 s 每文件,≈ 42× 实时

下载哪个

包 说明
tifa-label-windows-x64.zip Windows 图形标注工具,解压即用
tifa-cli-<平台>-full.tar.gz 命令行版,F16 全精度模型
tifa-cli-<平台>-q4.tar.gz 命令行版,Q4_0 最小模型

<平台> = windows-x64 / linux-x64 / macos-arm64。包内已含模型、词典与 MSVC 运行时,开箱即用;用法见包内 USAGE.md(中英双语)。注意命令行必须带 -l(如 -l zh),模型是多语言的、音素表按语言限定。

breathLab_models_dml.zip 是呼吸(FBL)模型的上游源文件,供自行转换用(scripts/convert_breath_to_gguf.py)。

CUDA 包由独立的 cuda workflow 构建,编译完成后会补挂到本 release 上。


English

Native C++ (ggml) inference for openvpi/TIFA, the multilingual forced aligner used to build singing-voice datasets — no Python at runtime. Runs on CPU / Vulkan / Metal / CUDA, with boundaries matching the reference implementation (≤ 1e-3 float drift).

Pipeline: TIFA align → FBL breath AP/SP detection merged into the phones tier → 2PASS re-align.

Packages: tifa-label-windows-x64.zip is the Windows GUI annotator (unpack and run; it picks up the bundled models automatically). tifa-cli-<platform>-{full,q4}.tar.gz for windows-x64 / linux-x64 / macos-arm64 — models, G2P dictionaries and the MSVC runtime are included, so they work out of the box; see USAGE.md inside. Pass -l (e.g. -l zh) — the model is multilingual and its symbols are language-qualified.

breathLab_models_dml.zip is the upstream source for the breath (FBL) model, for converting yourself with scripts/convert_breath_to_gguf.py.

CUDA packages are built by the separate cuda workflow and are attached to this release once that build finishes.

edge build

edge build Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 30 Sep 09:37