Qwen3TTS distillation #367
ShuvalovDR
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Hi everyone / Hi authors,
First of all, thanks for the great work on this repository!
I am currently trying to perform knowledge distillation from the larger 1.7B model (teacher) into the smaller 0.6B model (student). However, I’ve run into an architectural bottleneck and couldn't find any existing papers or documentation that cover this specific setup.
The Problem & What I Tried
The main challenge lies in how the local transformer operates. As I understand it, the local transformer is conditioned on the 0-level backbone and the hidden layer. Because of this conditional structure, it’s not entirely clear how to align the teacher and student models.
Here is what I have experimented with so far:
Independent per-level losses: I tried applying KL divergence + Cross-Entropy (CE) at each level independently.
Freezing modules: I experimented with freezing the depth and backbone modules separately.
Result: The audio quality turned out to be very low. My hypothesis is that training the levels independently leads to inconsistent/uncoordinated acoustic codes.
Questions (Focusing on Losses & Alignment)
Since this approach isn't working well, I really want to discuss the loss functions and what we should actually be approximating. I would love to hear your thoughts on the following:
Loss functions: Is standard KL + CE per level the wrong approach here? Should we introduce intermediate supervision (e.g., MSE or Cosine Embedding Loss on hidden states)? If so, how would you balance the weight of these intermediate losses against the final code prediction losses?
What to approximate: If we do need to align hidden states, how should we handle the local transformer's conditioning on the 0-level backbone?
Dimensionality mismatch: If we align the hidden layers between the 1.7B and 0.6B models, do we need projection layers to map the student's hidden states up to the teacher's dimensions before calculating the distillation loss?
Code coherence: How do we enforce coordination between the levels to avoid the low audio quality issue I’m currently facing?
I would really appreciate any pointers, pseudo-code, or just a general discussion on how you would approach the losses in this specific architecture.
Thanks in advance!
标题: 讨论:从 1.7B 到 0.6B 的蒸馏 —— 对齐局部 Transformer 的条件输入与损失函数
正文:
大家好 / 作者们好,
首先,感谢你们在这个仓库中做出的出色工作!
目前我正尝试将较大的 1.7B 模型(教师模型)的知识蒸馏到较小的 0.6B 模型(学生模型)中。然而,我遇到了一个架构上的瓶颈,并且找不到任何涵盖这种特定设置的现有论文或文档。
问题与我的尝试
主要挑战在于 局部 Transformer 的工作方式。据我了解,局部 Transformer 是以 0级主干网络 和 隐藏层 为条件的。由于这种条件结构,目前尚不清楚如何对齐教师和学生模型。
以下是我迄今为止所做的实验:
独立的逐级损失: 我尝试在每个层级独立地应用 KL 散度 + 交叉熵 (CE)。
冻结模块: 我分别尝试了冻结 depth(深度)和 backbone(主干网络)模块。
结果: 音频质量非常低。我的假设是,独立训练各个层级导致了不一致/不协调的声学码。
问题(聚焦于损失与对齐)
由于这种方法效果不佳,我非常想讨论一下损失函数以及我们实际上应该近似什么。我很想听听大家对以下问题的看法:
损失函数: 在这里每个层级使用标准的 KL + CE 是错误的方法吗?我们是否应该引入中间监督(例如,在隐藏状态上使用 MSE 或余弦嵌入损失)?如果是这样,您将如何平衡这些中间损失与最终码预测损失的权重?
近似对象: 如果我们确实需要对齐隐藏状态,我们应该如何处理局部 Transformer 依赖于 0级主干网络的问题?
维度不匹配: 如果我们要对齐 1.7B 和 0.6B 模型之间的隐藏层,在计算蒸馏损失之前,我们是否需要投影层将学生模型的隐藏状态映射到教师模型的维度?
码的一致性: 我们如何强制各个层级之间进行协调,以避免我目前面临的音频质量低下的问题?
我非常感谢任何提示、伪代码,或者只是关于你们将如何在这种特定架构中处理损失的一般性讨论。
提前致谢!
All reactions