如何调整生成的语速 #77

MisakaMikoto-o · 2023-12-18T08:19:31Z

我在prompt里用语速普通、语速很快和语速很慢发现区别不大，在代码里其他地方也没发现与速度相关的地方，请问有没有办法能调整生成的语音的语速

Ccj0221 · 2023-12-19T04:14:36Z

我想想知道语速如何调节，默认生成出来的音频需要加速到225%左右才是正常的，但是这样就带电子音了

MisakaMikoto-o · 2023-12-19T09:36:32Z

我现在用的python的第三方库pyrubberband来调的速，这样基本不会改变音调

syq163 · 2023-12-20T01:55:04Z

我现在用的python的第三方库pyrubberband来调的速，这样基本不会改变音调

Thank you for providing the information. Could you please assist in enhancing the related functions and contribute to EmotiVoice?

MisakaMikoto-o · 2023-12-22T01:44:14Z

Thank you for providing the information. Could you please assist in enhancing the related functions and contribute to EmotiVoice?

sure, it's my pleasure, I'll provide my code here, you can refer to it.

# 读取本地文件变速实现
import pyrubberband as pyrb
import soundfile as sf
import simpleaudio as sa
import numpy as np

def play_audio_at_speed(file_path, speed_factor=1.0):
    # 读取音频
    data, samplerate = sf.read(file_path)

    # 使用 pyrubberband 改变速度而不改变音调
    new_data = pyrb.time_stretch(data, samplerate, speed_factor)

    # 将音频数据转换为 simpleaudio 可接受的格式
    # 例如，转换为 16 位整数 (int16)
    if new_data.dtype != np.int16:
        new_data = (new_data * 32768).astype(np.int16)

    # 播放处理后的音频
    play_obj = sa.play_buffer(new_data, 1, 2, samplerate)
    #play_obj.wait_done()

# 调用函数以 1.5 倍速播放音频，不改变音调
play_audio_at_speed('/mnt/hgfs/reflection/EmotiVoice/outputs/prompt_tts_open_source_joint/test_audio/audio/g_00140000/16000.wav', speed_factor=1.5)

MisakaMikoto-o · 2023-12-22T02:07:06Z

Actual use can be written like this

from demo_page import get_models, tts
from frontend import g2p_cn_en, ROOT_DIR, read_lexicon, G2p
from config.joint.config import Config
import pyrubberband as pyrb
import simpleaudio as sa
import numpy as np

config = Config()
speakers = config.speakers
models = get_models()
lexicon = read_lexicon(f"{ROOT_DIR}/lexicon/librispeech-lexicon.txt")
g2p = G2p()

content = "hello"
text =  g2p_cn_en(content, g2p, lexicon)
path = tts(text, "开心", content, "8051", models)
# 转为0.75倍速
new_data = pyrb.time_stretch(path, 16_000, 0.75)
if new_data.dtype != np.int16:
    new_data = (new_data * 32768).astype(np.int16)
play_obj = sa.play_buffer(new_data, 1, 2, 16_000)
#play_obj.wait_done()

syq163 mentioned this issue Dec 20, 2023

How to control the speech speed in openAPI #67

Open

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

如何调整生成的语速 #77

如何调整生成的语速 #77

MisakaMikoto-o commented Dec 18, 2023

Ccj0221 commented Dec 19, 2023

MisakaMikoto-o commented Dec 19, 2023

syq163 commented Dec 20, 2023

MisakaMikoto-o commented Dec 22, 2023 •

edited

MisakaMikoto-o commented Dec 22, 2023 •

edited

如何调整生成的语速 #77

如何调整生成的语速 #77

Comments

MisakaMikoto-o commented Dec 18, 2023

Ccj0221 commented Dec 19, 2023

MisakaMikoto-o commented Dec 19, 2023

syq163 commented Dec 20, 2023

MisakaMikoto-o commented Dec 22, 2023 • edited

MisakaMikoto-o commented Dec 22, 2023 • edited

MisakaMikoto-o commented Dec 22, 2023 •

edited

MisakaMikoto-o commented Dec 22, 2023 •

edited