Skip to content

Releases: Pashtet495/AutoSpeechWriter

v.1.0.2

Choose a tag to compare

@Pashtet495 Pashtet495 released this 23 Jun 06:06
9d73f4a

v.1.0.2

EN

FIXED:

  • Fixed the animation of the stop recording button (red pulsation around the button depending on the volume level of the audio input).
  • Fixed the automatic reduction of the system volume level of microphones in case of high signal from the microphone.

RU

Исправления:

  • Исправлена анимация кнопки остановки записи (вокруг кнопки пульсация красного цвета зависящая от уровня громкости на входе аудио).
  • Исправлено автоматическое уменьшение системного уровня громкости микрофонов в случае высокого сигнала с микрофона.

v.1.0.1

Choose a tag to compare

@Pashtet495 Pashtet495 released this 22 Jun 16:15
499eff9

v1.0.1

EN

New features:

Input device settings, using system audio as an audio source, and a built-in mixer:

  • A model input device settings window has been added to the input device (microphone) settings. Each device now has a volume control (100% by default) and a real-time volume indicator. Multiple devices can be selected.

  • Selecting more than one device activates the mixer functionality. The program combines the output from the selected devices into a single audio stream.

  • In addition to real input devices, system audio has been added. System audio is also mixed with other selected devices.

004

Interface changes:

  • The main application interface has been optimized. Some button labels have been removed, leaving only icons and symbols.
  • The ability to edit text after recognition has been added. Text is cleared from the recognition window only upon user request.
  • A visual progress indicator has been added during text recognition from a file.

Video support:

  • Added the ability to recognize speech in video files (this can take a long time depending on the video file format and size. FFmpeg is responsible for separating and converting the audio track).

Creating subtitles in .SRT format:

  • Added the ability to create subtitles for audio and video files. An "srt" checkbox has been added to activate this feature. This feature generates text compatible with the *.SRT format in the output window. In subtitle creation mode, the "Copy" button is replaced with a "Save" button. The subtitle file is saved in the *.SRT format. To save the subtitle file with the name of the desired media file, select the media file name you want to clone into the new subtitle file when choosing a name and save location in the File Explorer window. File Explorer will warn you that you will overwrite the selected file, but this will not happen. The program only takes the file path and name to save the subtitle file so it can be recognized by video players as subtitles for that media file.
  • Subtitle timings are automatically split during pauses in the conversation. If there are no pauses longer than 15 seconds, the text is split by the current word.

Sound recognition mode without real-time output:

  • Added "record before recognition" mode. This mode does not use real-time text streaming. The recording is saved as a temporary file and processed after recording is complete. This feature allows the program to be used on devices that do not provide the required real-time performance. This mode provides the best text recognition quality. Dictation defects, repetitions, etc. are removed.

RU

Новые функции:

Настройка устройств ввода, использование системного звука как источника звука, встроенный микшер:

  • В настройках устройства ввода (микрофон) добавлено модельное окно настройки устройства ввода. Для каждого устройства добавлена настройка уровня громкости (по умолчанию 100%) и индикатор уровня громкости в реальном времени. Можно выбрать несколько устройств.
  • В случае выбора более 1 устройства активируется функционал микшера. Программа собирает выход с выбранных устройств в один звуковой поток.
  • Кроме реальных устройств ввода добавлен системный звук. Системный звук тоже микшируется с другими выбранными устройствами.

Изменения интерфейса:

  • Оптимизирован основной интерфейс приложения. Убраны некоторые подписи на кнопках, оставлены только пиктограммы и значки.
  • Добавлена возможность редактирования текста после распознавания. Текст стирается из окна распознавания только по требованию пользователя.
  • Во время распознавания текста из файла появился визуальный индикатор работы.

Поддержка работы с видео:

  • Добавлена возможность распознавать речь в видеофайлах (это может занимать много времени в зависимости от формата и размера видеофайла. За отделение и конвертацию аудидорожки отвечает FFmpeg).

Создание субтитров в формате .SRT:

  • Добавлена функция создания субтитров для аудио и видео файлов. Добавлен чекбокс "srt", активирующий эту функцию. С данной функцией в окне вывода генерируется текст, совместимый с *.SRT форматом. В режиме создания субтитров кнопка "копировать" заменяется кнопкой "сохранить". Файл субтитров сохраняется в формате *.SRT. Чтобы сохранить файл субтитров с именем нужного медиафайла в процессе выбора названия и места для сохранения в окне проводника выберите медиафайл название которое должно быть клонировано в новом файле субтитров. Проводник будет предупреждать, что вы перезапишете выбранный файл, но этого не произойдёт. Программа только забирает путь к файлу и его название, чтобы сохранить файл субтитров так чтобы он распознался видео проигрывателями как субтитры к этому медифайлу.
  • Тайминги субтитров автоматически разделяются во время пауз в разговоре. В случае отсутствия пауз более 15 секунд текст разделяется по текущему слову.

Режим распознавания звука без вывода в реальном времени:

  • Добавлен режим "запись перед распознаванием". В этом режиме не используется функционал стриминга текста в реальном времени. Запись сохраняется как временный файл и обрабатывается после окончании записи. Функция позволяет использовать программу на устройствах не обеспечивающих необходимую производительность в реальном времени. В этом режиме наилучшее качество распознавания текста. Убираются дефекты надиктовывая, повторения и т.д.

First version

Choose a tag to compare

@Pashtet495 Pashtet495 released this 20 Jun 14:51
60851ed

EN

  • Supported operating systems: Windows 10/11.

  • Supported recognition languages: Bulgarian (bg), Croatian (hr), Czech (cs), Danish (da), Dutch (nl), English (en), Estonian (et), Finnish (fi), French (fr), German (de), Greek (el), Hungarian (hu), Italian (it), Latvian (lv), Lithuanian (lt), Maltese (mt), Polish (pl), Portuguese (pt), Romanian (ro), Slovak (sk), Slovenian (sl), Spanish (es), Swedish (sv), Russian (ru), Ukrainian (uk).

  • Supported interface languages: En, Ru, De, Fr, Es, It, Ch.

Minimum system requirements:

  • SSD/HDD: 1.5 GB

  • Without a graphics card or with a graphics card that doesn't support the Vulkan API:

CPU 4 cores 3.5 GHz
RAM: 6 GB

  • With an integrated graphics card or with a discrete graphics card that supports the Vulkan API:

CPU 2 cores 2 GHz
RAM: 6 GB
VRAM: 1 GB

When working with an HDD, there may be a 2-3 second delay between pressing the record button and recording starting.

Main screen:
001

Settings screen
002

  • Recommended "BACKEND" setting: GPU Vulkan
  • If the application runs in parallel with neural networks or other workloads that consume all video memory, it is recommended to select a value (0, 1, etc.) in the "GPU ID" field that uses integrated graphics (if integrated graphics are available). This will ensure that the voice recognition model is loaded into RAM rather than VRAM (Intel UHD Graphics 770 provides approximately 15x real-time performance for the duration of the audio being loaded).

Context menu in the tray.
003
When you click on the cross in the main interface, the program is minimized to the tray.

The program has two recognition modes:

  • with minimal delay
  • with better correction.
    In correction mode, the text is accumulated and processed by the neural network until the sentence is completed. The model removes repetitions, hesitations, etc.

RU

  • Рекомендуемая настройка Бэкенда - GPU Vuklan.
  • Если распознавание запускается параллельно с другими задачами использующими VRAM полностью - рекомендуется запускать на встроенной графике, в таком случае модель загружается в оперативную память (выбор устройства производится в настройках в поле ID GPU).
  • Программа при нажатии на крестик сворачивается в трей.
  • В программе два режима распознавания с микрофона:
    С минимальной задержкой
    Лучшее распознавание.
    В режиме лучшего распознавания модель накапливает фразу и исправляет её, убирая запинки, повторы, слова паразиты и т.д.