The ULTIMATE Local AI Update Just Dropped! ACE-Step 1.5, Paints-Undo & Whisper Premium Will WOW You #383
FurkanGozukara
announced in
Tutorials
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
The ULTIMATE Local AI Update Just Dropped! ACE-Step 1.5, Paints-Undo & Whisper Premium Will WOW You
Full tutorial: https://www.youtube.com/watch?v=hzKSt5WUAm0
In this tutorial I show the newer ACE-Step XL 1.5 Premium features, especially the corrected remix workflow that was missing from the previous video. You will see how to remix songs properly, convert lyrics and language, tune remix strength and melody retention, regenerate only selected parts, use the Library metadata system, install the upgraded Paints-Undo pipeline, reduce VRAM usage, and fix repeating lines in Whisper Premium transcriptions.
Important links:
Download ACESTEP XL Premium files: [ https://www.patreon.com/posts/ace-step-1-5-xl-premium-157675060 ]
Discord: https://discord.com/invite/software-engineering-courses-secourses-772774097734074388
Google AI Studio: https://aistudio.google.com/
Previous ACE-Step tutorial: [ https://youtu.be/9C_6qNKjgpA ]
Windows requirements/setup tutorial: [ https://youtu.be/DrhUHnYfwC0 ]
Whisper Premium tutorial: [ https://youtu.be/4lAk6sf1qF8 ]
Download Whisper Premium App Files: [ https://www.patreon.com/posts/whisper-webui-premium-145395299 ]
Video chapters:
00:00:00 ACE-Step XL 1.5 Premium update, remix focus, and Paints-Undo preview
00:00:49 Billie Jean remix example: checking style preservation and vocal change
00:01:16 Lower remix strength idea and Gangnam Style Korean-to-English demo
00:01:36 Hearing the Korean-to-English result and why extreme remixes sound strange
00:01:53 Stronger remix settings with higher strength and melody retention values
00:02:12 How to update ACE-Step XL 1.5: download ZIP, extract, and overwrite
00:02:37 Run Windows install/update.bat and rebuild the virtual environment if needed
00:02:56 Library tab overview: daily categories, saved generations, and metadata
00:03:18 Loading old songs from Library with lyrics, parameters, and JSON restored
00:03:33 Starting a proper remix: select the SFT model and open Advanced -) Remix
00:03:45 Uploading the source song and learning the two key remix parameters
00:04:04 Remix presets explained: different lyrics, same lyrics, medium and big change
00:04:19 Why lower melody retention lets the model generate completely new lyrics
00:04:30 Torch compile speed tip and preparing the target style caption and lyrics
00:04:42 Using Gemini in Google AI Studio with the ACE-Step lyric instruction file
00:04:59 Editing or writing lyrics and matching the vocal language accurately
00:05:15 Launching a live remix generation and measuring local generation speed
00:05:27 Live timing result: around 33 seconds for a complete remix generation
00:05:49 Use generated result as source to repair or improve selected song parts
00:06:03 Selecting remix start/stop points and regenerating only the chosen section
00:06:17 How section patching works: full remix generated, only selected part replaced
00:06:36 When to keep lyrics/style the same and when to change them for a section
00:06:47 Comparing the full output against the newly generated section preview
00:06:58 Iterative remix workflow for perfecting each part of the composition
00:07:14 LoRA training progress, future voice accuracy, and language-swap limits
00:07:27 Fast local iterations, multiple attempts, no watermark, and usable outputs
00:07:55 Final remix reminders: SFT model, proper Windows setup, and compile mode
00:08:11 Tuning dramatic changes with percentage values and fixed seed comparisons
00:08:41 Paints-Undo upgraded intro: new pipeline, faster speed, and better results
00:08:52 Download, install, and start Paints-Undo with windows_startup.bat
00:09:08 First launch model downloads plus new xFormers Triton attention support
00:09:21 GPU compatibility, Torch 2.12.1, CUDA 13, and upgrades over the original
00:09:34 Upload an image, generate the prompt, and tag it with the WD14 tagger
00:09:52 Operating steps, keyframes, Tiled VAE options, and low VRAM preparation
00:10:04 24GB vs 7GB VRAM usage and how the new memory-saving options help
00:10:27 How keyframes become the drawing video before final video generation
00:10:52 CUDA 13, Torch 2.12, Triton attention, diagnostics, and speed improvements
00:11:10 Automatic attention fallback plus Linux/cloud installer compatibility notes
00:11:27 Full generation time, possible torch compile addition, and result preview
00:11:46 Why keyframes matter, why results vary, and why not every image works
00:11:57 Fixing out-of-VRAM errors by lowering resolution and using supported ratios
00:12:14 Whisper Premium update: transcribing videos with all Whisper model options
00:12:31 Faster Whisper quality mode and the repeated sentence problem
00:12:42 Using large-v1 and repetition penalty to prevent repeated subtitle lines
00:12:54 How to tune repetition penalty carefully so transcription is not skipped
00:13:06 Example result: highly accurate subtitles generated from the new video
00:13:19 29-minute video transcribed in 1.5 minutes, 20x real-time, and closing
This video is for users who want fast local AI music remixing, better generation iteration, image-to-drawing animation, and high quality subtitle transcription. Follow the timestamps to jump directly to ACE-Step remix settings, Paints-Undo installation, low VRAM options, or Whisper repetition penalty tuning.
Video Transcription
00:00:00 Greetings everyone. Today I am going to show you newer features of the ACE-Step XL 1.5 premium
00:00:06 application and how to do remixes properly. It was lacking and inaccurate in our previous tutorial.
00:00:15 I assume that you have watched that. If you didn't yet, please watch it. This application is amazing.
00:00:20 It has so many unique features that you would like. Today you will learn how to do remixes
00:00:25 properly, and also I will show another application called as Paints-Undo upgraded. This application
00:00:32 is very unique. You provide an image, and it reconstructs it like you have hand-drawn it. It is
00:00:38 not perfect, but as you do more iteration, you may get a really good video. So it is like starting
00:00:43 from the 0 to repainting it into the full perfect quality. Okay, so let me show you a few remixes.
00:00:49 The first 1 will be famous Michael Jackson song, Billie Jean. So let's listen a little bit of it.
00:01:00 So you see it is keeping the style. The vocal has changed because we are doing heavy remix
00:01:06 right now. However, it is working really, really good. You can also do lesser degree remixes.
00:01:16 So if you make lesser degree, it won't be too much changed. I will show all of them. Another
00:01:21 demonstration will be entirely changing Korean song into English song. The famous Gangnam Style.
00:01:36 It sounds weird because the song is weird, and we are changing from Korean to English,
00:01:41 but still successful. As we have LoRA training, hopefully, it will be even better. You can even
00:01:47 do more powerful remixes by increasing the remix strength and the remix melody retention values,
00:01:53 like this case. So this is even more powerful remix. These are the lyrics of the song itself.
00:02:12 Yeah, it is weird, but this is how it is done. So how you are going to update to this version? The
00:02:18 link of the application will be in the description of the video. You already know it. Download the
00:02:22 latest zip file, put it where your installation is. If you are making a fresh installation,
00:02:27 it is the same. Right-click, right-click, and extract all the files, and make sure
00:02:32 that you are overwriting all the files. This is the proper way of upgrade. For upgrading,
00:02:37 Windows install or update.bat file, it will fully upgrade. If you encounter any issues,
00:02:42 you can also delete the virtual environment folder and run the installer again. So this will remake
00:02:48 the virtual environment. Everything will be kept, only the virtual environment where the
00:02:51 libraries are installed will be remade. It will be really fast. This is the way of upgrading.
00:02:56 So before I show you how to make remixes from scratch, there is 1 another thing that I want to
00:03:01 show you which I missed in the previous tutorial. When you go to Library, you will see all of your
00:03:06 generations. You can category them by the day, as you are seeing right now. It is very useful,
00:03:12 and when you click on the song, it will update all the lyrics, all the parameters that were used,
00:03:18 the full metadata JSON. So this is very useful. You can play the songs from here. So the library
00:03:28 feature is really, really useful. I had forgotten that in the previous tutorial, but you know it
00:03:33 now. So how to make remixes? Make sure that you have selected SFT model. The remix works best,
00:03:40 maybe only works with SFT model, we can say that. Then go to Advanced tab, select the Remix tab,
00:03:45 upload your song. Your base song will be here. Let me show you. So this is the original song. The
00:03:53 most important 2 parameters are remix strength and remix melody retention. As you make these lower,
00:03:58 it will make more dramatic changes. And we already have 4 presets: different lyrics,
00:04:04 same lyrics medium change, and same lyrics big change. Since this is a Korean song and I want to
00:04:09 generate it in English, I am selecting different lyrics. So you see it is setting the remix melody
00:04:14 retention to 15 percentage. Therefore, it will be able to generate entirely new lyric. If you
00:04:19 make this like 25 percentage, you won't see lyrics changed. This is how it works. You need
00:04:24 to make a few tests to understand how it works. I am also using torch compile which speeds up
00:04:29 my generation significantly. Then you need to type your style caption and lyrics according to
00:04:34 whatever you want to turn the song into. So I used Gemini and this file to generate the target style
00:04:42 and the lyrics. You can use the Gemini for free on Google AI Studio. So type Google AI Studio to
00:04:47 Google, go to this link, aistudio.google.com, and then upload your files, and also upload
00:04:53 this file that comes with the zip file, which is ACE-Step lyric generation instruction for LLMs.
00:04:59 Then tell whatever you want to do. So this is the style, and there is this lyrics. This lyrics
00:05:05 is modified version of the English lyrics that were published. You can also type your lyrics
00:05:10 yourself. Then select your vocal language. This is super important. Whatever the language of the
00:05:15 song that it will be turned into or the original, you need to set it accurately from here. This made
00:05:21 huge difference. Then you are all set. Let's live generate 1. The generation is really fast. Let's
00:05:27 see how much time it will take. Okay, we are almost done. So 20 seconds, and in 25... okay,
00:05:34 in 30, 31, and 33. In 33 seconds, I got the remix generated. So I can generate as many as needed
00:05:43 remixes and pick the best 1 that I love. Okay, let's say you didn't like a certain part of the
00:05:49 song. Now we have a new feature: Use generated result as a source. Click this button. It will
00:05:56 update the input audio and also this preview audio. Then from here, select the remix start
00:06:03 and stop wherever the part you want to regenerate. Once you select it, hit Generate Music again. It
00:06:09 will regenerate the remix and only change the part you selected. This way, you can quickly iterate
00:06:17 different sections of the song and get the perfect output as you want. So how this works is that it
00:06:23 regenerates entire remix but only changes the selected part from newer remix into your previous
00:06:30 song. So don't change your lyrics or style unless you want to change your lyrics, of course. If you
00:06:36 want to change your lyrics, change entirely, keep entirely, and that section will be regenerated and
00:06:42 copy-pasted. So you see this is the entire song, and if you want to only listen the new generated,
00:06:47 it is also here. It is shown here. Let me show you. So this way you can iteratively regenerate
00:06:57 and regenerate and get the perfect composition of each part you want. This is super useful. This
00:07:03 is really, really good feature, and hopefully, with the LoRAs, we will get much better quality,
00:07:09 much more accurate audio, much more accurate voice. It will be like the original singer.
00:07:14 The LoRA training is hopefully coming soon. I am working on it, but we have pretty much perfected
00:07:20 the inference. Now everything is ready. Of course, changing the entire singing from Korean to English
00:07:27 is hard. It is far from perfect. You need to work on it, make multiple generations. But the
00:07:32 generation speed is incredible, like 30 seconds on my local Windows computer, so that you can
00:07:38 iteratively improve song until you get a perfect song. And 1 advantage of this application is
00:07:43 that it won't have any watermarks, so it will be fully original to you. You can use it anywhere you
00:07:50 want. I will say 1 more time, make sure that you have selected the SFT model. This is for remix.
00:07:55 If you have proper setup of your computer with following the Windows tutorials, you can enable
00:08:00 compile mode. This will really speed up subsequent generations. The first 1 will be slow since it
00:08:06 will compile, but then it will become very fast. And if you want more dramatic changes, reduce
00:08:11 these values. As you reduce it, it will do more dramatic changes. Like 10 percentage will change
00:08:17 it entirely, like 15 percentage will change it a little bit lesser degree. So you see we have also
00:08:23 same lyrics medium change, like 25 percentage. Play with the percentages and see the impact it
00:08:29 is making. If you want to see direct impact on the same song, you can also disable random seed
00:08:34 and set a certain seed to keep trying on the same seed value to get similar results. Okay, the next
00:08:41 application I'm going to show is Paints-Undo upgraded. This is an older application,
00:08:46 but today I have updated it entirely into new pipeline. So it is much faster, much better now.
00:08:52 Download the latest zip file, the link will be in the description of the video as usual. Put it into
00:08:56 wherever you want to install. Everything is same in my applications actually. It is using Python
00:09:02 3.11. Just install it. So the installation has been completed already very quickly. For starting,
00:09:08 use the windows_startup.bat file. It will download the necessary models automatically for you. It
00:09:14 is using special new xFormers Triton attention, so it is really fast. It supports all the GPUs,
00:09:21 so this application that I made is much better than the official Paints-Undo application, many
00:09:27 times better. And it works with Torch 2.12.1 with CUDA 13. Upload your image. I have a test image in
00:09:34 the zip file. Generate prompt. It will quickly tag the image. So you see it is first time downloading
00:09:40 the model automatically. The tagging is like almost instant with WD14 tagger. Okay, it is done.
00:09:46 So this tag will be used for operating steps. You can change these steps. It is like generating
00:09:52 keyframes. And if you are on low VRAM, if you get out of VRAM errors, enable Tiled VAE keyframes and
00:09:58 Tiled VAE video. I have entirely coded these 2 options. And then there is also Low VRAM mode.
00:10:04 If you have a 24 gigabytes GPU, you shouldn't need to enable any of these, but you can also
00:10:09 enable. They don't cause massive speed degrades. It is almost same speed but much lower VRAM usage.
00:10:15 When I enable all 3, it is as low as 7 gigabytes of VRAM. And the base repository was like 24 plus
00:10:22 gigabytes of VRAM. So you will see your keyframes. You can also upload your keyframes if you want.
00:10:27 What these keyframes do is that then it will turn that into a video with Generate Video. It is like
00:10:34 you are drawing it slowly. This was developed by the legendary developer of the ControlNet,
00:10:40 lllyasviel. You know him maybe if you are in the generative AI projects. You can read more
00:10:46 description here. Our application is many times better than the original version with supporting,
00:10:52 as I said, newest Triton attention, CUDA 13, Torch 2.12. I needed to do a lot of programming
00:10:59 to upgrade all of them. And the speed is amazing, as you are seeing, compared to before. It is also
00:11:05 showing peak VRAM and other stuff. So you see it is also showing whichever the attention is being
00:11:10 used. It has automatic trying of the attentions. If 1 fails, it uses the next 1. But this
00:11:16 attention is really, really fast. And also the installers are working on RunPod, Massed Compute,
00:11:21 SimplePod as well. So if you are a Linux user, you can use the Massed Compute installer easily. Okay,
00:11:27 we are almost done. So the entire thing is taking like 100 seconds. I may even add torch compile if
00:11:34 anyone requests. So it is like, you see, drawing from scratch to it. It is like sketches to the
00:11:40 final video like this. Okay, we are watching it right now. So the keyframes are very crucial
00:11:46 here because it is trying to match them. It is far from perfect, and not every image will work. Also,
00:11:52 if you get out of VRAM error, you can reduce the resolution. This makes really significant impact
00:11:57 on the speed and VRAM. It is like this. You cannot set it any ratio as you want. It is like this. You
00:12:03 need to use this to set your resolution. And generate multiple keyframes, multiple generate
00:12:09 video to get the perfect output that you want, like you are drawing it. Oh, 1 more thing that
00:12:14 I want to show. You know that we have Whisper premium application which supports all Whisper
00:12:19 models with extra features to transcribe videos. If you are using this with the latest version,
00:12:26 you see we have Whisper, Faster Whisper. This is the really good quality. And if you get repeating
00:12:31 output, sometimes because you are getting repeating output, like repeating sentences, I
00:12:36 recommend you to select model large version 1. And the most crucial thing that you are going to make
00:12:41 additionally to default settings is repetition penalty. Set this like 1.2, and it should prevent
00:12:47 your repeating sentences. You can try 1.1, 1.2, 1.3. As you increase it, it may skip some of the
00:12:54 transcriptions. So making this very high is not recommended. So slowly increase this until you do
00:13:00 not get any repeating sentences. This way you can generate amazing accurate transcriptions,
00:13:06 amazing accurate subtitles. For example, the new video that I have transcribed is looking perfect.
00:13:11 I don't notice any errors at all. And the speed was amazing. 29 minutes video was transcribed in
00:13:19 1.5 minutes. You can see that, 29 minutes video transcribed in 1.5 minutes. So the real-time
00:13:26 speed was 20x with the highest quality options. Thank you so much. Hopefully see you later.
All reactions