Skip to content

Ideogram 4 is HERE The Ultimate JSON Prompting Masterclass

FurkanGozukara edited this page Jul 19, 2026 · 1 revision

Ideogram 4 is HERE: The Ultimate JSON Prompting Masterclass!

Ideogram 4 is HERE: The Ultimate JSON Prompting Masterclass!

image Hits Patreon BuyMeACoffee Furkan Gözükara Medium Codio Furkan Gözükara Medium

YouTube Channel Furkan Gözükara LinkedIn Udemy Twitter Follow Furkan Gözükara

Learn how to run Ideogram 4 locally with SwarmUI and ComfyUI, download the required model bundle, use ready Turbo/Balanced/Highest Quality presets, and create accurate structured JSON prompts with Ultimate Image Captioner Pro for image recreation, text rendering, batch captioning, and training dataset preparation.

Tutorial Links:

🖼️ Ultimate Image Captioner Pro:

https://www.patreon.com/SECourses/posts/ultimate-image-captioner-pro-162527725

🐝 SwarmUI installer, model downloader and presets:

https://www.patreon.com/SECourses/posts/swarm-ui-installer-model-downloader-114517862

🧩 ComfyUI installer:

https://www.patreon.com/SECourses/posts/comfyui-installer-105023709

🎬 Windows requirements tutorial:

https://youtu.be/DrhUHnYfwC0

⚙️ Requirements post with links/screenshots:

https://www.patreon.com/SECourses/posts/requirements-tutorial-step-by-step-written-111553210

💬 Discord support:

https://discord.com/invite/software-engineering-courses-secourses-772774097734074388

⭐ SECourses GitHub:

https://github.com/FurkanGozukara/Stable-Diffusion

Chapters:

00:00:00 Ideogram 4 overview: JSON prompting, SwarmUI presets, ComfyUI workflows, and model bundle

00:00:53 Ultimate Image Captioner Pro for turning reference images into Ideogram JSON prompts

00:01:10 Editing JSON elements, bounding boxes, wanted text fields, captions, and prompt layout

00:02:02 Regeneration examples showing structure, objects, scene layout, and image text matching

00:03:18 Captioner Pro feature tour: Qwen, JoyCaption, saved outputs, and JSON builder

00:04:30 Dataset workflow: prompt presets, batch folder captioning, and automatic VRAM presets

00:05:13 Tutorial roadmap: ComfyUI update, SwarmUI update, model download, app install, usage

00:05:40 Updating ComfyUI by extracting the latest installer zip and overwriting old files

00:05:56 Optional fresh ComfyUI venv rebuild for fixing outdated or broken installations

00:06:15 Running the ComfyUI update script, Python choice, UV speed, and quant support

00:07:05 Installing recommended custom nodes bundle 100 for ComfyUI and SwarmUI compatibility

00:07:50 Launching fresh ComfyUI and testing the Ideogram Turbo preset workflow

00:08:47 Setting width, height, resolution, and matching prompt aspect ratio

00:09:07 Updating SwarmUI with the latest zip, overwrite method, and safe folder paths

00:09:48 Automatic .NET SDK 10 install and why SwarmUI needs the correct SDK version

00:10:51 SwarmUI backend setup: ComfyUI backend, Triton, Sage Attention cautions, extra args

00:11:44 Downloading the Ideogram 4 core bundle with hash verification

00:12:28 16-connection parallel downloads, target folders, ComfyUI mode, and URL downloader

00:13:20 Merging model parts and sharing SwarmUI models through extra_model_paths.yaml

00:13:51 Setting SwarmUI model root to reuse another model folder and avoid duplicates

00:14:12 Updating SwarmUI presets with delete import, normal import, overwrite, and backup

00:14:58 Refreshing presets and confirming Ideogram Turbo, Balanced, and Highest Quality

00:15:14 First simple Ideogram prompt, false safety filter block, and weak plain prompting

00:15:34 Using Realism Engine Ideogram 5 LoRA to fix the blocked car prompt

00:15:57 Why detailed JSON prompts are needed and downloading Captioner Pro

00:16:23 Installing Captioner Pro with Windows install update app, venv, and model downloads

00:16:34 Windows requirements: Python, CUDA, cuDNN, C++ tools, FFmpeg, Git, and setup guide

00:17:03 Cloud/Linux notes plus Massed Compute interface, creator image, GPU, and coupon

00:17:34 Captioner installer downloader: 16 connections, hash checks, and accurate setup

00:17:57 Starting Ultimate Image Captioner Pro and saving custom user presets

00:18:14 Loading the Bugatti reference image and generating official Ideogram JSON

00:18:39 Prompt generation speed, copying the prompt, and understanding VRAM usage

00:19:09 Subprocess mode to release all VRAM and RAM after each captioning run

00:19:54 Reviewing generated JSON: high level description, visible text, boxes, and details

00:20:21 Pasting JSON into SwarmUI and matching the custom 5:3 aspect ratio

00:20:43 Aspect ratio calculator, side length control, and high resolution generation

00:21:36 Comparing with and without aspect ratio metadata and avoiding false safety blocks

00:21:58 Realism Engine LoRA strength, when to use it, and output comparison

00:22:34 Choosing Turbo, Balanced, or Highest Quality and testing Turbo speed

00:22:54 Ideogram 4 image to image, inpainting, image creativity, and image prompts

00:23:19 Captioner Pro batch folder processing: subfolders, overwrite, and append modes

00:23:35 Post processing captions with prefixes, suffixes, replacements, and sensitivity

00:24:07 Final options, auto quantization by GPU VRAM, support channels, and closing

Covered in this video: local Ideogram 4 installation, SwarmUI and ComfyUI preset usage, automatic model downloads, JSON prompt creation, bounding box editing, image recreation, safety filter fixes, LoRA realism settings, batch captioning, and VRAM friendly caption generation.

Video Transcription

  • 00:00:00 Greetings everyone, today I am going to show you  everything about newest Ideogram 4 model. This  

  • 00:00:06 model is extremely powerful, it works and supports  with JSON prompts, and how to use it is not easy  

  • 00:00:14 as other models, but it is very powerful. I have  prepared 3 different presets, they are all ready:  

  • 00:00:21 balanced preset, highest quality preset, and  turbo preset. These are SwarmUI presets but  

  • 00:00:27 their same versions exist on ComfyUI as well as  workflows. Downloading the necessary models is  

  • 00:00:34 also ready with our model downloader Ideogram 4  core bundle, you can download them right away or  

  • 00:00:41 you can go to the image generation models and  see the additional models here as well. So it  

  • 00:00:47 works with JSON prompts but how to use this  model with JSON prompts easily? To generate  

  • 00:00:53 JSON prompts I have developed a new application  called as Ultimate Image Captioner Pro. This  

  • 00:00:59 application is a powerhouse, I will show all  features of it. You see this is the input image,  

  • 00:01:05 this is the JSON prompt it generated, these are  the JSON elements that you can modify, and this  

  • 00:01:10 is the visualization of the JSON prompt. Another  example, this is input image, this is JSON prompt,  

  • 00:01:15 you see these are the captions of the boxes,  when we scroll down we can see the boxes here,  

  • 00:01:21 you can move these boxes very easily like this.  If you want to work with the boxes more easier,  

  • 00:01:26 hide them, and drag and drop, make the changes  like this, and once you done with everything,  

  • 00:01:32 including changing the captions or everything,  apply box edits and it will save the edited  

  • 00:01:38 boxes. So you see as you click here and here  it will update all these values for you, you  

  • 00:01:44 can change them. One of the big advantage of this  model is that it has both caption and text area.  

  • 00:01:50 So you see it is taking the text differently.  You see this is wanted text and the wanted text  

  • 00:01:55 is written like this. Caption is the regular  caption, text is if there is a text. The model  

  • 00:02:01 works without any JSON prompt as well, it is not  mandatory but the power of the model comes from  

  • 00:02:07 JSON. So this is another example as you see, and  I have used these examples and regenerated images  

  • 00:02:13 like this one. So this and this image is perfectly  matching as a structure, the style is different,  

  • 00:02:19 obviously you need to change the style, but you  see there is a moon here and we can see that in  

  • 00:02:25 the prompt there is a crescent moon visible in the  dark sky. So it is perfectly matching. There are  

  • 00:02:30 2 mouses, we can see how it is captioned, it says  a white mouse wearing a beach coat and blue scarf  

  • 00:02:37 holding a snowball, smiling with snow on its head.  So I need to change the style if I want to change  

  • 00:02:42 the style, there is no style information here,  but the structure is fully matching. Mouse colors,  

  • 00:02:48 what they are wearing, what they are doing, the  overall scene. This is another example, this is  

  • 00:02:53 generated from this, you see it was like this and  this is the generated example. So the model is  

  • 00:02:59 perfectly able to regenerate the original input  image, as you wish. So this is the regeneration  

  • 00:03:05 of this original image. You see it was this one  and this is regeneration. So this model is very  

  • 00:03:12 powerful as I said, you can use it as you wish.  With this captioner application you won't have any  

  • 00:03:18 issues to regenerate any images. This application  has so many features, not only Qwen Instruct,  

  • 00:03:24 but we also support JoyCaption as well. So you  can use JoyCaption captioners and you will see  

  • 00:03:31 the generated captions and everything. Everything  is automatically saved inside outputs folder.  

  • 00:03:37 Moreover, we have JSON builder as well. So with  JSON prompt builder you can start from beginning  

  • 00:03:42 or you can load your existing generations and make  changes and save everything. This is amazing. The  

  • 00:03:49 JSON prompt builder is so easy to use, when you  click generate JSON it will start empty like this,  

  • 00:03:54 so you can keep adding boxes, type captions,  whatever you want, and save them. So it will  

  • 00:04:00 be saved in a new fresh folder, but I prefer to  load the generated JSON and work on it. It is up  

  • 00:04:07 to you, you can do both ways. Finally you can also  use the saved outputs. So with saved outputs you  

  • 00:04:14 can refresh and when you click it will load very  easily all the generations. There is filtering,  

  • 00:04:20 daily filtering, number of displayed results per  page, it is very advanced, you can use this very  

  • 00:04:26 easily. So if you are going to make training,  this application is perfect for that. You see we  

  • 00:04:30 support all these prompting presets: Ideogram,  photorealistic, art style, character subject,  

  • 00:04:36 product object. So you can use these according to  your dataset and you can do batch captioning. In  

  • 00:04:42 the bottom we have batch folder captioning, it  will use these set parameters and batch caption  

  • 00:04:48 the given folder and save all the outputs.  Moreover, we support all these VRAMs. So  

  • 00:04:53 if you have a low VRAM GPU, don't worry we cover  you. These low VRAM presets exist for JoyCaption  

  • 00:05:00 as well. So whether you are using JoyCaption or  Qwen, you can select your VRAM preset. It will  

  • 00:05:06 automatically select it according to your GPU, but  you can go to the lower or higher VRAM presets as  

  • 00:05:13 well. How you are going to use this model? I will  begin with showing you how to update your ComfyUI,  

  • 00:05:20 then how to update your SwarmUI, then how to  download models, then how to install the image  

  • 00:05:26 captioner app, and the rest is as usual using.  All older tutorials are valid, there aren't many  

  • 00:05:33 new stuff, so let's begin. So first of all, check  below description and go to ComfyUI installer and  

  • 00:05:40 download the latest version. Move the downloaded  zip file into your ComfyUI installation folder.  

  • 00:05:46 Right click and extract and overwrite all  the files. This is super important. Once you  

  • 00:05:51 overwritten all files, you are ready to update. I  recommend you to delete your virtual environment  

  • 00:05:56 folder if you didn't update for a long time, so  that we will get all freshly installed virtual  

  • 00:06:02 environment. This is not mandatory step, but this  is perfect way of fixing your ComfyUI installation  

  • 00:06:08 and updating it. Okay, the folder is deleted.  Then run the Windows install or update ComfyUI.bat  

  • 00:06:15 file. Choose your Python version, currently I  am still using 3.10 but we may move to 3.11 as  

  • 00:06:22 a preferred Python version soon because a lot of  packages are moving to that. This will generate  

  • 00:06:27 the virtual environment or update the libraries  if it is needed and it will update your ComfyUI to  

  • 00:06:32 the latest version. Moreover, it will update the  some of the custom nodes that we use by default  

  • 00:06:38 with our workflows and presets. We are using UV  installer, therefore the updates or installations  

  • 00:06:45 are super fast. Moreover, I am automatically  installing quant operations, therefore our  

  • 00:06:51 ComfyUI automatically supports all of the quants  versions that you might find out there, like int8  

  • 00:06:58 block-based quantizations, different quantizations  there you might find. Okay, it is all done,  

  • 00:07:05 everything is set. One more thing that I recommend  you to do is use the Windows custom nodes bundles  

  • 00:07:12 installer, run, and my recommended bundle is  100. You can also choose them 1 by 1 from here  

  • 00:07:19 with comma separation, but I am going to choose  bundle 100 and hit yes. This will update the most  

  • 00:07:26 commonly used nodes to the latest versions. You  see these nodes will get installed, this is what I  

  • 00:07:33 recommend to use with ComfyUI and SwarmUI with the  maximum quality and performance. So it is updating  

  • 00:07:39 my nodes and everything should be ready. If you  use other different custom nodes, then these ones,  

  • 00:07:45 they may conflict with your installation, they  may break your installation, but these are the my  

  • 00:07:50 recommendation. So my ComfyUI is now ready. I am  using the extra model paths.yaml file, therefore  

  • 00:07:58 it is seeing all of the models downloaded into  my SwarmUI. So now I can start and use right  

  • 00:08:04 away. Actually let me show you, run GPU.bat  file, it will start the ComfyUI freshly set,  

  • 00:08:09 everything is freshly set right now. We are  still using PyTorch 2.9.1 but I plan to upgrade  

  • 00:08:16 to PyTorch 2.12 soon, once the TorchAudio is also  updated. Then inside the presets you will see the  

  • 00:08:24 Ideogram presets. You can use any of them. Turbo  is also working very well. When I drag and drop it  

  • 00:08:30 will be loaded like this. You see all the models  are automatically seen accurately. When I run it,  

  • 00:08:35 it should pretty fast generate output. The turbo  preset is really fast. And it is done. The turbo  

  • 00:08:41 preset already generated. One more thing that  I need to mention is that you need to set the  

  • 00:08:47 width and height from here. These width and  heights are not important, so set your width  

  • 00:08:53 and height from here to set your resolution.  Moreover, the prompt generator uses aspect ratio,  

  • 00:09:00 therefore try to match your aspect ratio with your  resolution to the prompt, or change both of them  

  • 00:09:07 accordingly. Now as a next step we will update our  SwarmUI. So go to the description below and go to  

  • 00:09:14 the SwarmUI link and download the SwarmUI model  downloader zip file. Move it wherever you are  

  • 00:09:20 going to install or your existing installation.  So I recommend you to not have any special  

  • 00:09:27 characters in your folder paths, including spaces  or non-English base characters. Make your folder  

  • 00:09:31 paths like this. Then right click and extract and  overwrite all the files. This is super important,  

  • 00:09:37 overwriting all the files. Once it is done, you  can just use install SwarmUI or update SwarmUI.  

  • 00:09:43 There is 1 more thing that I want to mention  before I update it. When you start the SwarmUI,  

  • 00:09:48 you may have noticed that it is telling you  this: please install .NET SDK 10. So the SwarmUI  

  • 00:09:55 is going to update to SDK 10 version. I updated  our installer and updater, when I update SwarmUI,  

  • 00:10:02 now it will automatically install the accurate  SDK version. It is also going to ask permission,  

  • 00:10:09 that is why I deleted it from my computer to  show you. Okay, now it is asking the permission,  

  • 00:10:14 I click yes, and it will open this screen and  install. It will first download the exe file,  

  • 00:10:19 then it will start the installation process like  this, that you need to click and continue. This  

  • 00:10:24 way you will get the accurate .NET SDK version  and you will have the latest version. So your  

  • 00:10:30 SwarmUI will keep working. This is necessary since  SwarmUI is actually programmed with C# rather than  

  • 00:10:38 Python. We are using ComfyUI as a backend if you  remember my previous tutorials. So the backend  

  • 00:10:43 is ComfyUI but SwarmUI is basically a wrapper that  lets you use the ComfyUI with much easiness. Okay,  

  • 00:10:51 it is done. Then it will continue updating, it  will update the necessary other stuff if there is  

  • 00:10:56 anything, it will compile and start the SwarmUI.  Once SwarmUI started, make sure that you are using  

  • 00:11:02 our ComfyUI backend installation, and now I am  using enable Triton backend. I don't recommend  

  • 00:11:08 you to add Sage attention by default, because in  some models, in newer models, it may not work very  

  • 00:11:14 well. So you need to test whether it is working  or not on each model that you are using. Enable  

  • 00:11:20 Triton backend is working amazing. Enable Triton  backend is also added to the ComfyUI starter,  

  • 00:11:26 when you edit the run.bat file you will see  that it is using the enable Triton backend.  

  • 00:11:32 So it also uses Sage attention, so you can remove  it from there as well if you need. This is how you  

  • 00:11:37 add extra arguments to your ComfyUI backend from  SwarmUI interface. So as a next step you need to  

  • 00:11:44 download necessary models. To download necessary  models we are going to use start download models  

  • 00:11:50 app.bat file. It will start the model downloader  application. And then in the SwarmUI bundles you  

  • 00:11:56 will see that Ideogram 4 core bundle. Download all  the models, it will download if they are missing,  

  • 00:12:02 if they are already downloaded it will just hash  verify them. SwarmUI may modify your downloaded  

  • 00:12:09 models and it will cause mismatch of the hash  files. In that case it will redownload. But this  

  • 00:12:16 downloader is made very well, it verifies hash  files, so with this downloader you will never have  

  • 00:12:22 corrupted model issues. Moreover, it starts 16  different parallel downloads, therefore it is able  

  • 00:12:29 to download with maximum speed that your internet  service provider supports. Currently you see it is  

  • 00:12:34 downloading with 100 megabytes per second, this  is my maximum speed, I have 1 gigabits internet.  

  • 00:12:40 And we can also see the downloaded model parts  here, so you see it is downloading as a 16 parts,  

  • 00:12:47 16 connections. All the models, this application  download, downloaded same way. You can set your  

  • 00:12:52 target model folder, it also supports ComfyUI  model structure, just enable this checkbox. It  

  • 00:12:58 also supports Forge WebUI Automatic1111 folder  structure or lower case folder names. Moreover,  

  • 00:13:04 it also has URL downloader, I had explained all  of this in previous tutorials, so you can download  

  • 00:13:09 from Civitai or Hugging Face into target folder.  And it will be very fast and hash verified. I  

  • 00:13:15 really recommend to use this model downloader.  Once the model downloaded, they will be merged  

  • 00:13:20 into single part. And I am not duplicating the  models, I am using the extra model paths.yaml, you  

  • 00:13:27 need to copy this and paste it into your ComfyUI  folder. It is coming with our zip file. And when  

  • 00:13:32 you edit this file you will see your base folder  path, you need to change this according to your  

  • 00:13:39 SwarmUI installation, therefore it will see all  the models that was downloaded into your SwarmUI,  

  • 00:13:45 so that you can use it inside ComfyUI as well.  Or in the SwarmUI, I think it also supports that,  

  • 00:13:51 so go to server configuration and you see there  is model root, so you can give another root folder  

  • 00:13:57 like your ComfyUI models, and when you save it in  the bottom, or auto save it, yes it is probably  

  • 00:14:02 auto saved as you change them, so it will see the  models from that another folder as well. You don't  

  • 00:14:07 need to duplicate any models. Then what you need  to do is, you need to update presets. For updating  

  • 00:14:13 presets I recommend you to use Windows preset  delete import. It will clear all of your presets  

  • 00:14:20 and update them to our latest versions. You can  alternatively also use import, so choose file,  

  • 00:14:26 go to the folder and pick the amazing SwarmUI  presets, currently version 51. It will say that  

  • 00:14:33 are you want to overwrite or not, then you can  overwrite and import everything. Alternatively,  

  • 00:14:38 you can click the Windows preset delete import,  you need to run this when the SwarmUI is running,  

  • 00:14:45 then it will ask you whether you are sure or  not, yes, and it will clear all of your presets  

  • 00:14:50 and import them like this. It also backups your  presets inside utilities folder as presets backup  

  • 00:14:57 before deleting them. So once the presets are  refreshed you will see them like this, so refresh  

  • 00:15:02 presets, and the Ideogram presets have arrived.  Now they are ready to use. Once the models are  

  • 00:15:08 downloaded, yes I see they are already downloaded  and verified. So for using as always, as usual  

  • 00:15:14 quick tools, reset params to default, then select  the preset that you want to download, like let's  

  • 00:15:19 select turbo direct apply. And you can type your  prompt. Amazing car going fast on a road. This  

  • 00:15:27 is a very simple prompt. And unfortunately it  is blocked. So we have a LoRA, Realism Engine  

  • 00:15:34 Ideogram 5, it is also downloaded with the core  bundle, let's select this and try again. Let's  

  • 00:15:39 see if it will fix this issue. Yes, you see the  previous image was blocked by the safety filter,  

  • 00:15:45 but when I enabled the Realism Engine Ideogram  version 5, it is working. However, it is not a  

  • 00:15:51 good quality because this model wants you to have  detailed JSON prompts. So I am going to do that,  

  • 00:15:57 how? Check below and you will see Ultimate Image  Captioner Pro link, open the page and download  

  • 00:16:04 the Ultimate Image Captioner Pro latest zip file.  I will show a fresh installation into my Q drive,  

  • 00:16:11 so I will paste it there, right click and I will  extract. Then enter inside the extracted folder  

  • 00:16:16 and all you need to do is just use the Windows  install update app.bat file. It will install,  

  • 00:16:23 it will install the application with a virtual  environment and download all the necessary  

  • 00:16:27 models. This application is using Python 3.11.  Always pay attention to the Windows requirements,  

  • 00:16:34 if you didn't watch the requirements tutorial  previously please watch it, its link is fully up  

  • 00:16:39 to date, so when you open the link of the tutorial  you will see everything with images as you are  

  • 00:16:45 seeing right now. So you won't have any issues  how to follow tutorial, how to use the latest  

  • 00:16:51 updated libraries. When you follow this tutorial  you will be fully ready to run any AI application  

  • 00:16:57 on your Windows computer. As always we also have  RunPod and Massed Compute instructions, so if you  

  • 00:17:03 are a Linux user you can use the Massed Compute  instructions. Everything is fully up to date,  

  • 00:17:09 even the tutorial videos are up to date in the  instructions read.txt file. There is only 1 thing  

  • 00:17:14 that I want to show you, Massed Compute updated  its interface, so if you use Massed Compute make  

  • 00:17:20 sure that category creator, image SECourses,  select your GPU, enter your coupon as SECourses  

  • 00:17:28 and verify. You see you will get amazing discount  in all of the GPUs that you can use on Massed  

  • 00:17:34 Compute. So the application is getting installed  and it is downloading the models, I already have  

  • 00:17:40 it in somewhere else. This is also using 16  connection download and hash verification. All  

  • 00:17:46 of my installers uses this specific special  downloader that I have developed, therefore  

  • 00:17:52 they will be always fully accurate. So the other  application was installed here, I will just run  

  • 00:17:57 it with Windows start Ultimate Image Captioner  Pro. Once your installation has been completed,  

  • 00:18:02 this is the interface. It supports custom user  presets as well, so you can make changes and save  

  • 00:18:08 them. So let's find an image and replicate it  in the Ideogram 4 model. For example, let's try  

  • 00:18:14 this image. I will pick the image, so for picking  image click here or you can drag and drop. Okay,  

  • 00:18:20 this is the image. Then I will caption image, I  am using the Ideogram official version 1 preset,  

  • 00:18:26 you can also use other presets as I have said.  It also supports text generation, not only JSON  

  • 00:18:32 based prompts but also text based prompts as well.  Moreover, our application is super optimized, it  

  • 00:18:38 will also even show you the token speed, let's see  the token speed, it is about 25 token per second.  

  • 00:18:45 It is up to 4000 tokens, but usually it ends much  faster depending on the image. So it is generated  

  • 00:18:51 in 11 seconds. Let's copy this prompt. By the way,  now it is keeping my VRAM busy, how? When I open  

  • 00:18:59 it I can see that if you don't want to keep your  VRAM busy, what you can do is, let me close this  

  • 00:19:05 and show you again. So other applications are also  keeping VRAM busy. Let's terminate them and let's  

  • 00:19:11 start the image captioner Pro again. This way  you can run both of the applications at the same  

  • 00:19:15 time. Okay, so our VRAM is like this right now.  This is SwarmUI started the ComfyUI backend. So  

  • 00:19:22 for Ultimate Image Captioner to not use any VRAM,  there is this option: run single and batch in sub  

  • 00:19:29 process. So what will this do? This will generate  caption and then terminate the process. Therefore,  

  • 00:19:35 it will leave absolutely 0 VRAM and RAM usage.  It will fully terminate process. This way you  

  • 00:19:41 can keep your captioner open, run it anytime you  want, and it will fully terminated later. It won't  

  • 00:19:48 have any VRAM usage, any VRAM leakage, we will see  in a moment, yes, it terminated and generated the  

  • 00:19:54 caption. So let's verify that it is accurate, yes,  we can see that a black, orange Bugatti Chiron,  

  • 00:20:00 and it even shows its text here. You see, amazing.  We can see also generated caption. So there is 1  

  • 00:20:08 high level description, it says a black and orange  Bugatti Chiron Super Sport 300 plus sports car.  

  • 00:20:16 This is the main structure of the image, then it  shows the text as well. So let's return back to  

  • 00:20:21 our SwarmUI. And remember last time it had given  us image blocked by safety filter, just paste  

  • 00:20:29 this. Let's see the aspect ratio to be sure, okay,  it is 5 to 3. So this is a pretty custom aspect  

  • 00:20:36 ratio, perhaps we can, yeah, there is not exactly  5 to 3. So how we gonna do that 5 to 3? I asked  

  • 00:20:43 the SwarmUI developer to add custom aspect ratio  here, I hope he adds, you can also tell him. There  

  • 00:20:50 is a website that I have found, aspect ratio  calculator, so let's enter here, let's set our  

  • 00:20:56 aspect ratio, and let's generate a high resolution  image. So it is around, yeah, 2045 to this one,  

  • 00:21:04 so I will enter custom, this will be this, and  this will be that. If it was a standard aspect  

  • 00:21:10 ratio it would be much easier, or alternatively we  can remove the aspect ratio from here, but it may  

  • 00:21:17 break our, yes, B boxes probably, let's try it. So  let's also select, then you can have side length,  

  • 00:21:24 so this is very useful, why? This way I can change  the resolution of the image while keeping the  

  • 00:21:30 aspect ratio, this is a new feature and it is  amazing. Okay, we got our image, pretty cool.  

  • 00:21:35 Let's try this one and see without aspect ratio  how it works. Okay, it is I think generating  

  • 00:21:41 almost same image. This image was, let's see the  generation, yeah, 2045 to 1227, this will be 2016  

  • 00:21:50 to 1152, yeah, and this is another image. So when  you do JSON prompting, the stupid safety filter  

  • 00:21:58 should not be triggered unnecessarily. And in the  LoRAs you can set your Realism Engine Ideogram 5,  

  • 00:22:04 this is automatically downloaded, if you want  to change its impact strength, you can change it  

  • 00:22:09 from here, you see, with my mouse, or you can just  type it, like let's try 1.1. And this is improving  

  • 00:22:16 the realism, I tried it, it is pretty useful. So  if your prompt is related to realism, use this,  

  • 00:22:22 if it is not, then don't use it. And it is ready.  So with realism I got this, without realism I got  

  • 00:22:28 this. It is up to you, you can try them. You can  change the presets, this was the turbo preset,  

  • 00:22:34 there is also balanced preset and highest  quality preset, you can use any of them,  

  • 00:22:38 but turbo is pretty fast and especially fast at  1024 pixel, let's see real time how fast it is  

  • 00:22:45 being generated. Okay, almost ready, yes. So it  took like, let's see, 7 or 8 seconds to generate.  

  • 00:22:54 The model supports image to image as well, so you  can use init image, set image creativity, and use  

  • 00:23:00 the prompt from this to this, so you can do image  to image as well, it is pretty consistent, it also  

  • 00:23:06 supports inpainting. It also even supports image  prompt from here, upload prompt image, however I  

  • 00:23:13 didn't test it and it is not very good like Qwen  2512, so you need to play with it, you can provide  

  • 00:23:19 image prompts as well. And about image captioner  app, for folder batch processing enter your  

  • 00:23:25 folder, input folder here and output here, you can  also process sub folders if you want, you can also  

  • 00:23:30 overwrite captions or append captions, it supports  all of them. And 1 another thing is that we  

  • 00:23:35 support text prefix and suffixes, so you can add  OHWX to all of your captions automatically, or you  

  • 00:23:42 can even have replace word, like man replace it  with OHWX and add, and it will add it as a replace  

  • 00:23:49 word. And it will post process all generated  captions and replace these words. You can make  

  • 00:23:54 it case sensitive or single word sensitive. When  you make it single word sensitive actually it will  

  • 00:24:00 only find man word and not the wildcard prefix.  So it is all up to you, we have all the options.  

  • 00:24:07 You can also disable auto save box image. So check  all the options we have, these are all set after  

  • 00:24:15 deep research and everything is automatically set.  You see when I change the preset it changes the  

  • 00:24:20 quantization fully automatically, everything is  fully working automatic depending on your GPU. So  

  • 00:24:25 you see 16 gigabytes preset using int8 version, 24  uses BF16 highest quality, 10 gigabyte uses NF4.  

  • 00:24:34 So this application is really good. You can always  ask me any questions that you have from Discord  

  • 00:24:39 or from YouTube as a comment or from Patreon. I  hope you have enjoyed, hopefully see you later.

Clone this wiki locally