Repository navigation
Generation
AI generation is back, now on the Diffusion Studio API. Make images, video, music, sound effects and speech from the prompt bar, which shows your credits and what a generation will cost before you start. The model menu is searchable, and video models get a duration slider. New voice model: Gemini 3.8 Flash TTS.
generateandjobtools: agents can run any model, price it first withestimate, and poll the job until its files are saved to the library. Every model and its fields are listed in the models reference.
Changes
- Tool names: the
media_prefix is gone.media_grabis nowgrab(diffusion grab), and the same goes forprobe,transcribe,filmstrip,waveformandlisten. Object segmentation moved frommedia_segmenttogenerate sam-2.1, and themodelsandvoicestools are replaced by the models reference. - Agent chat: paste files straight into the message box.
- Captions: a shimmer shows while captions are being generated.
Full Changelog: v0.209.1...v0.210.0