
Models for video, images, voice and music
The world's best models. One studio.
Make films, photos, voices and music with Veo, Kling, Seedance, Nano Banana, ElevenLabs and every other top model, side by side in one studio. Find the one that fits your idea, then make it yours.
- Start free, no card
- Every top model, one account
- Pay per run, priced before it starts
Video
Make it move
Direct a scene, animate a still, film a product. The models behind the clips people share.


Seedance 1.5 Pro
ByteDance
Sharp, steady video at a very low price, with optional sound.
- Text to video
- Image to video



FLUX.3 Video
Black Forest Labs
Black Forest Labs' first video model, from text or a first frame.
- Text to video
- Image to video

Gemini Omni Flash 1.1
Google
Google's fast video model, from 360p to 4K.
- Text to video
- Image to video
More video models
- Try
Fast video with sound, from text or a first frame.
- TryFLUX.3 Video EditBlack Forest Labs
Restyles or edits a video from a prompt at 720p, at three cents a second.
- TryOmniHuman 1.5ByteDance
The person in a photo speaks, sings or raps your recording, with the face, hands and body moving to it. Billed by the length of the recording.
- TryByteDance Video UpscalerByteDance
Upscales a video to 1080p at 30 fps and cleans it up, for under a cent a second.
- TryLucy Edit ProDecart
Changes one thing in a video from a prompt (an outfit, an object, a product, the background) and keeps the rest as filmed.
Image
Make it real
Posters, products, portraits and whole worlds, from one sentence or your own photo.


Seedream 4.5
ByteDance
2K images with crisp typography, and edits of up to 4 reference images.
- Text to image
- Image edit

Seedream 5.0 Pro
ByteDance
ByteDance's newest: detailed 2K images and edits of up to 4 reference images.
- Text to image
- Image edit

Recraft V4.1 Pro
Recraft
Recraft V4.1 at four megapixels, for print, and vectors as SVG.
- Text to image

xAI's second Grok Imagine, with a 2K option and a cheaper draft quality.
- Text to image

Nano Banana Pro
Google
Gemini 3 Pro Image: Google's best for legible text, infographics and edits of up to 6 references.
- Text to image
- Image edit
One prompt, 6 models
“A hand-painted shop sign that reads OPEN LATE above a tiny ramen bar on a rainy night, warm light spilling onto wet cobblestones, cinematic 35mm photo”
More image models
- TryFLUX.1 [schnell]Black Forest Labs
Fast, low-cost text to image. Good for drafts and volume.
- TryFLUX.2 [pro]Black Forest Labs
Black Forest Labs' second generation: sharper detail and text than FLUX1.1, at a lower price.
- TryFLUX.2 [flex]Black Forest Labs
FLUX.2 tuned for typography and fine detail, such as posters and layouts.
- TryFLUX.2 [klein] 9BBlack Forest Labs
A small, fast FLUX.2, a step up in quality from klein 4B.
- TryFLUX.2 [klein] 4BBlack Forest Labs
The smallest, fastest FLUX.2, for volume.
- TrySeedream 4.0ByteDance
The lowest-priced Seedream: 2K images and edits of up to 4 reference images.
Voice and sound
Give it a voice
Narration, characters, songs and sound effects for everything you make. Press play.
TTS-1
OpenAI
0 / 7 sRead the transcript
Hi, I'm testing Leap's voices. Can you hear the difference between each one? Let's find out together.
OpenAI's fast text to speech, in nine voices.
- Text to speech
TTS-1 HD
OpenAI
0 / 7 sRead the transcript
Hi, I'm testing Leap's voices. Can you hear the difference between each one? Let's find out together.
OpenAI's higher-quality text to speech, in nine voices.
- Text to speech
Grok TTS
xAI
0 / 7 sRead the transcript
Hi, I'm testing Leap's voices. Can you hear the difference between each one? Let's find out together.
Expressive speech in five voices.
- Text to speech
MAI-Voice-2.1
Microsoft
0 / 8 sRead the transcript
Hi, I'm testing Leap's voices. Can you hear the difference between each one? Let's find out together.
Microsoft's most natural voices, for narration, across 23 languages.
- Text to speech
More voice and sound models
- TryQwen Audio 3 TTSAlibaba
Alibaba's speech model, strong in Chinese and English.
- TrySeed Speech 2ByteDance
ByteDance's speech model, in Chinese and English.
- TryCassetteAI MusicCassetteAI
Royalty-free music in seconds, from a prompt.
- TryCassetteAI Sound EffectsCassetteAI
Sound effects up to 30 seconds long, in about a second.
3D
Make it solid
Turn a picture or a sentence into a 3D model you can spin, print or drop into a game.
- TryRodin 2.5 FastDeemos
Rodin 2.5 in seconds, at ten cents a model.
- TryMeshy 6 Text to 3DMeshy
Detailed, textured 3D models from a prompt.
- TryHunyuan3D 2Tencent
A textured 3D model from one image, as a GLB.
Use them together
One model makes the still. Another makes it move.
"5 da manhã, shot 4", replayed: Nano Banana Pro makes the still, then FLUX.3 Video brings it to life, in the same studio and the same library.


- 01Nano Banana ProMakes the stillDone
Low-angle wide shot, 24mm, dawn: the dancer from the reference image, same clothes, on the terrace of a Lisbon miradouro, terracotta rooftops and the...
- 02FLUX.3 VideoBrings it to lifeDone
He breaks the freeze into a fluid wave through his arms and body as the sun touches the river behind him; pigeons lift off the balustrade. Slow dolly...
"5 da manhã, shot 4", made on Leap: two models, one studio, one library.


![FLUX1.1 [pro] Ultra's answer to the same prompt: a ramen bar sign reading OPEN LATE on a rainy night](/_next/image?url=https%3A%2F%2Fgpdiatw6a9ifgvnx.public.blob.vercel-storage.com%2Fmedia%2Fcompare%2Fflux-1.1-pro-ultra.d5028b17.jpg&w=3840&q=75)




