One tool.
Every kind of
generation.
Text to image, image to video, reference to video, face and character swap, lip-sync, upscale — one balance, one interface, no separate tools to learn. Open models on GPUs we run ourselves.
$5 is 1,200 Takes on your first top-up — hundreds of images, or dozens of clips. Browse every model and its price before you sign up.
text → image
text → image
text → imageHow it works
No install. Pick a model, drop in a prompt or a source image, and you're looking at results in a few minutes — or seconds if a GPU is already warm.
Give it something
A prompt, a photo, a short clip, or an audio track — whatever the mode calls for. Your own LoRAs load per model.
Pick the mode
Text to image, image to video, reference to video, swap, lip-sync, upscale — same interface, same balance.
Keep what you make
Full resolution, no watermark, yours to use. Unused Takes never expire. Failed jobs refund themselves.
Under the hood
Several open models, picked per mode, so you're not stuck with one engine for everything. Each one runs on hardware we rent by the hour — which is also why it's priced the way it is.
MiniMax H3
Video with synchronised audio, in a single pass.
Wan 2.2
Video. Animate a still, or put your character into a performance.
Krea2 Turbo
Images. Our default — fast, and the one model with a style LoRA.
Flux 2 Klein
Images. 4-step draft speed.
Qwen-Image
Images. Generate from text, or edit an image you already have.
InfiniteTalk
Lip-sync. Make a photo or a clip talk, driven by audio.
FaceFusion
Face swap. Graft one face onto an image or video.
SeedVR2
Upscale. Restore and enlarge video to 1080p, 2K or 4K.
LoRAs load per model — pick your base first, then the LoRA trained for it. Civitai and Hugging Face links, or a .safetensors upload.
Not a fixed pipeline
Everything runs through ComfyUI on our fleet — the same graphs you could run yourself, with the queue, the GPUs and the pricing handled. New modes and checkpoints ship when we add them, not when a vendor does.
Generation types get added as we build them — not gated behind someone else's API update.
Resolution, duration, aspect, seed, step-level settings where the model exposes them — tuned per mode instead of one default for everything.
Bring a style, a character or a subject from Civitai or Hugging Face. Every generation on that model leans toward what it was trained on.
$5 = 1,200 Takes.
An image is 1. A 4-second clip is 16.
Takes are prepaid credit. Every generation shows its price before you run it; that is what you pay. Bigger top-ups buy more per dollar: 120/$ from $5, 140/$ from $20, 160/$ from $50.
Start with $5USD shown at the best rate (160 Takes/$). API use is priced above the web rate.
Made with $5





