AI Video Generation Models

43 models available on BudgetPixel

Every text-to-video and image-to-video model on BudgetPixel. Video models vary most on the things that decide whether a shot is usable: how well motion holds together, whether a start and end frame are respected, whether audio comes with it, and how long a single generation can run.

Video is billed per second of output rather than per generation, so the credit figure on each card is a rate — a short clip from an expensive model can cost less than a long one from a cheap model. Each card links to a full breakdown with real sample clips.

Open the Video Workshop

All video generation models

  • Wan 3.0 Primecredits

    High-speed tier of Alibaba's all-in-one Wan 3.0 — the same T2V, I2V (first+last frame), and reference-to-video capabilities with significantly faster generation. 480p/720p/1080p with optional audio. Any integer 2-30s duration at 30fps.

  • Wan 3.0credits

    Alibaba's all-in-one video model. T2V, I2V (first+last frame), and reference-to-video mixing up to 10 reference images, 5 reference video clips (editing, replication & extension), and 5 reference audio clips. 480p/720p/1080p with optional audio. Any integer 2-30s duration at 30fps.

  • SeeDance 2.5credits

    ByteDance's flagship multimodal video model. T2V, I2V (start+end frame), and reference-to-video mixing up to 15 reference images, 5 reference video clips (editing & extension), and 5 reference audio clips. 480p/720p/1080p with optional audio generation. Any integer 4-30s duration.

  • MiniMax H3credits

    MiniMax's flagship multimodal video model. Text-to-video, image-to-video (start + end frame), and reference-to-video.

  • SeeDance 2.0 Minicredits

    Best cost-performance variant of SeeDance 2.0. Supports T2V, I2V (start+end frame), and reference-to-video (up to 9 reference images, optional reference audio). 480p/720p with optional audio generation. Any integer 4-15s duration.

  • HappyHorse 1.1credits

    Newest HappyHorse video model with highly realistic dynamic rendering. Supports text-to-video, image-to-video, and reference-to-video (up to 9 reference images for subject-consistent generation). 720p/1080p at 3-15s. No video-edit mode (1.0 only).

  • Kling 3.0 Turbocredits

    Kuaishou's Kling 3.0 Turbo via the official Kling AI API — fast generations with native audio always included. Text-to-video and image-to-video (first frame) at 720p or 1080p, any duration from 3 to 15 seconds.

  • PixVerse C1credits

    PixVerse's cinematic model for high-quality, reference-driven generation. Supports text-to-video, image-to-video (start + end frame), and reference-to-video (up to 3 images). 360p/720p/1080p at 5s, 10s, or 15s.

  • Kling V3 Omnicredits

    Kuaishou's Kling V3 Omni model via the official Kling AI API.

  • Kling 3.0 4Kcredits

    Kuaishou's newest Kling model at native 4K resolution via the official Kling AI API. Flagship tier for hero shots and premium output. Text-to-video and image-to-video at any duration from 3 to 15 seconds with start/end frame control and optional audio.

  • Kling 3.0 Procredits

    Kuaishou's newest Kling model in Professional mode (1080p) via the official Kling AI API. Best balance of quality and cost in the v3.0 family. Text-to-video and image-to-video at any duration from 3 to 15 seconds with start/end frame control and optional audio.

  • Kling 3.0 Standardcredits

    Kuaishou's newest Kling model in Standard mode (720p) via the official Kling AI API. Cost-effective tier with the same prompt adherence and start/end frame control as Pro. Text-to-video and image-to-video at any duration from 3 to 15 seconds.

  • SeeDance 2.0 Fastcredits

    Faster and more affordable variant of SeeDance 2.0. Supports T2V, I2V (start+end frame), and reference-to-video (up to 9 reference images, optional reference audio). 480p/720p with optional audio generation. Any integer 4-15s duration.

  • SeeDance 2.0credits

    ByteDance's latest video generation model. Supports T2V, I2V (start+end frame), and reference-to-video (up to 9 reference images, optional reference audio). 480p/720p/1080p/4K (10-bit) with optional audio generation. Any integer 4-15s duration.

  • PixVerse V6credits

    PixVerse's latest model with multi-clip narrative and audio. Supports text-to-video, image-to-video (start + end frame), and reference-to-video (up to 3 images). 360p/720p/1080p at 5s, 10s, or 15s.

  • Wan 2.7credits

    Latest Wan model with prompt-driven narrative, start+end frame I2V, video editing, and audio. Supports 720p/1080p at 5s, 10s, or 15s.

  • Wan 2.2 Animate Movecredits

    Animate a character from a reference image using a driving video. The character in the image will perform the movements shown in the video.

  • Wan 2.2 Animate Replacecredits

    Replace a character in a driving video with one from a reference image. The replacement character will perform the same actions as the original.

  • Vidu Q3 Procredits

    Vidu's Q3 Pro model with 540p/720p/1080p resolution, optional audio, and image-to-video. Supports 4s and 8s durations.

  • P Videocredits

    PrunaAI's fast video generation with native audio, text-to-video and image-to-video. Supports audio input for lip-sync and rhythm matching.

  • Grok Imagine Videocredits

    xAI's video generation model with flexible 1-15 second durations. Supports text-to-video and image-to-video at 480p or 720p.

  • Wan 2.6 I2V Flashcredits

    Fast and affordable image-to-video with multi-shot cinematic generation. Half the cost of Wan 2.6 I2V with same quality. Supports 720p/1080p at 5s, 10s, or 15s duration with audio.

  • SeeDance 1.5 Procredits

    ByteDance's latest video generation model with optional audio. Supports T2V and I2V (with start+end frame) at 480p, 720p, or 1080p resolution.

  • Wan 2.6credits

    Wan model with multi-shot cinematic video generation. Supports 720p/1080p at 5s, 10s, or 15s duration with audio.

  • Kling v2.6 Procredits

    Latest Kling model with optional audio generation. High quality text-to-video and image-to-video with start/end frame control and sound support.

  • Wan 2.5credits

    Advanced text-to-video and image-to-video with audio synchronization support.

  • MiniMax Hailuo 2.3 Fastcredits

    Faster and more affordable variant of Hailuo 2.3. Image-to-video only. Supports 768p (6s/10s) and 1080p (6s only).

  • MiniMax Hailuo 2.3credits

    Latest MiniMax video model with enhanced quality and motion. Supports 768p (6s/10s) and 1080p (6s only).

  • LTX-2 Fastcredits

    Lightricks' fast video generation model with high-quality output. Supports 1080p, 2k, and 4k resolutions with optional audio generation. Only supports 16:9 aspect ratio.

  • Wan 2.5 I2V Fastcredits

    Fast image-to-video with audio synchronization. Supports 720p and 1080p at 5s or 10s duration.

  • Wan 2.5 T2V Fastcredits

    Fast text-to-video with audio synchronization. Supports 720p and 1080p at 5s or 10s duration.

  • Veo 3.1 Fastcredits

    Faster video generation with audio output at 1080p (no reference images).

  • Veo 3.1credits

    Enhanced video generation with reference images and audio output at 1080p.

  • SeeDance 1 Litecredits

    ByteDance video generation model with text-to-video and image-to-video support for 5s or 10s videos at 480p, 720p, and 1080p resolution.

  • MiniMax Hailuo 02credits

    Balanced quality and speed from MiniMax.

  • Wan 2.2 I2V A14Bcredits

    High‑quality image‑to‑video pipeline.

  • Wan 2.2 I2V Fastcredits

    Fast and affordable Image‑to‑video conversion.

  • Wan 2.2 T2V Fastcredits

    Fast and inexpensive text‑to‑video.

  • Veo 3 Fastcredits

    Speed‑optimized Veo variant of veo-3.

  • SeeDance 1 Pro 1080pcredits

    1080p high quality video generation and stylization.

  • SeeDance 1 Pro 480pcredits

    480p video generation with a balance of quality and speed.

  • Veo 3credits

    Creative video generation with cinematic style.

  • HappyHorse 1.0credits

    New text-to-video, image-to-video, and reference-to-video model with highly realistic dynamic rendering. Supports up to 3 reference images for subject-consistent generation. 720p/1080p at 3-15s.

AI Video Generation Models — common questions

Which AI video model is best?
They trade off differently. Some hold motion together better across a longer shot, some respect a start and end frame, some generate synchronised audio, and some are simply far cheaper per second. Each model's page lists its durations, resolutions and whether it supports image-to-video.
How is AI video priced?
Per second of output, not per generation — so the credit figure on each card is a rate. A 5-second clip from an expensive model can cost less than a 15-second clip from a cheaper one. The pricing block on each model page shows a worked example.
What is the difference between text-to-video and image-to-video?
Text-to-video generates a clip from a written prompt alone. Image-to-video animates a still you supply, which gives you much tighter control over composition and character appearance. Most models here do both; some also accept an end frame to steer where the shot finishes.
How long can a generated video be?
It varies by model, typically between 4 and 30 seconds per generation. Longer pieces are made by generating several clips and assembling them, which the Clip Editor in BudgetPixel Studio is built for.
Can I add sound?
Some video models generate audio natively. For everything else, the Audio Workshop generates music and sound effects you can lay under the clip, and the sound-effect models can sync effects to an existing video.

Browse other model types

  • AI Image Generation Models
  • AI Music Generation Models
  • AI Voice & Text-to-Speech Models
  • All AI models