Reelina
← All articles
Video Generation·By Reelina

Reference to Video AI: Every Model, Its Limits and Its Price

Reference-to-video AI keeps your characters, products and voices in a new shot. Every model on Reelina, how many references each takes, and its price.

Reference to Video AI: Every Model, Its Limits and Its Price

Image-to-video starts from one picture. Reference-to-video is different: you give the model several things to keep — a character, a product, an outfit, a place, sometimes a voice — and it builds a new shot around them. It is the closest thing AI video has to casting.

Reelina runs 26 reference-to-video models. They differ less in what they promise than in what they accept: how many images, whether a reference video is allowed, whether sound comes with it, and how long the result can be. This guide lists them from the Reelina catalogue as of October 2026, with the price of one generation in coins at each model's default length.

How reference-to-video works

You upload references — usually one image per subject — and write a prompt that says what they do. The model is not trained on your references; it reads them for that one generation. That has two consequences:

  • Consistency comes from reusing the same references. Keep a clean portrait of each character and give it to every shot.
  • One subject per image works best. A plain background and even light keep the model from copying the lighting into the scene.

The models, by what they accept

Model Coins References Length
PixVerse v6 Reference-to-Video 38 References 1–15 s
PixVerse c1 Reference-to-Video 46 References 1–15 s
Wan 3.0 Reference-to-Video 60 References 2–30 s
Vidu Q3 Reference-to-Video 64 References, optional audio —
Seedance 2.0 Mini Reference-to-Video 68 Up to 9 images, 3 videos (15 s total), reference audio 4–15 s
Grok Imagine Video Reference-to-Video 76 1–7 images (people, objects, styles) —
Wan 3.0 Prime Reference-to-Video 92 References 2–30 s
Vidu Q2 Reference-to-Video 96 1–7 subjects, optional audio —
Kling Video O3 Std Reference-to-Video 108 Subject "elements" 3–15 s
Seedance 2.0 Fast Reference-to-Video 110 Up to 9 images, 3 videos, reference audio 4–15 s
MiniMax H3 Fast Reference-to-Video 112 References 5–15 s (8 s default)
Grok Imagine Video v1.5 Reference-to-Video 122 1–7 images —
Vidu Q2-Pro Reference-to-Video 128 1–7 subjects, optional audio —
Seedance 2.0 Reference-to-Video 136 Up to 9 images, 3 videos, reference audio 4–15 s
Kling Video O3 Pro Reference-to-Video 144 Subject "elements" 3–15 s
Wan 2.7 Reference-to-Video 152 Up to 3 videos — one per character, which can carry its voice —
MiniMax H3 Reference-to-Video 152 References 4–15 s (8 s default)
Seedance 2.5 Reference-to-Video 204 Up to 30 images, 10 videos, reference audio 4–30 s
Veo 3.1 Reference-to-Video 484 References, optional audio 8 s

Also in the catalogue: Vidu Q3-Mix (160), Gemini Omni Flash (204), HappyHorse 1.1 (106) and 1.0 (212), Vidu Reference-to-Video 2.0 (304), Vidu Q1 (516) and Vidu Reference-to-Video Q1 (606). "—" means the length is fixed by the model rather than chosen. Prices are for each model's default length; longer clips cost proportionally more, and the exact cost is shown before you generate.

Which one to pick

  • Cheapest way to try the idea: PixVerse v6 (38 coins) or Wan 3.0 (60 coins, up to 30 seconds).
  • A cast of characters plus a moving reference: Seedance 2.0 — up to 9 images and 3 reference videos, with reference audio. Draft on Mini (68), finish on the full model (136).
  • Long, reference-heavy shots: Seedance 2.5 — up to 30 images and 10 videos, clips up to 30 seconds.
  • Characters that should keep their own voice: Wan 2.7, where each reference video is one character and can carry its voice.
  • Up to seven separate subjects: Vidu Q2 / Q2-Pro or Grok Imagine Video.
  • Kling's take on it: Kling Video O3, with subject "elements".

Reference images for images, too

If what you need is a still, Nano Banana 2 Reference-to-Image (24 coins) and its Lite version (13 coins) build an image from references, and both accept a video clip as a reference.

How to do it on Reelina

  1. Open the Video studio and pick a reference-to-video model — for example Seedance 2.0 Reference-to-Video or Seedance 2.5 Reference-to-Video.
  2. Upload your references and describe the shot: who does what, where, and how the camera moves.
  3. Check the price shown on the button and generate. Keep the references — reusing them is what keeps the next shot consistent.

For a whole episode rather than a single shot, the AI drama generator does this for you: it casts each character as a reference portrait and continues every shot from the last frame of the one before.

FAQ

What is reference-to-video AI? A video model that takes several reference inputs — images, and in some models videos and audio — and generates a new shot that keeps those subjects, instead of animating a single start image.

How many reference images can I use? It depends on the model: up to 30 on Seedance 2.5, 9 on Seedance 2.0, 1–7 on Vidu Q2 and Grok Imagine Video. Wan 2.7 takes up to 3 reference videos instead.

Which reference-to-video model is cheapest? On Reelina, PixVerse v6 Reference-to-Video at 38 coins per generation.

Does the model learn my face? No. References are read for that one generation only; nothing is trained. Consistency comes from giving the same references to every shot.

Do I need an API key? No. Every model runs in the browser on one Reelina coin balance. There is no free tier.

#reference to video#AI video#Seedance#model comparison

Try it yourself

Every model on one credit balance, straight in your browser. No install, no watermark.

Start creating

Keep reading