Reference to Video AI: Every Model, Its Limits and Its Price
Reference-to-video AI keeps your characters, products and voices in a new shot. Every model on Reelina, how many references each takes, and its price.

Image-to-video starts from one picture. Reference-to-video is different: you give the model several things to keep — a character, a product, an outfit, a place, sometimes a voice — and it builds a new shot around them. It is the closest thing AI video has to casting.
Reelina runs 26 reference-to-video models. They differ less in what they promise than in what they accept: how many images, whether a reference video is allowed, whether sound comes with it, and how long the result can be. This guide lists them from the Reelina catalogue as of October 2026, with the price of one generation in coins at each model's default length.
How reference-to-video works
You upload references — usually one image per subject — and write a prompt that says what they do. The model is not trained on your references; it reads them for that one generation. That has two consequences:
- Consistency comes from reusing the same references. Keep a clean portrait of each character and give it to every shot.
- One subject per image works best. A plain background and even light keep the model from copying the lighting into the scene.
The models, by what they accept
| Model | Coins | References | Length |
|---|---|---|---|
| PixVerse v6 Reference-to-Video | 38 | References | 1–15 s |
| PixVerse c1 Reference-to-Video | 46 | References | 1–15 s |
| Wan 3.0 Reference-to-Video | 60 | References | 2–30 s |
| Vidu Q3 Reference-to-Video | 64 | References, optional audio | — |
| Seedance 2.0 Mini Reference-to-Video | 68 | Up to 9 images, 3 videos (15 s total), reference audio | 4–15 s |
| Grok Imagine Video Reference-to-Video | 76 | 1–7 images (people, objects, styles) | — |
| Wan 3.0 Prime Reference-to-Video | 92 | References | 2–30 s |
| Vidu Q2 Reference-to-Video | 96 | 1–7 subjects, optional audio | — |
| Kling Video O3 Std Reference-to-Video | 108 | Subject "elements" | 3–15 s |
| Seedance 2.0 Fast Reference-to-Video | 110 | Up to 9 images, 3 videos, reference audio | 4–15 s |
| MiniMax H3 Fast Reference-to-Video | 112 | References | 5–15 s (8 s default) |
| Grok Imagine Video v1.5 Reference-to-Video | 122 | 1–7 images | — |
| Vidu Q2-Pro Reference-to-Video | 128 | 1–7 subjects, optional audio | — |
| Seedance 2.0 Reference-to-Video | 136 | Up to 9 images, 3 videos, reference audio | 4–15 s |
| Kling Video O3 Pro Reference-to-Video | 144 | Subject "elements" | 3–15 s |
| Wan 2.7 Reference-to-Video | 152 | Up to 3 videos — one per character, which can carry its voice | — |
| MiniMax H3 Reference-to-Video | 152 | References | 4–15 s (8 s default) |
| Seedance 2.5 Reference-to-Video | 204 | Up to 30 images, 10 videos, reference audio | 4–30 s |
| Veo 3.1 Reference-to-Video | 484 | References, optional audio | 8 s |
Also in the catalogue: Vidu Q3-Mix (160), Gemini Omni Flash (204), HappyHorse 1.1 (106) and 1.0 (212), Vidu Reference-to-Video 2.0 (304), Vidu Q1 (516) and Vidu Reference-to-Video Q1 (606). "—" means the length is fixed by the model rather than chosen. Prices are for each model's default length; longer clips cost proportionally more, and the exact cost is shown before you generate.
Which one to pick
- Cheapest way to try the idea: PixVerse v6 (38 coins) or Wan 3.0 (60 coins, up to 30 seconds).
- A cast of characters plus a moving reference: Seedance 2.0 — up to 9 images and 3 reference videos, with reference audio. Draft on Mini (68), finish on the full model (136).
- Long, reference-heavy shots: Seedance 2.5 — up to 30 images and 10 videos, clips up to 30 seconds.
- Characters that should keep their own voice: Wan 2.7, where each reference video is one character and can carry its voice.
- Up to seven separate subjects: Vidu Q2 / Q2-Pro or Grok Imagine Video.
- Kling's take on it: Kling Video O3, with subject "elements".
Reference images for images, too
If what you need is a still, Nano Banana 2 Reference-to-Image (24 coins) and its Lite version (13 coins) build an image from references, and both accept a video clip as a reference.
How to do it on Reelina
- Open the Video studio and pick a reference-to-video model — for example Seedance 2.0 Reference-to-Video or Seedance 2.5 Reference-to-Video.
- Upload your references and describe the shot: who does what, where, and how the camera moves.
- Check the price shown on the button and generate. Keep the references — reusing them is what keeps the next shot consistent.
For a whole episode rather than a single shot, the AI drama generator does this for you: it casts each character as a reference portrait and continues every shot from the last frame of the one before.
FAQ
What is reference-to-video AI? A video model that takes several reference inputs — images, and in some models videos and audio — and generates a new shot that keeps those subjects, instead of animating a single start image.
How many reference images can I use? It depends on the model: up to 30 on Seedance 2.5, 9 on Seedance 2.0, 1–7 on Vidu Q2 and Grok Imagine Video. Wan 2.7 takes up to 3 reference videos instead.
Which reference-to-video model is cheapest? On Reelina, PixVerse v6 Reference-to-Video at 38 coins per generation.
Does the model learn my face? No. References are read for that one generation only; nothing is trained. Consistency comes from giving the same references to every shot.
Do I need an API key? No. Every model runs in the browser on one Reelina coin balance. There is no free tier.
Try it yourself
Every model on one credit balance, straight in your browser. No install, no watermark.
Start creatingKeep reading
Kling 3.0 vs Veo 3.1: Clip Length, Audio and Cost ComparedKling 3.0 vs Veo 3.1: 3–15 s clips vs 4, 6 or 8 s, aspect ratios, resolutions, which tiers make sound, and cost per second for every variant.
Seedance 2.0 vs Kling 3.0: Specs, Prices and When to Use EachSeedance 2.0 vs Kling 3.0: tiers, clip length, aspect ratios, audio, reference-to-video vs multi-shot and motion control, and cost per second in coins.
Seedance vs Seedream: Which ByteDance Model Do You Need?Seedance makes video, Seedream makes images — both from ByteDance. Variants, sizes, audio and coin prices side by side, and how to use them together.
