Choose a category

3 models available
4 models available

3 models available



2 models available
4 models available

8 models available
HappyHorse 1.1 Reference to Video
Keep faces and products consistent across a brand-new shot.
Creative Vision
Name each subject where you use it, e.g. "the woman in the red dress in [Image 1] hands a cup to the man in [Image 2]".
Video Parameters
Longer videos cost more.
Cinema Preview
Generation History0 records
HappyHorse 1.1 Reference to Video Generator
Upload up to 9 reference images, name them in the prompt as [Image 1] through [Image 9], and HappyHorse 1.1 puts those subjects into a scene none of the photos ever contained. 1080p, native audio, 3 to 15 seconds.
Your reference images never appear as a frame: they tell the model what each subject looks like, not where the shot starts
Why Choose HappyHorse 1.1 Reference to Video
Image to video is limited to what is already in the frame. Reference to video is not: it takes the identity of each subject from your reference images and composes an entirely new shot around them, which is how you keep a character or a product recognisable across clips.
Up to 9 Subjects, One Scene
Bring a cast together: characters, a product, a location reference. Upload each one and refer to them as [Image 1] through [Image 9] in the prompt, in the order you uploaded them.
Consistency Across Runs
Subject consistency is what version 1.1 improved most. Reuse the same reference images across several generations and the same faces and products carry between them, which a text prompt alone cannot do.
Native Audio and Lip Sync
Sound is generated with the picture. Give a referenced character a line of dialogue and it comes back spoken, in sync, in the language you wrote it in.
Nine Aspect Ratios
This page composes the frame itself, so you choose the shape: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 21:9 or 9:21. A reference image never constrains the ratio.
3 to 15 Seconds at 1080p
A per-second slider from 3 to 15, at 720p or 1080p. Three seconds costs 20 credits at either resolution, which makes checking a likeness cheap.
Failed Runs Are Refunded
Credits are held at submission and returned automatically if generation fails or the provider never responds. A run that produces nothing costs nothing.
Build a Series, Not a One-Off
The point of reference to video is the second clip and the third: same subjects, new scene, recognisably the same cast. Keep your reference images and reuse them.
How to Use HappyHorse 1.1 Reference to Video
This model has one rule that decides whether it works at all: the prompt has to say which subject in which image you mean. Everything else is ordinary prompting.
Upload the References
Up to 9 images, 20MB each, with the shortest side at least 400 pixels; 720p or better is recommended. Blurry or heavily compressed references degrade the whole result.
Name Them in the Prompt
Upload order is reference order: the first image is [Image 1], the second [Image 2], and so on. Say which object you mean, as in "the woman in the red dress in [Image 1]", rather than a bare "[Image 1]" — naming the subject is what makes the likeness carry.
Describe the New Scene
Now write the shot you actually want: setting, action, camera and any spoken line. The references supply the subjects, so the prompt should spend its words on everything else.
Test, Then Render
Run 3 seconds first to check the likenesses carried over. If they did, run the full length at 1080p. If one did not, swap in a clearer reference for that subject rather than rewriting the prompt.
A Reference Is Not a First Frame
Reference to video takes only the identity of each subject and builds the frame from scratch. If you need the clip to open on your actual photo, that is a different tool on this site.
HappyHorse 1.1 Reference to Video
FAQ
Reference image requirements, prompt syntax, audio, aspect ratios and credit costs for the HappyHorse 1.1 reference-to-video model.
Image to video uses your photo as the literal first frame, so the clip begins on that exact picture. Reference to video never shows your images: it learns what the subjects look like and composes a completely new shot around them. That is why this page has an aspect ratio menu and image to video does not.
Write [Image 1], [Image 2] and so on, matching the order you uploaded them, and say which object in that image you mean. "The woman in the red qipao in [Image 1] walks through a night market" works. A bare "[Image 1]" with no description is much weaker.
Up to 9. That said, more is not automatically better: every extra subject is one more thing the prompt has to place and the model has to keep straight. Start with the one or two that matter and add others only when the shot genuinely needs them.
The shortest side must be at least 400 pixels and 720p or higher is recommended, up to 20MB each, in JPEG, PNG or WebP. Use clear, well-lit shots where the subject is unobstructed. Small, blurry or over-compressed images degrade the likeness, and no amount of prompting recovers it.
Between 20 and 100 credits, priced per second: about 5.1 credits per second at 720p and 6.5 at 1080p, rounded up. Three seconds costs 20 credits at either resolution, 5 seconds at 1080p is 35, and 15 seconds at 1080p is 100. The exact figure is shown before you submit.
Yes. HappyHorse 1.1 produces audio with the picture, including spoken dialogue with matching lip movement, in the language you wrote the line in. It is always on and there is no setting for it.
Close, but not identical. Version 1.1 improved subject consistency specifically, so reused reference images carry a recognisable likeness across runs, which is the point of this model. There is no seed on HappyHorse, so exact frame-level reproduction is not available.
Try HappyHorse 1.1 Reference to Video
Upload a reference, spend 20 credits on 3 seconds, and see how well the likeness carries into a new scene.