Flow Veo 3 Apps
Choose a category
Model

Grok Imagine
2 models available
Wan
2 models available

HappyHorse
3 models available
1.1 Text to Video
1.1 Image to Video
1.1 Reference to Video
Hailuo
2 models available
HappyHorse 1.1 Text to Video
Cinematic 1080p with sound baked in, up to 15 seconds.
Creative Vision
Describe the scene, the camera and the motion. Dialogue in the prompt is spoken in the video.
Video Parameters
Longer videos cost more.
Cinema Preview
Generation History0 records
HappyHorse 1.1 Text to Video Generator
HappyHorse 1.1 is the model on this site that comes back with sound. Alibaba's cinematic video model generates its own audio, so a prompt with dialogue returns a clip where the dialogue is spoken and the lips match. Nine aspect ratios, 3 to 15 seconds, 720p or 1080p.
Best for shots that need to be heard as well as seen: dialogue, ambience and footsteps arrive with the video
Why Choose HappyHorse 1.1 Text to Video
Released by Alibaba in June 2026, HappyHorse 1.1 was built for short cinematic pieces rather than clip demos. Its distinguishing feature is native audio: sound is generated with the picture, not added afterwards.
Native Audio and Lip Sync
Write dialogue into the prompt and it comes back spoken, in sync, in the language you wrote it in. Ambience and effects are generated alongside. There is no audio setting to switch on because it is not optional.
720p or 1080p
720p for drafts, 1080p for the take you keep. At 3 seconds both tiers round to the same 20 credits, so the shortest test run costs nothing extra at full resolution.
Nine Aspect Ratios
16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 21:9 and 9:21. The two extremes are the useful ones: 21:9 for a cinematic cut, 9:21 for a full-bleed phone screen.
3 to 15 Seconds
A per-second slider. Three seconds is enough to check framing and voice for 20 credits; fifteen is long enough to carry a complete beat of a scene.
Multi-Shot Prompts
The model handles prompts written as a shot list, with cuts described in order. Version 1.1 improved subject consistency, so a face stays the same face across those cuts.
Failed Runs Are Refunded
Credits are held when you submit and returned automatically if the generation fails or the provider never responds. A failed clip costs you nothing.
Test in Seconds, Render in Full
A 3-second 1080p test costs 20 credits and already tells you how the voice and motion land. Extend to the full length once the prompt is right.
How to Use HappyHorse 1.1 Text to Video
Because audio is generated with the picture, the prompt is doing two jobs at once. Writing for both is what separates a usable clip from a silent-looking one.
Describe the Shot
Scene, camera and motion first. The model responds to camera language: slow push in, handheld follow, static wide, rack focus. Multiple shots in sequence work well.
Write the Sound Too
Add the line you want spoken, in quotes, and the ambience you want under it. "She says: we should go back" gives you speech; "rain on a tin roof" gives you the bed under it.
Test Short
Run 3 seconds first. It costs 20 credits at either resolution and shows you the subject, the framing and the voice, which is where most prompts need fixing.
Render the Full Take
Same prompt, now at full length and 1080p. Generation takes a few minutes, you can leave the page, and the finished clip lands in your history with audio attached.
One Thing at a Time
There is no seed on this model, so two runs of the same prompt will differ. Change one element per run rather than rewriting the prompt, or you will not know what caused the change.
HappyHorse 1.1 Text to Video
FAQ
Audio, resolutions, aspect ratios, clip length and credit costs for the HappyHorse 1.1 text-to-video model.
Try HappyHorse 1.1 Text to Video
Three seconds at 1080p costs 20 credits and tells you how the picture and the voice land together.