Choose a category

3 models available
4 models available

3 models available



2 models available
4 models available

8 models available
HappyHorse 1.1 Text to Video
Cinematic 1080p with sound baked in, up to 15 seconds.
Creative Vision
Describe the scene, the camera and the motion. Dialogue in the prompt is spoken in the video.
Video Parameters
Longer videos cost more.
Cinema Preview
Generation History0 records
HappyHorse 1.1 Text to Video Generator
HappyHorse 1.1 is the model on this site that comes back with sound. Alibaba's cinematic video model generates its own audio, so a prompt with dialogue returns a clip where the dialogue is spoken and the lips match. Nine aspect ratios, 3 to 15 seconds, 720p or 1080p.
Best for shots that need to be heard as well as seen: dialogue, ambience and footsteps arrive with the video
Why Choose HappyHorse 1.1 Text to Video
Released by Alibaba in June 2026, HappyHorse 1.1 was built for short cinematic pieces rather than clip demos. Its distinguishing feature is native audio: sound is generated with the picture, not added afterwards.
Native Audio and Lip Sync
Write dialogue into the prompt and it comes back spoken, in sync, in the language you wrote it in. Ambience and effects are generated alongside. There is no audio setting to switch on because it is not optional.
720p or 1080p
720p for drafts, 1080p for the take you keep. At 3 seconds both tiers round to the same 20 credits, so the shortest test run costs nothing extra at full resolution.
Nine Aspect Ratios
16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 21:9 and 9:21. The two extremes are the useful ones: 21:9 for a cinematic cut, 9:21 for a full-bleed phone screen.
3 to 15 Seconds
A per-second slider. Three seconds is enough to check framing and voice for 20 credits; fifteen is long enough to carry a complete beat of a scene.
Multi-Shot Prompts
The model handles prompts written as a shot list, with cuts described in order. Version 1.1 improved subject consistency, so a face stays the same face across those cuts.
Failed Runs Are Refunded
Credits are held when you submit and returned automatically if the generation fails or the provider never responds. A failed clip costs you nothing.
Test in Seconds, Render in Full
A 3-second 1080p test costs 20 credits and already tells you how the voice and motion land. Extend to the full length once the prompt is right.
How to Use HappyHorse 1.1 Text to Video
Because audio is generated with the picture, the prompt is doing two jobs at once. Writing for both is what separates a usable clip from a silent-looking one.
Describe the Shot
Scene, camera and motion first. The model responds to camera language: slow push in, handheld follow, static wide, rack focus. Multiple shots in sequence work well.
Write the Sound Too
Add the line you want spoken, in quotes, and the ambience you want under it. "She says: we should go back" gives you speech; "rain on a tin roof" gives you the bed under it.
Test Short
Run 3 seconds first. It costs 20 credits at either resolution and shows you the subject, the framing and the voice, which is where most prompts need fixing.
Render the Full Take
Same prompt, now at full length and 1080p. Generation takes a few minutes, you can leave the page, and the finished clip lands in your history with audio attached.
One Thing at a Time
There is no seed on this model, so two runs of the same prompt will differ. Change one element per run rather than rewriting the prompt, or you will not know what caused the change.
HappyHorse 1.1 Text to Video
FAQ
Audio, resolutions, aspect ratios, clip length and credit costs for the HappyHorse 1.1 text-to-video model.
HappyHorse 1.1 is a video generation model released by Alibaba in June 2026, developed at the Future Life Lab inside its Taotian Group. Version 1.1 improved motion, subject consistency, prompt adherence and audio over 1.0. This page runs its text-to-video endpoint: you supply a prompt and settings, and we handle generation, credit accounting and result storage. FlowVeo3 is an independent platform and is not affiliated with Alibaba.
Yes, and it is not a setting you turn on. Sound is produced with the picture: spoken dialogue with matching lip movement, plus ambience and effects that follow the scene. That is the main reason to pick this model over the others on the site, which return silent video.
Between 20 and 100 credits. Pricing is per second: about 5.1 credits per second at 720p and 6.5 at 1080p, rounded up to a clean figure. So 3 seconds costs 20 credits at either resolution, 5 seconds at 1080p is 35, and the full 15 seconds at 1080p is 100. The exact cost is shown before you submit.
Because prices are rounded up to a clean figure and at three seconds both tiers land in the same bracket. It is a genuine quirk of the rounding, not a promotion, and it makes a 3-second 1080p test the cheapest useful run on this page.
No. HappyHorse 1.1 exposes only prompt, resolution, aspect ratio and duration, so there is no negative prompt and no seed. If you need seed-repeatable output, Wan 2.7 text to video on this site has both.
The model accepts prompts in any language and handles multilingual lip sync. Write the line in the language you want it spoken and it will be delivered in that language rather than translated.
Expect a few minutes, longer for 15-second 1080p runs. You can leave the page while it works and the clip appears in your history when it is done. If the run fails, your credits are returned automatically.
Try HappyHorse 1.1 Text to Video
Three seconds at 1080p costs 20 credits and tells you how the picture and the voice land together.