Choose a category

2 models available
4 models available

3 models available
2 models available
2 models available

3 models available
Kling 2.6 Text to Video
Kling's cinematic motion, now with native dialogue and sound.
Creative Vision
Describe the scene and motion. With sound on, dialogue in quotes gets voiced.
Video Parameters
Adds voices and sound effects. Doubles the credit cost.
Cinema Preview
Generation History0 records
Kling 2.6 Text to Video Generator
Kling 2.6 generates picture and sound in the same pass, so a prompt containing a spoken line comes back as a clip where that line is actually spoken and the lips match. Released by Kuaishou in December 2025. Five or ten seconds, three aspect ratios, audio optional.
Write the shot under Visual: and the spoken lines under Dialog: — that is the format this model was tuned on
Why Choose Kling 2.6 Text to Video
Kuaishou released Kling 2.6 in December 2025 around one idea it calls simultaneous audio-visual generation: the soundtrack is produced with the frames rather than dubbed onto them afterwards. That is what this page gives you access to.
Dialogue That Lands on the Lips
Speech, narration, singing, ambience and effects are generated in the same pass as the picture, so timing and mouth shapes agree. Kuaishou trained the release for both English and Chinese lines.
Audio Is a Switch, Not a Tax
Leave sound off and a 5-second clip costs 15 credits. Turn it on and the same clip costs 25. You decide per run whether the shot needs to be heard, instead of paying for audio you will mute.
Consistency Across Cuts
Version 2.6 was benchmarked by Kuaishou for cross-shot character consistency and roughly 15% better instruction following than 2.5. A prompt describing two shots keeps the same face in both.
Landscape, Vertical or Square
16:9, 9:16 and 1:1. Vertical is the one worth knowing about: Kling comes out of a short-video platform, and 9:16 framing behaves like a native format rather than a crop.
Five or Ten Seconds
Two fixed lengths, not a slider. Five seconds fits a single line of dialogue comfortably; ten fits a short exchange or a line plus a reaction shot.
Failed Runs Are Refunded
Credits are held when you submit and returned automatically if the generation fails or the provider never answers. A clip that does not arrive costs you nothing.
Draft Silent, Finish With Sound
A silent 5-second run costs 15 credits and settles framing, motion and pacing. Switch audio on only for the take you intend to keep.
How to Use Kling 2.6 Text to Video
This model reads a structured prompt better than a paragraph. The examples Kuaishou ships split the description into a visual block and a dialogue block, and following that shape is the single biggest quality difference on this page.
Write the Visual Block
Open with the setting, the subject and the camera: "Visual: in a bright rehearsal room, sunlight through the window, the camera slowly circles the band." Camera verbs work — push in, circle, rack focus, static wide.
Write the Dialogue Block
Then name the speaker and the tone in brackets and put the line in quotes: Dialog: [female host, cheerful voice] says: "double-sided fleece, thirty dollars off today." Bracketed speakers are how the model keeps two voices apart.
Test With Sound Off
Run 5 seconds silent first for 15 credits. Composition, motion and subject are all visible without audio, and they are what usually needs another pass. Fix those before you pay for the voice.
Turn Sound On for the Keeper
Same prompt, audio enabled, 25 credits for five seconds or 50 for ten. Generation takes a few minutes, you can leave the page, and the finished clip appears in your history with the audio track attached.
Keep Lines Short
The prompt box holds 2,500 characters, but the constraint that bites is time, not characters: a line that takes eight seconds to say will not fit in a five-second clip. Count the words out loud before you submit.
Kling 2.6 Text to Video
FAQ
Audio, dialogue, aspect ratios, clip length and credit costs for the Kling 2.6 text-to-video model.
Try Kling 2.6 Text to Video
Five silent seconds costs 15 credits and tells you whether the shot works before you pay for the voice.