Choose a category

2 models available
4 models available

3 models available
2 models available
2 models available

3 models available
Kling 2.6 Image to Video
Animate a photo with Kling's motion — and let it speak.
Creative Vision
Describe the motion. With sound on, dialogue in quotes gets voiced.
Video Parameters
Adds voices and sound effects. Doubles the credit cost.
Cinema Preview
Generation History0 records
Kling 2.6 Image to Video Generator
Upload one photo and Kling 2.6 animates it — and, if you want, gives the person in it a voice. Kuaishou's December 2025 model generates speech and picture in the same pass, so a line you write in the prompt comes back spoken, in sync. Five or ten seconds, from 15 credits.
The clip keeps your image's aspect ratio — crop the photo before you upload it, not after
Why Choose Kling 2.6 Image to Video
Most image-to-video models give you motion and silence. Kling 2.6 is the pairing on this site where a still photograph can end up speaking: audio is generated alongside the frames, which is why the mouth and the voice agree.
One Photo Is the Whole Setup
Upload a single JPEG or PNG up to 10MB and describe the motion. The image sets the subject, the lighting and the framing, so the prompt only has to carry what changes.
Give the Subject a Voice
Write a line in quotes with the speaker in brackets and the person in your photo delivers it, lips matching. Kuaishou tuned the 2.6 release for English and Chinese speech.
Your Image Sets the Frame
There is no aspect ratio control here, and that is deliberate: the clip inherits the shape of what you upload. A vertical phone photo returns a vertical video without cropping or letterboxing.
Pay for Sound Only When You Need It
Silent five seconds is 15 credits, with audio 25. Ten seconds is 25 silent and 50 with sound. Motion tests cost the low number; only the final take pays for the voice.
The Face Stays the Face
Cross-shot character consistency was one of Kuaishou's stated targets for 2.6, which matters more here than in text-to-video: the reference is a real person you can compare the output against.
Failed Runs Are Refunded
Credits are held at submission and returned automatically if generation fails or the provider never responds — including runs rejected for an unusable input image.
Same Price as Text to Video
Adding an image costs nothing extra. A photo-driven clip and a prompt-only clip of the same length and audio setting are billed identically.
How to Use Kling 2.6 Image to Video
Two things decide the result here: the photo you feed it and how you word the motion. The prompt is not describing a scene from scratch — it is describing what changes in a scene that already exists.
Upload a JPEG or PNG
One image, up to 10MB. WebP is not accepted by this endpoint — convert it first. Sharp, well-lit photos where the subject is unobstructed animate far more reliably than dark or busy ones.
Describe the Motion, Not the Photo
The model can already see the picture. Spend the prompt on movement and camera: "she turns toward the window as the camera pushes in slowly." Re-describing the outfit and the room wastes the instruction budget.
Add a Line if It Should Speak
Name the speaker and the tone in brackets, then the line in quotes: [young woman, calm voice] says: "I'll be there in five minutes." Keep it short enough to actually say inside the clip length.
Test Silent, Then Commit
Run five seconds without audio for 15 credits to check that the motion reads and the face holds. Once it does, enable sound and rerun. Generation takes a few minutes and you can leave the page.
One Change per Run
There is no seed on this model, so identical settings still produce different clips. Change one thing between runs — the motion verb, the camera, the line — or you will not know which edit helped.
Kling 2.6 Image to Video
FAQ
Image requirements, audio, aspect ratio, clip length and credit costs for the Kling 2.6 image-to-video model.
Try Kling 2.6 Image to Video
Upload one photo, describe the motion, and see it move for 15 credits.