Choose a category

3 models available
8 models available

3 models available
5 models available
4 models available

8 models available
Kling 2.6 Image to Video
Animate a photo with Kling's motion — and let it speak.
Creative Vision
Describe the motion. With sound on, dialogue in quotes gets voiced.
Video Parameters
Adds voices and sound effects. Doubles the credit cost.
Cinema Preview
Generation History0 records
Kling 2.6 Image to Video Generator
Upload one photo and Kling 2.6 animates it — and, if you want, gives the person in it a voice. Kuaishou's December 2025 model generates speech and picture in the same pass, so a line you write in the prompt comes back spoken, in sync. Five or ten seconds, from 15 credits.
The clip keeps your image's aspect ratio — crop the photo before you upload it, not after
Why Choose Kling 2.6 Image to Video
Most image-to-video models give you motion and silence. Kling 2.6 is the pairing on this site where a still photograph can end up speaking: audio is generated alongside the frames, which is why the mouth and the voice agree.
One Photo Is the Whole Setup
Upload a single JPEG or PNG up to 10MB and describe the motion. The image sets the subject, the lighting and the framing, so the prompt only has to carry what changes.
Give the Subject a Voice
Write a line in quotes with the speaker in brackets and the person in your photo delivers it, lips matching. Kuaishou tuned the 2.6 release for English and Chinese speech.
Your Image Sets the Frame
There is no aspect ratio control here, and that is deliberate: the clip inherits the shape of what you upload. A vertical phone photo returns a vertical video without cropping or letterboxing.
Pay for Sound Only When You Need It
Silent five seconds is 15 credits, with audio 25. Ten seconds is 25 silent and 50 with sound. Motion tests cost the low number; only the final take pays for the voice.
The Face Stays the Face
Cross-shot character consistency was one of Kuaishou's stated targets for 2.6, which matters more here than in text-to-video: the reference is a real person you can compare the output against.
Failed Runs Are Refunded
Credits are held at submission and returned automatically if generation fails or the provider never responds — including runs rejected for an unusable input image.
Same Price as Text to Video
Adding an image costs nothing extra. A photo-driven clip and a prompt-only clip of the same length and audio setting are billed identically.
How to Use Kling 2.6 Image to Video
Two things decide the result here: the photo you feed it and how you word the motion. The prompt is not describing a scene from scratch — it is describing what changes in a scene that already exists.
Upload a JPEG or PNG
One image, up to 10MB. WebP is not accepted by this endpoint — convert it first. Sharp, well-lit photos where the subject is unobstructed animate far more reliably than dark or busy ones.
Describe the Motion, Not the Photo
The model can already see the picture. Spend the prompt on movement and camera: "she turns toward the window as the camera pushes in slowly." Re-describing the outfit and the room wastes the instruction budget.
Add a Line if It Should Speak
Name the speaker and the tone in brackets, then the line in quotes: [young woman, calm voice] says: "I'll be there in five minutes." Keep it short enough to actually say inside the clip length.
Test Silent, Then Commit
Run five seconds without audio for 15 credits to check that the motion reads and the face holds. Once it does, enable sound and rerun. Generation takes a few minutes and you can leave the page.
One Change per Run
There is no seed on this model, so identical settings still produce different clips. Change one thing between runs — the motion verb, the camera, the line — or you will not know which edit helped.
Kling 2.6 Image to Video
FAQ
Image requirements, audio, aspect ratio, clip length and credit costs for the Kling 2.6 image-to-video model.
One image per run, JPEG or PNG, up to 10MB. WebP is rejected by this endpoint even though other models on this site accept it — if your file is a WebP, convert it to PNG or JPEG before uploading. The upload control enforces this, so an unsupported file is caught before any credits are held.
Not here. This page takes exactly one image. Wan 2.7 image to video on this site has explicit first-frame and last-frame slots if you need to define both ends of a shot, and Seedance 2.0 takes up to five reference images you can address individually in the prompt.
Between 15 and 50 credits, driven by length and the audio switch: 15 for five silent seconds, 25 for five with sound or ten silent, 50 for ten seconds with sound. It is the same price as Kling 2.6 text to video — the input image does not add anything.
No, and you do not need to. The output follows the aspect ratio of the image you upload, so a 9:16 photo gives a 9:16 clip. To control the shape of the video, crop the photo before uploading it.
The model picks a voice consistent with the subject it sees and syncs the mouth to it, but you cannot supply a specific voice sample on this endpoint. Describing the speaker in brackets — age, gender, tone — is the control you have over how it sounds.
Only with that person's permission. Every submission is screened before it reaches the model, and uploads of identifiable people used without consent, public figures placed in fabricated situations, or anything sexual involving a real likeness will be rejected. Your own photos, people who have agreed, and generated or licensed images are all fine.
A few minutes, longer for ten-second runs with audio. The finished clip appears in your history and stays available for seven days, so download anything you plan to reuse. If a run fails, the held credits are returned automatically.
Try Kling 2.6 Image to Video
Upload one photo, describe the motion, and see it move for 15 credits.