Launching October 2026 · Kling 4.0

Kling 4.0: 30-Second AI Video in One Generation

Up to 15 references and 10 keyframes, stereo lip-sync, and output up to 4K. Until it lands on Pilio, start with Kling 3.0 in the same Studio

Create with Kling 3.0 now

Kling 4.0 is not live on Pilio yet; generation uses Kling 3.0 · Free credits on signup

Continue from a first frame, or lock characters with up to 15 references

Generate with Kling 3.0 now; your prompts carry over to Kling 4.0

  1. Image to video

    One first frame, and the scene plays on

    Characters, light, and framing all start from your frame

  2. Text to video

    One sentence, a full performance

    Name the subject, action, and sound; short prompts still work

  3. Reference to video

    A few references lock the cast and moves

    Kling 3.0 on Pilio takes up to 2 reference images

What's new in Kling 4.0

Longer, more controllable, closer to the final cut

Stereo audio, tighter lip-sync

Two-channel stereo with dialogue in many languages, accents, and dialects

Up to 10 keyframes

Pin character states and story turns on the timeline, not just the first and last frame

30 seconds in one generation

Pick 3–30 seconds for long takes and continuous stories without stitching

Omni Reference with up to 15 assets

Mix images, videos, and subjects to recreate a proven video with your own cast and product

Kling 4.0 in action

Motion, styles, and aspect ratios at a glance

Close-quarters action in ultrawide that stays stable at speed

Kling 4.0 specs as announced

Options on Pilio are confirmed at launch

Duration

3–30 seconds

Video extension up to 2 minutes is coming soon

Resolution

720p / 1080p / 4K

10-bit HDR at 1080p and 4K is coming soon

Aspect ratio

21:9 / 16:9 / 1:1 / 9:16

Adds 21:9 ultrawide

References

Up to 15

10 images, 5 videos (30 seconds total), 7 subjects; audio is voice reference only

Keyframes

Up to 10

Set frames and story beats at specific points in time

Audio

Two-channel stereo

Multiple languages, accents, and dialects with more accurate lip-sync

Prompt

Up to 8,000 tokens

Even a one-line idea gets expanded into a director-level brief

Video editing

Up to 5 input clips

Change expressions, motion, camera, style, or background while the rest stays put

Kling 4.0 Flash

3–20 seconds · 720p

8-bit SDR, faster generation for high-volume work

Write your shots on Kling 3.0 until 4.0 arrives

Generate with Kling 3.0 until Kling 4.0 lands; free credits on signup

Kling 4.0 FAQs

What we know about Kling 4.0 and how to use it on Pilio

When will Kling 4.0 be released?

Kling 4.0 launches in October 2026. Kling 4.0 Flash opened to a limited group of users for early access on September 28.

Can I use Kling 4.0 on Pilio now?

Not yet. The generator on this page uses Kling 3.0 today. We will add Kling 4.0 as soon as access opens and update this page.

What is new in Kling 4.0 compared with Kling 3.0?

Single-generation length goes from 15 to 30 seconds, and Kling 4.0 adds up to 10 keyframes, 15 multimodal references, 21:9, stereo audio, and clearer on-screen text. Kling 3.0 on Pilio supports 3–15 seconds, 1:1 / 9:16 / 16:9, and up to 2 reference images.

What is the difference between Kling 4.0 and Kling 4.0 Flash?

Flash is built for speed and value: text-to-video, first-frame image-to-video, and Omni Reference at 3–20 seconds, 720p, 8-bit SDR. Full Kling 4.0 adds 30-second generation, up to 4K, multi-keyframes, and video editing.

Does Kling 4.0 support 4K HDR?

It outputs up to 4K. 10-bit HDR at 1080p and 4K is coming soon.

How long can Kling 4.0 videos be?

Each generation is 3–30 seconds. An upcoming video extension feature lets you extend a clip several times, up to 2 minutes.

Which languages and accents does Kling 4.0 support?

Dialogue works in Chinese, English, Japanese, Korean, Spanish, Portuguese, German, French, Hindi, and more. Chinese dialects include Beijing, Taiwanese Mandarin, Northeastern, Sichuanese, and Cantonese; English accents include American, British, and Indian.

Will my Kling 3.0 prompts work with Kling 4.0?

Yes. The structure is the same: subject, action, camera, lighting, and sound. Prompts and references you prepare now carry over when you switch models.