<aside> ๐
Real-Time Long Video Generation (Fastest Variant)
Most AI video tools top out at short clips (5โ10 seconds) or slow to a crawl for longer content โ and they often need an expensive multi-GPU cluster to keep up. Helios Distilled is the fastest, most efficient member of the Helios family: it turns text prompts, images, or existing videos into fluid, high-quality clips of up to ~60 seconds, generated in real time on a single GPU. Purpose-built for teams that need production-ready video at speed and scale. Built by PKU-YuanGroup.
</aside>
Helios Distilled is an AI video generation model that turns text prompts, images, or existing videos into fluid, high-quality video content. It is the fastest and most efficient version in the Helios model family, purpose-built for teams that need production-ready video output at speed and scale. Unlike most competing video AI tools that struggle with longer clips or require expensive multi-GPU infrastructure, Helios Distilled generates up to 60 seconds of coherent, high-quality video in real time.
Input:
Text to Video: Describe the scene you want in plain language and Helios generates the video

Image to Video: Provide a starting image and a description of the motion or scene
Video to Video: Provide an existing video clip and a description to transform or extend it
Output: A fully generated MP4 video file
Up to ~60 seconds in length
High visual quality with smooth, coherent motion throughout

Parameters:
Resolution โ fixed at 640 ร 384 (5:3) โ not user-selectable in this build
Frame count (chunk multiples, default 99 frames (~4s @ 24fps)) โ pick from a fixed list of chunk-multiple values that maps to standard clip lengths at 24 fps:
| Frames | Duration @ 24 fps |
|---|---|
| 99 (default) | ~4 s |
| 198 | ~8 s |
| 297 | ~12.5 s |
| 396 | ~16.5 s |
| 495 | ~21 s |
| 594 | ~25 s |
| 693 | ~29 s |
| 792 | ~33 s |
| 891 | ~37 s |
| 990 | ~41 s |
| 1,089 | ~45 s |
| 1,188 | ~49.5 s |
| 1,287 | ~54 s |
| 1,386 | ~58 s |
| 1,452 | ~60.5 s |
Frame counts must be chunk multiples โ the model generates in 33-frame autoregressive chunks (99 = 3 chunks, 198 = 6 chunks, and so on), which is why the dropdown offers only these specific values.
Output FPS (dropdown, default 24) โ choose 16 or 24 fps
24 โ standard cinematic frame rate16 โ lighter alternative for shorter or more animation-style outputAdvanced Options:
| Real-time generation speed | 19.5 frames per second on a single H100 GPU |
| Maximum video length | ~ 60 seconds |
| Infrastructure needed | Single GPU (no multi-GPU cluster required) |
HuggingFace: