<aside> ðŸ›
Real-Time Long Video Generation (Highest Quality Variant)
Most AI video tools either top out at a few seconds or need an expensive multi-GPU cluster to go longer — and quality drifts the moment a clip runs long. Helios-Base is the flagship, highest-quality member of the Helios family: it generates long-form video (up to ~45 seconds) from a text description, a starting image, or an existing clip, and it does it in real time on a single GPU at 19.5 FPS. When output quality is the priority, this is the variant to reach for. Built by PKU-YuanGroup.
</aside>
Helios-Base is the flagship, highest-quality variant of the Helios video generation model family, developed by PKU-YuanGroup. It generates long-form, visually rich video content from a text description, a reference image, or an existing video clip. While the distilled version of Helios prioritizes speed, Helios-Base is the go-to choice when output quality is the top priority, producing the most detailed, coherent, and visually polished results in the family. It still runs on a single GPU and generates video in real time at 19.5 FPS, making it both high-quality and cost-effective compared to larger competing models.
Text to Video: Describe the scene you want in plain language and Helios generates the video

Image to Video: Provide a starting image and a description of the motion or scene
Video to Video: Provide an existing video clip and a description to transform or extend it
Up to ~45 seconds in length
High visual quality with smooth, coherent motion throughout

| Real-time generation speed | 19.5 frames per second on a single H100 GPU |
| Maximum video length | ~ 45 seconds |
| Infrastructure needed | Single GPU (no multi-GPU cluster required) |
| Output quality | Highest in the Helios model family |
| Version | Best For |
|---|---|
| Helios-Base | Highest output quality - ideal for premium and production use cases |
| Helios-Mid | Intermediate quality/speed balance |
| Helios-Distilled | Fastest generation - best for high-volume or real-time use cases |
Note: Image-to-Video and Video-to-Video modes may produce slightly less consistent results than Text-to-Video, as the model was primarily trained on text-based generation.