Welcome! This page helps you find the right documentation for each AI model in CNAPS Studio.


🔗 Open-Source Models

<aside>

🎵 Audio Models

Text to Speech (Qwen3)

Turn text into a natural spoken voiceover

</aside>

<aside>

🎙️ Audio Understanding

Speech Recognition (Clips)

Transcribe Every Clip in a Video List — Timed Transcripts

</aside>

<aside>

🎙️ Audio Understanding

Speech Recognition (Video)

Transcribe Speech in a Video into Timed Text

</aside>

<aside>

🏷️ Image Classification

Adult Content Detection

Flag Explicit Images Automatically

</aside>

<aside>

🏷️ Image Classification

Gender Recognition

Classify Pedestrian Gender in Photos

</aside>

<aside>

🏷️ Image Classification

Object Classification

Identify What's in a Photo (1,000 Categories)

</aside>

<aside>

🎛️ Image Control

ControlNet XL Canny

Generate Images That Follow an Edge Map

</aside>

<aside>

🎛️ Image Control

ControlNet XL Union

Generate Images from Any Control Map — 6 Modes in One Model

</aside>

<aside>

🎛️ Image Control

Depth Anything Annotator

Turn Any Photo into a Depth Map for ControlNet

</aside>

<aside>

🎨 Image Edit

FireRed-1.1

Edit with Text Instructions — Best-in-Class Identity Preservation

</aside>

<aside>

🎨 Image Edit

LatentDiffusion (Object Removal)

Remove People, Objects, or Blemishes Seamlessly

</aside>

<aside>

🎨 Image Edit

QWEN-Image-Edit-2511

Edit Images with Natural-Language Instructions

</aside>

<aside>

🎨 Image Edit

QWEN-Inpaint

Edit One Region, Leave the Rest Untouched

</aside>

<aside>

🎨 Image Edit

QWEN-Layered

Split a Flat Image into Editable Photoshop-Style Layers

</aside>