model_name: phoenix-3 and train_video_url. The video must be at a publicly accessible URL.
See Choose a Model and Training Path if you are still deciding.
Video training automatically creates a voice from your recording audio and sets it as the face’s
default_voice_id. You can change it later with Update Face.Recording structure
Phoenix-3 uses a 1 minute speaking + 1 minute listening format, not the Phoenix-4 / Phoenix-4.5 30-second structure.- 1 minute of natural speaking
- 1 minute of still / listening footage
- One continuous shot
Camera & Framing
- Place the camera at eye level so your face fills at least 25% of the frame
- Keep your face centered and clearly visible
- Realistic, cartoon, anime, and Pixar-style humans are all fine. Animals, mascots, and other non-human figures are not
- Well-lit space with a simple background
- Record in 1080p using a desktop app - avoid browser recording and low resolution cameras
Supported Video Formats
- Minimum frame rate: 25 FPS
- Minimum resolution: 1080p
- Maximum file size: 750 MB
- Supported formats:
.webmand.mp4(H.264 video codec + AAC audio codec)

