Skip to main content
Use this path when you call Create Face with model_name: phoenix-4.5 and train_video_url. The video must be at a publicly accessible URL. See Choose a Model and Training Path if you are still deciding. Phoenix-4.5 does not support full-body, head-to-toe framing. For that, use Phoenix-4 Video Requirements.
Video training automatically creates a voice from your recording audio and sets it as the face’s default_voice_id. You can change it later with Update Face.
A zero-shot preview becomes usable in minutes, then tunes in the background. The preview carries a visible watermark; the later tuned version does not. Footage takes more effort than a photo, but the tuned face usually moves more naturally because tuning learns from your real movement.

Camera & Framing

Record a seated or standing chest-up portrait for best results: mid-chest to just above the head, both shoulders in frame. Slightly wider (waist-up) or slightly closer both work. Do not record head to toe. For full body, use Phoenix-4 Video Requirements.
  • Place the camera at eye level so your face fills at least 25% of the frame
  • Use a stable camera with your face centered and clearly visible
  • Sit at least 3 feet from the camera in a natural, Zoom-style setup: head, shoulders, and upper chest clearly visible
  • Record in 1080p using a desktop app - avoid browser recording and low resolution cameras
  • Keep your neck visible with clear separation between your head and clothing
Glasses, jewelry, headphones, and hair in front of the shoulders are allowed as long as they do not cover the face. Shorter, crisp beards that leave the mouth and teeth visible work best - results may vary with dense or long beards. Realistic, cartoon, anime, and Pixar-style humans are all fine. Animals, mascots, and other non-human figures are not.

Supported Video Formats

Whether recording through the PAL Maker or uploading a pre-recorded training video via the API, ensure your video meets the following requirements:
  • Minimum frame rate: 25 FPS
  • Minimum resolution: 1080p
  • Maximum file size: 750 MB
  • Supported formats: .webm and .mp4 (H.264 video codec + AAC audio codec)

Training Data Policy

All training data uploaded to Tavus must comply with our Terms of Service and Acceptable Use Policy. You are responsible for confirming you have permission to use any submitted content, including visual, audio, and identity-related assets.

Recording Structure

Your video must be one continuous shot, containing 30 seconds of speaking followed by 30 seconds of still footage. You can use a script provided by Tavus or speak on any topic of your choice.
1

Speaking Segment (30 Seconds)

  • Speak naturally on any topic - the content itself does not matter
  • Speak clearly and enunciate well - keeping your teeth visible while talking
  • Keep head and body movement minimal
  • Avoid hand gestures or sudden head turns
Sample script (optional):
2

Still Segment (30 Seconds)

  • Keep your head still and maintain eye contact with the camera
  • Keep lips neutral and closed throughout
  • Do not lick lips or form unusual mouth shapes
  • Avoid any head tilting or movement
Phoenix-4.5 becomes usable in minutes, then continues tuning in the background for a few hours. You can track progress by:
  • Providing a callback_url when creating the Face via API
  • Using the Get Face API
  • Checking the PAL Maker

High-Quality Training Example