Skip to main content
Phoenix-3 is legacy. Use it only for edge cases. Almost always train with Phoenix-4.5 or Phoenix-4 instead. Image training is not available on Phoenix-3. Use this path when you call Create Face with model_name: phoenix-3 and train_video_url. The video must be at a publicly accessible URL. See Choose a Model and Training Path if you are still deciding.
Video training automatically creates a voice from your recording audio and sets it as the face’s default_voice_id. You can change it later with Update Face.
Phoenix-3 faces are ready when training finishes.

Recording structure

Phoenix-3 uses a 1 minute speaking + 1 minute listening format, not the Phoenix-4 / Phoenix-4.5 30-second structure.
  • 1 minute of natural speaking
  • 1 minute of still / listening footage
  • One continuous shot

Camera & Framing

  • Place the camera at eye level so your face fills at least 25% of the frame
  • Keep your face centered and clearly visible
  • Realistic, cartoon, anime, and Pixar-style humans are all fine. Animals, mascots, and other non-human figures are not
  • Well-lit space with a simple background
  • Record in 1080p using a desktop app - avoid browser recording and low resolution cameras

Supported Video Formats

  • Minimum frame rate: 25 FPS
  • Minimum resolution: 1080p
  • Maximum file size: 750 MB
  • Supported formats: .webm and .mp4 (H.264 video codec + AAC audio codec)

Training Data Policy

All training data uploaded to Tavus must comply with our Terms of Service and Acceptable Use Policy. You are responsible for confirming you have permission to use any submitted content, including visual, audio, and identity-related assets.