model_name: phoenix-4 and train_video_url. The video must be at a publicly accessible URL.
See Choose a Model and Training Path if you are still deciding. Phoenix-4 is the only model that supports full-body framing.
Video training automatically creates a voice from your recording audio and sets it as the face’s
default_voice_id. You can change it later with Update Face.Camera & Framing
- Use a stable, eye-level camera setup with your face centered and clearly visible. Your face must fill at least 25% of the frame.
- Sit at least 3 feet from the camera in a natural, Zoom-style setup - head, shoulders, and upper chest clearly visible. Ensure well-lit space with a simple background
- Record in 1080p using a desktop app - avoid browser recording and low resolution cameras
- Realistic, cartoon, anime, and Pixar-style humans are all fine. Animals, mascots, and other non-human figures are not.
Hair & Clothing Separation
Keep a clear gap between head, neck, and clothing:- Keep your neck and jawline fully visible with clear separation from your clothing
- Avoid high collars or clothing that covers the neck
- Keep hair away from the face and positioned behind the shoulders
- Avoid bangs, loose strands, or complex hairstyles that obscure the face, neck, or shoulder
Supported Video Formats
Whether recording through the PAL Maker or uploading a pre-recorded training video via the API, ensure your video meets the following requirements:- Minimum frame rate: 25 FPS
- Minimum resolution: 1080p
- Maximum file size: 750 MB
- Supported formats:
.webmand.mp4(H.264 video codec + AAC audio codec)
Training Data Policy
All training data uploaded to Tavus must comply with our Terms of Service and Acceptable Use Policy. You are responsible for confirming you have permission to use any submitted content, including visual, audio, and identity-related assets.Recording Structure
Your video must be one continuous shot, containing 30 seconds of speaking followed by 30 seconds of still footage. You can use a script provided by Tavus or speak on any topic of your choice.1
Speaking Segment (30 Seconds)
- Speak naturally on any topic - the content itself does not matter
- Speak clearly and enunciate well - keeping your teeth visible while talking
- Keep head and body movement minimal
- Avoid hand gestures or sudden head turns
2
Still Segment (30 Seconds)
- Keep your head still and maintain eye contact with the camera
- Keep lips neutral and closed throughout
- Do not lick lips or form unusual mouth shapes
- Avoid any head tilting or movement
High-Quality Training Example
Full Body Face
Full body is Phoenix-4 only. Send"model_name": "phoenix-4" with train_video_url.
To create a full body face for conversational video, follow these guidelines.
- Record in vertical format, with the full body visible from head to toe
- Stand still throughout the recording and avoid large movements or hand gestures
- Use consistent lighting with minimal shadows or exposure changes
- 4K resolution is recommended for best quality
All standard recording requirements for Phoenix-4 face training still apply.

