model_name plus either train_video_url or train_image_url (never both). The PAL Maker asks the same questions with live validation.
Step 1. Pick a model
Phoenix-4.5
Newest model. Start using your face in minutes.Choose this if
- You want a zero-shot preview you can use immediately, followed by a tuned version
- You want conventional ‘talking head’ framing: chest-up or waist-up, not head to toe
- You want movement below the neck (torso, shoulders, hair, & clothing) that is more natural
- You only have a photo to work from
- You need a full-body face, head to toe (that is Phoenix-4)
- You need clothing with complex texture to stay exact, like a small logo or lots of decorations (that is Phoenix-4)
- You cannot ship anything with a watermark on it, even briefly
Phoenix-4
The full-body option.Choose this if
- You need full body or a wider body shot
- You need clothing with complex texture to stay as it is, like a small logo or lots of decorations
- You want to start using the face right away (Phoenix-4 is only usable when the whole training job finishes)
- Your photo or video footage has glasses, jewelry, headphones, hair over the shoulders, or a heavy beard (Phoenix-4 rejects those; Phoenix-4.5 does not)
Phoenix-3 (legacy)
Available for edge cases. Almost always use 4.5 or 4 instead.Choose this if
- You have a specific reason to pin this older model. Those cases are rare.
- You want to use the latest, most expressive, and most realistic model
- You want to train from a photo (Phoenix-3 only takes video)
Step 2. Pick an input for that model
Phoenix-4.5
A photo is faster to set up. Video takes more work, and the finished face usually moves more naturally.Start from a photo
Easier. One photo, no recording.
train_image_url. You also attach a voice. See Voices for Image-Based Faces.- Preview usable in about a minute, then tunes in the background
- Chest-up or waist-up, not full body
Start from a video
Better result. Your real movement, and your own voice.
train_video_url- Preview takes a few minutes, then tunes in the background
- Voice is created from the recording
- Chest-up or waist-up, not full body
Phoenix-4
Start from a photo
Upload one photo. Ready when training finishes.
train_image_url. You also attach a voice. See Voices for Image-Based Faces.- No early usable version and no watermark
- Stricter photo checks: hair behind the shoulders, neck visible, sleeves or wide straps, no glasses, no necklace, no large earrings, no over-ear headphones, no heavy beard covering the neck
Start from a video
Upload footage. Highest fidelity today, and the only full-body path.
train_video_url- Thirty seconds speaking, then thirty seconds listening
- Full body: record vertically, head to toe, standing still
- A voice is created from the video’s audio
- Ready when training finishes
Phoenix-3 (legacy)
Start from a photo
Not available on Phoenix-3.Training from a photo was introduced on Phoenix-4. If you want to start from a photo, use Phoenix-4.5 or Phoenix-4.
Start from a video
Video only. Legacy.
model_name: phoenix-3- Ready when training finishes
- Older recording rules apply
- Almost always use Phoenix-4.5 or Phoenix-4 instead

