Skip to main content
Create a face with two choices, in order: which Phoenix model, then whether you start from a photo or a video. Use the Create Face API with model_name plus either train_video_url or train_image_url (never both). The PAL Maker asks the same questions with live validation.

Step 1. Pick a model

Phoenix-4.5

Newest model. Start using your face in minutes.Choose this if
  • You want a zero-shot preview you can use immediately, followed by a tuned version
  • You want conventional ‘talking head’ framing: chest-up or waist-up, not head to toe
  • You want movement below the neck (torso, shoulders, hair, & clothing) that is more natural
  • You only have a photo to work from
Do not choose this if
  • You need a full-body face, head to toe (that is Phoenix-4)
  • You need clothing with complex texture to stay exact, like a small logo or lots of decorations (that is Phoenix-4)
  • You cannot ship anything with a watermark on it, even briefly

Phoenix-4

The full-body option.Choose this if
  • You need full body or a wider body shot
  • You need clothing with complex texture to stay as it is, like a small logo or lots of decorations
Do not choose this if
  • You want to start using the face right away (Phoenix-4 is only usable when the whole training job finishes)
  • Your photo or video footage has glasses, jewelry, headphones, hair over the shoulders, or a heavy beard (Phoenix-4 rejects those; Phoenix-4.5 does not)

Phoenix-3 (legacy)

Available for edge cases. Almost always use 4.5 or 4 instead.Choose this if
  • You have a specific reason to pin this older model. Those cases are rare.
Do not choose this if
  • You want to use the latest, most expressive, and most realistic model
  • You want to train from a photo (Phoenix-3 only takes video)
All models accept a human or clearly human-shaped subject: realistic, cartoon, anime, and Pixar-style faces are fine. Animals, mascots, and other non-human figures are not.

Step 2. Pick an input for that model

Phoenix-4.5

A photo is faster to set up. Video takes more work, and the finished face usually moves more naturally.

Start from a photo

Easier. One photo, no recording.train_image_url. You also attach a voice. See Voices for Image-Based Faces.
  • Preview usable in about a minute, then tunes in the background
  • Chest-up or waist-up, not full body

Start from a video

Better result. Your real movement, and your own voice.train_video_url
  • Preview takes a few minutes, then tunes in the background
  • Voice is created from the recording
  • Chest-up or waist-up, not full body

Phoenix-4

Start from a photo

Upload one photo. Ready when training finishes.train_image_url. You also attach a voice. See Voices for Image-Based Faces.
  • No early usable version and no watermark
  • Stricter photo checks: hair behind the shoulders, neck visible, sleeves or wide straps, no glasses, no necklace, no large earrings, no over-ear headphones, no heavy beard covering the neck

Start from a video

Upload footage. Highest fidelity today, and the only full-body path.train_video_url
  • Thirty seconds speaking, then thirty seconds listening
  • Full body: record vertically, head to toe, standing still
  • A voice is created from the video’s audio
  • Ready when training finishes

Phoenix-3 (legacy)

Start from a photo

Not available on Phoenix-3.Training from a photo was introduced on Phoenix-4. If you want to start from a photo, use Phoenix-4.5 or Phoenix-4.

Start from a video

Video only. Legacy.model_name: phoenix-3
  • Ready when training finishes
  • Older recording rules apply
  • Almost always use Phoenix-4.5 or Phoenix-4 instead