Skip to content

Consistent characters across generated images

Diffusion models have no memory between calls, so the same prompt gives you a different person every time. There are two ways to fix that, and they suit different products. This guide covers both, and when to reach for each.

Face ID
1 photo

Instant, approximate likeness

Training
4–30 photos

~10 minutes, tight likeness

Reuse
By id

Weights never leave the server

Why seeds are not the answer

The usual first attempt is to fix the seed. That reproduces an identical image for an identical prompt, but change one word — the setting, the outfit, the pose — and the face drifts, because the seed controls the noise, not the identity.

What you actually need is to condition generation on a specific face. Two mechanisms do that: an adapter that reads a face embedding from a reference photo, or a small set of trained weights that teach the model the person.

Option 1: Face ID, from a single photo

Fastest path. Post one clear photo alongside your prompt and get back the same person in that scene. Nothing to train, nothing to store, and it works the first time a user uploads a selfie.

The reference drives the face only — describe the rest of the person in the prompt, because hair colour, build and clothing come from your text, not the photo.

curl
curl https://api.fastgencloud.com/v1/images/face-id \
  -H "Authorization: Bearer $FASTGEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
      "image": "https://example.com/selfie.jpg",
      "prompt": "a woman with dark curly hair in a green raincoat, standing on a harbour wall, overcast light",
      "style": "realistic",
      "identity_strength": 0.9,
      "n": 4
    }'

Option 2: train a character

When the same person recurs — a brand mascot, a user avatar, a series — training gives a much closer likeness that holds up across wildly different scenes. Send 4–30 photos, wait about ten minutes, get an id.

train, then poll until ready
# 1. start training
curl https://api.fastgencloud.com/v1/characters \
  -H "Authorization: Bearer $FASTGEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
      "name": "Ada",
      "class": "woman",
      "base_model": "general-image",
      "images": ["https://…/1.jpg", "https://…/2.jpg", "https://…/3.jpg", "https://…/4.jpg"]
    }'
# → { "id": "chr_0123456789ab", "status": "training", "progress": { … } }

# 2. poll until status is "ready"
curl https://api.fastgencloud.com/v1/characters/chr_0123456789ab \
  -H "Authorization: Bearer $FASTGEN_API_KEY"

Generating with a trained character

Reference the character by id on an ordinary generation call. The server resolves the weights and the trigger token — you never see either.

curl
curl https://api.fastgencloud.com/v1/images/generations \
  -H "Authorization: Bearer $FASTGEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
      "model": "general-image",
      "prompt": "a woman in a linen shirt reading at a cafe table, soft afternoon light",
      "character": { "id": "chr_0123456789ab", "strength": 1.0 },
      "size": "832x1216",
      "n": 2
    }'

Which one to build on

A pattern worth stealing: offer Face ID immediately on upload so the user sees a result within seconds, and kick off training in the background. Swap to the trained character once it is ready.

User uploads a selfie, wants 4 images now
Face ID
Recurring character in an ongoing product
Training
Likeness must be convincing to someone who knows the person
Training
You cannot store biometric-derived assets long term
Face ID
Anime or cartoon output
Either — train against that family, or set style

Getting a better likeness

  • Vary the training set — angle, distance, lighting, expression. Ten varied photos beat thirty near-identical ones
  • One face per photo, no sunglasses, no heavy filters, no motion blur
  • On Face ID, raise identity_strength toward 1.1 for closer resemblance, lower toward 0.7 if faces look stiff or waxy
  • Generate several and pick — identity pipelines vary more between seeds than plain text-to-image does
  • Likeness is always strongest on the realistic style; anime and cartoon carry a person more loosely by design

Both features operate on real people’s faces. Generating sexual or otherwise harmful depictions of a real, identifiable person without their consent is prohibited by the Acceptable Use Policy and will end an account.

If you are building a consumer product, get explicit consent at upload, and be clear with users about how long you keep their photos. Outputs here are deleted after 7 days; your own storage is your responsibility.

Frequently asked questions

How many photos do I need to train a character?
Four is the minimum and thirty is the maximum. Ten to twenty varied photos is the sweet spot — beyond that you get diminishing returns unless the extra photos add genuinely new angles or lighting.
Can I use a trained character on any model?
On any model in the family it was trained against. Train against general-image for photoreal, or against anime or cartoon for those styles.
Can I download the trained weights?
No. Characters are referenced by id and the weights stay server-side. The public API never accepts or returns LoRA files.
What does it cost?
Face ID is billed per image like any generation. Training is billed once per run, and storing characters uses slots — two free per account, with monthly packages beyond that. See the pricing page for current rates.

Related

Start generating

Create an account, add a prepaid balance, and call the API with a key from the console. No subscription, no minimum, no per-seat pricing.

Get an API key