Skip to content

FastGen Cloud vs Replicate

Replicate is a general marketplace for running community models — enormous breadth, and you pay for the compute time each run consumes. FastGen Cloud is the opposite trade: a small curated set of image and video models at a fixed, published price per output.

Side by side

DimensionFastGen CloudReplicate
Model catalogueA curated set of image and video models, each behind a stable model id.A very large open catalogue of community-published models, plus your own.
Custom modelsNot supported — the API never accepts checkpoints, LoRA files or workflow JSON. Custom subjects go through the characters training API instead.Supported — you can package and push your own model and run it on their hardware.
Unit of billingPer image, or per second of video output. The cost of a call is knowable before you make it.Predominantly per second of compute on the hardware tier the model runs on.
Payment modelPrepaid balance, debited at submit, auto-refunded when a job fails.Postpaid — usage accrues against a card on file.
Cold startsSame trade-off: serverless GPUs scale to zero, so an idle model pays a load cost on the first call.Same trade-off, with the option to pay for always-warm dedicated capacity.
Job lifecycle202 with a job id, then poll or receive a signed webhook. Idempotency-Key replays instead of double-charging.Prediction object you poll or receive by webhook.
Prompt handlingPrompts and negatives reach the model verbatim — nothing injected, no rating tags added.Depends entirely on the individual model you pick.

Comparison of platform design, not of prices — Replicate sets its own rates and changes them independently, so check its current pricing page before deciding. Last reviewed 20 August 2026.

When Replicate is the better choice

Be honest about this: if you need a model that is not in a curated catalogue, or you have your own fine-tune you want to host, a marketplace is the right tool and this platform cannot help you.

  • You need a specific research model or a niche architecture
  • You want to deploy your own packaged model and call it over an API
  • Your workload is bursty and experimental across many different models

When a fixed-price API is the better choice

Per-second compute billing means the price of a generation depends on how long the GPU happened to take. That is fine for experiments and awkward once you are reselling generations to end users, because your unit economics move underneath you.

  • You are charging your own users per image and need a stable cost per unit
  • You want a hard ceiling on spend — a prepaid balance cannot be overrun
  • You want one contract across text-to-image, editing, identity and video rather than seven model-specific schemas
  • You need generation without a platform rewriting your prompts

What migrating actually involves

The shape is the same everywhere: authenticate, submit a job, wait for a callback, download the output. In practice a migration is one client module and a model-name mapping.

submit + wait
const res = await fetch("https://api.fastgencloud.com/v1/images/generations", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.FASTGEN_API_KEY}`,
    "Content-Type": "application/json",
    "Idempotency-Key": requestId,
  },
  body: JSON.stringify({ model: "general-image", prompt, size: "1024x1024" }),
});
const job = await res.json();          // 202 → { id, status: "queued" }

// then either poll GET /v1/jobs/{id}, or pass webhook_url and get a signed callback

How billing works here

Money is a prepaid balance in your account. You top it up with a card, each job debits it at submit time, and a failed job is refunded automatically — there is no invoice at the end of the month and no way to run up a bill you did not intend.

Rates are per image or per second of video, published on the pricing page. Images above 1 megapixel bill pro rata by pixel count. Outputs are retained for up to 7 days.

Frequently asked questions

Can I run my own fine-tuned model here?
No. The API deliberately never accepts checkpoint filenames, LoRA files or raw workflow JSON — every model maps through a server-side allowlist. If your goal is a custom subject rather than a custom architecture, the characters API trains one from photos and gives you an id to reference.
Is the underlying hardware different?
Both run serverless GPU workers, and both therefore have cold starts when a model has been idle. The difference is what you are billed for: output here, compute time there.
How hard is the migration?
One client module and a model-name mapping in most codebases. Both are async job APIs with webhooks, so your existing "submit, wait, download" flow carries over unchanged.

Related

Start generating

Create an account, add a prepaid balance, and call the API with a key from the console. No subscription, no minimum, no per-seat pricing.

Get an API key