FastGen Cloud vs Replicate
Replicate is a general marketplace for running community models — enormous breadth, and you pay for the compute time each run consumes. FastGen Cloud is the opposite trade: a small curated set of image and video models at a fixed, published price per output.
Side by side
| Dimension | FastGen Cloud | Replicate |
|---|---|---|
| Model catalogue | A curated set of image and video models, each behind a stable model id. | A very large open catalogue of community-published models, plus your own. |
| Custom models | Not supported — the API never accepts checkpoints, LoRA files or workflow JSON. Custom subjects go through the characters training API instead. | Supported — you can package and push your own model and run it on their hardware. |
| Unit of billing | Per image, or per second of video output. The cost of a call is knowable before you make it. | Predominantly per second of compute on the hardware tier the model runs on. |
| Payment model | Prepaid balance, debited at submit, auto-refunded when a job fails. | Postpaid — usage accrues against a card on file. |
| Cold starts | Same trade-off: serverless GPUs scale to zero, so an idle model pays a load cost on the first call. | Same trade-off, with the option to pay for always-warm dedicated capacity. |
| Job lifecycle | 202 with a job id, then poll or receive a signed webhook. Idempotency-Key replays instead of double-charging. | Prediction object you poll or receive by webhook. |
| Prompt handling | Prompts and negatives reach the model verbatim — nothing injected, no rating tags added. | Depends entirely on the individual model you pick. |
Comparison of platform design, not of prices — Replicate sets its own rates and changes them independently, so check its current pricing page before deciding. Last reviewed 20 August 2026.
When Replicate is the better choice
Be honest about this: if you need a model that is not in a curated catalogue, or you have your own fine-tune you want to host, a marketplace is the right tool and this platform cannot help you.
- You need a specific research model or a niche architecture
- You want to deploy your own packaged model and call it over an API
- Your workload is bursty and experimental across many different models
When a fixed-price API is the better choice
Per-second compute billing means the price of a generation depends on how long the GPU happened to take. That is fine for experiments and awkward once you are reselling generations to end users, because your unit economics move underneath you.
- You are charging your own users per image and need a stable cost per unit
- You want a hard ceiling on spend — a prepaid balance cannot be overrun
- You want one contract across text-to-image, editing, identity and video rather than seven model-specific schemas
- You need generation without a platform rewriting your prompts
What migrating actually involves
The shape is the same everywhere: authenticate, submit a job, wait for a callback, download the output. In practice a migration is one client module and a model-name mapping.
const res = await fetch("https://api.fastgencloud.com/v1/images/generations", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.FASTGEN_API_KEY}`,
"Content-Type": "application/json",
"Idempotency-Key": requestId,
},
body: JSON.stringify({ model: "general-image", prompt, size: "1024x1024" }),
});
const job = await res.json(); // 202 → { id, status: "queued" }
// then either poll GET /v1/jobs/{id}, or pass webhook_url and get a signed callbackHow billing works here
Money is a prepaid balance in your account. You top it up with a card, each job debits it at submit time, and a failed job is refunded automatically — there is no invoice at the end of the month and no way to run up a bill you did not intend.
Rates are per image or per second of video, published on the pricing page. Images above 1 megapixel bill pro rata by pixel count. Outputs are retained for up to 7 days.
Frequently asked questions
- Can I run my own fine-tuned model here?
- No. The API deliberately never accepts checkpoint filenames, LoRA files or raw workflow JSON — every model maps through a server-side allowlist. If your goal is a custom subject rather than a custom architecture, the characters API trains one from photos and gives you an id to reference.
- Is the underlying hardware different?
- Both run serverless GPU workers, and both therefore have cold starts when a model has been idle. The difference is what you are billed for: output here, compute time there.
- How hard is the migration?
- One client module and a model-name mapping in most codebases. Both are async job APIs with webhooks, so your existing "submit, wait, download" flow carries over unchanged.
Related
Start generating
Create an account, add a prepaid balance, and call the API with a key from the console. No subscription, no minimum, no per-seat pricing.
Get an API key