Skip to content

Moderating user-generated AI images

A generation API that passes your prompts through verbatim hands you both the capability and the responsibility. This guide is the other half of that trade: where to put your own gate, what each position can and cannot catch, and what to keep for the day someone files a report.

Gate positions
3

Prompt, output, publish

Cheapest check
Prompt

Before you spend a credit

Only reliable one
Output

Prompts lie, pixels do not

Decide what you are actually moderating for

Teams often start by shopping for a classifier and never write down the policy it is supposed to enforce. Do it the other way around. Your policy is driven by four things, and they rarely point the same way.

  • The law where your users are — which varies enormously and changes faster than your roadmap.
  • Your app store, if you have one. Mobile store rules are usually stricter than the law and are enforced by removal.
  • Your payment processor. This is the one teams discover last and it is frequently the binding constraint — read your processor’s prohibited-business list before you build, not after your first hold.
  • Your own brand. Perfectly legal output can still be something you do not want on your front page.

Three places to put a gate

Each position catches a different failure and costs a different amount. Most products end up with two of the three.

Before submit (prompt)
Cheapest — rejects before you spend a credit. Misses anything phrased innocuously.
On output (image)
The only check that sees what was actually produced. Costs one classifier call per image.
Before publish (share, feed, export)
Catches what the first two missed and scales with exposure, not generation.

Gate at the webhook, not in the request path

Generation is asynchronous here: you submit, you get a job id, and a signed webhook fires when the output is ready. That callback is the natural place to run your classifier, because it is off the user’s critical path and you are holding the URL before anyone else has seen it.

Verify the signature first — an unauthenticated handler that trusts its body is a way to get arbitrary URLs into your feed.

app/api/hooks/fastgen/route.ts
import { verifyFastgenSignature } from "@/lib/fastgen";

export async function POST(req: Request) {
  const raw = await req.text();
  if (!verifyFastgenSignature(raw, req.headers)) {
    return new Response("bad signature", { status: 401 });
  }

  const event = JSON.parse(raw);
  if (event.status !== "succeeded") return new Response("ok");

  // Your policy, your classifier. Hold the asset until it returns.
  const verdict = await classify(event.output[0].url);

  await db.generation.update({
    where: { jobId: event.id },
    data: {
      state: verdict.allowed ? "ready" : "withheld",
      labels: verdict.labels,
      reviewedAt: new Date(),
    },
  });

  return new Response("ok");
}

Copy anything you intend to keep

Output URLs expire after 7 days. That is a retention feature, not a CDN — if your product shows a user their history, copy the bytes into your own storage inside that window. Do the copy after your classifier passes, so a withheld asset never lands in your bucket in the first place.

Keep records you can actually answer a report with

The day someone reports an image, you need to answer three questions quickly: who generated it, what did they send, and what did you do about it. Store enough to answer them and no more than you are willing to be responsible for.

  • The job id and your internal user id, so a reported asset resolves to an account.
  • The prompt as submitted, and the classifier verdict with its labels and timestamp.
  • The action you took and who took it — withheld automatically, or cleared by a named reviewer.
  • A takedown path a stranger can find and use without an account. A contact address on a public page is the minimum.
  • A retention window for all of the above, written down and actually enforced by a job.

What the platform does and does not do for you

Worth being unambiguous, because assuming a safety net that is not there is the expensive mistake. The platform blocks one thing at the prompt and passes everything else through, so any policy narrower than the legal floor is yours to implement.

Prompt filtering for adult content
Not performed. Yours.
Output nudity or violence detection
Not performed. Yours.
Minor-safety prompt gate
Enforced by the platform, always on.
Age verification of your end users
Yours.
Takedown handling for your users’ content
Yours.
Output retention
Deleted after 7 days by the platform.

Frequently asked questions

Does the API moderate output for me?
No. Generated images and video frames are returned without being classified. If your product needs nudity, violence or celebrity-likeness detection, you run it — most teams call a third-party vision classifier from their webhook handler.
Where should I put the gate if I can only afford one?
On the output, at the webhook. A prompt filter is cheaper but a determined user routes around it with innocuous phrasing, and the thing you are actually accountable for is the image, not the sentence.
Can I get the moderation verdict from the job record?
The job carries a moderation object reflecting the platform’s own prompt gate. It is not a general content classification of your image and should not be used as one.
What if a user generates something illegal?
Preserve the record, remove the asset from your product, and handle it under your own terms of service and your legal obligations in the relevant jurisdiction. Report abuse of the API to the contact address on the site so it can be investigated on this side too.

Related

Start generating

Create an account, add a prepaid balance, and call the API with a key from the console. No subscription, no minimum, no per-seat pricing.

Get an API key