FAQ

The questions that come up before the first deployment and just after it.

When should I use this instead of the existing API?

When the model is yours. A fine-tune, a proprietary pipeline, an architecture nobody hosts, or a chain of steps that has to happen on one machine next to the weights.

If the model already exists on Runware, calling it by identifier is simpler, cheaper and needs none of this.

What do I pay for when nothing is running?

Nothing, if minWorkers is 0. An app with no workers costs no GPU time, and the first request after a quiet period starts one.

Holding minWorkers above zero is how you buy away the cold start, and you pay for that capacity whether or not anything arrives.

How long is a cold start?

It depends on your image and your weights, and the worker states tell you which one is costing you. A worker in pulling is fetching the image. A worker in loading is running your load.

The lever that matters most is putting weights on a volume, because an unmounted download is fetched again on every cold start.

Can one app expose several endpoints?

Yes. Decorate one method per endpoint and they all serve from the same app.

@serve
class ImageTools:
    @endpoint
    def generate(self, prompt: str) -> dict: ...

    @endpoint
    def upscale(self, image: str) -> dict: ...

Can I type my requests instead of taking loose fields?

Yes. Annotate a parameter with a Pydantic model and it arrives as an instance, validated before your handler runs, with the model's schema published as that endpoint's contract.

class Transcribe(BaseModel):
    audio_url: str
    language: str | None = None

@endpoint
def transcribe(self, req: Transcribe) -> dict: ...

See Writing a model.

Can different endpoints use different GPUs?

No. Every endpoint on an app shares one queue and one worker pool, and hardware is set once on the app.

The API enforces it: the decorator takes no arguments, the create and update routes reject an endpoints field, and a container config carrying a hardware key is a 422.

If two endpoints genuinely need different hardware, they are two apps.

Can I use more than one GPU per worker?

Yes. gpusPerWorker takes 1, 2, 4 or 8, because a worker holds its GPUs as one group the cluster grants indivisibly.

The size has to be one the cluster behind your GPU type grants, and a size it does not answers 422. Raising the number on an existing app also means redeploying it, because an image built for a single rank holds every GPU it was granted and uses one.

Can I test before deploying?

A container app, yes. It is your image and your server, so docker run exercises the whole thing. Local runs miss only the envelope, because Runware hands your server the payload member rather than the whole body.

A code app, partly. serve hands your class straight back, so your own tests can import it and call the methods directly. What that does not exercise is the platform: the schema derived from your signatures, the dispatch, or the cold start.

Nothing checks your file before you upload it, either. A duplicate path, a parameter that cannot be filled by name or a method name that is not a legal path surface as a failed build, not as an error on your machine.

What happens when my handler raises?

The task reaches failed and carries the reason in error as a plain string, with no code and no structured cause. Read status to detect it.

That string is what the task carries. Make your exceptions say something useful, and read apps logs for what the worker printed around it.

The app's failed-request listing shows that a call failed and when, filterable by 4xx against 5xx, which is the fastest way to tell a caller sending the wrong shape from your handler breaking.

How do I roll back?

Activate an earlier version. It is the same operation as going forward, nothing is rebuilt, and the image that version pinned is re-applied.

runware serverless apps versions activate my-app 7

Can I pin a version when I invoke?

No. The invoke routes carry no version segment. Which version serves is app state, set by deploying one.

When does an environment variable reach a worker?

Setting one records a new version and rolls the workload, so it arrives on the workers that rollout starts rather than on the ones already running. Writing the value it already had records nothing and rolls nothing.

An app that is stopped pins the version instead and picks it up on the next resume.

Can I pull from a private registry?

No. A container deploy submits a Dockerfile and its build context, and the platform builds the image from it.

Which Python version and base image?

The builder takes a baseImage, defaulting to a slim Python image, and installs the requirements you list on top. Pin the base image if your model needs a specific interpreter or CUDA build.

Leave the platform's own client libraries out of your requirements. They are installed into the image at build time and again when the container starts, so a version pinned by you would drift from the services it talks to.

Can I change a volume later?

No. Volumes are fixed when the app is created, and the update route rejects them. Changing your mount layout means a new app, so decide the paths before the first deploy.

Does changing the scaling cost me a build?

No. A configuration change records a version and carries the previous image forward, so it takes seconds rather than minutes. It also means your scaling survives a code deploy, because those values live on the app.