Bringing a container

Deploy your own Dockerfile and HTTP server, and declare the endpoints the builder cannot see inside your image.

Introduction

Bring a container when the environment matters as much as the model: a CUDA build you have already fought with, a pipeline with system packages in it, or a server that works and that you would rather not rewrite.

You provide a Dockerfile, a container.yaml beside it, and whatever files the Dockerfile copies. Everything after the build is identical to a code app: the same versions, the same scaling, the same invoke routes.

Your server is the app

The image runs your CMD and your server answers requests, exactly as it does on your machine.

FROM python:3.12-slim

WORKDIR /app
COPY app.py /app/app.py

ENV PORT=8080
EXPOSE 8080
CMD ["python", "/app/app.py"]

What the platform requires of it is short. Bind the port you declare, answer POST /{endpointPath} for every endpoint you declare, and serve the two probe routes.

The platform's own containers reserve ports 8000, 8888, 8889 and 9090. Binding one of those collides with something already running beside you, so pick anything else.

Declaring what the builder cannot see

A code app has its endpoints read out of the source. A builder cannot look inside your image and find them, so a container app declares them itself.

configVersion: 1
port: 8080
probes:
  readiness: /ready
  liveness: /health
timeouts:
  requestSeconds: 60
  startupSeconds: 300
endpoints:
  - path: generate
    input:
      type: object
      additionalProperties: false
      required: [prompt]
      properties:
        prompt: { type: string }
        steps: { type: integer }
  - path: upscale

port is where your server listens. probes are the routes the platform calls to decide whether a worker is ready for traffic and whether it is still alive. timeouts bound one request and the startup that precedes it.

Only port and endpoints are required. The example above spells probes and timeouts out, but every one of those values is the default you get by omitting them.

SettingDefaultMost you can ask for
probes.readiness/readyAny absolute HTTP path
probes.liveness/healthAny absolute HTTP path
timeouts.requestSeconds60900
timeouts.startupSeconds3002400
endpointsnone, you declare them20 per app
Schema sizenone200 properties across every input and output together

The platform waits requestSeconds plus 30 seconds for each answer. A request still running after that fails its task with inference timed out after N seconds, and the late answer is discarded.

Schemas you write rather than derive

This is the real difference between the two paths. A code app derives its request schema from your handler signature. A container app declares it, in JSON Schema, per endpoint.

An endpoint with an input schema gets the same treatment as a code app: a mismatched body is rejected before it reaches your container. An endpoint with none, like upscale above, is a passthrough and takes whatever JSON you send it.

You can declare an output schema too, which documents the response rather than gating it.

Both paths converge here. The builder turns your container.yaml into the same manifest a code build produces, so the queue manager routes, validates and generates documentation without knowing which kind of app it is looking at.

Readiness is your cold start

A worker takes no traffic until its readiness route answers. That is where you load your model, and returning 503 until you are done is the whole mechanism.

def do_GET(self):
    if self.path == "/ready":
        self.send_response(200 if model_loaded else 503)
        self.end_headers()

Getting this wrong is the most common container failure. A server that reports ready before its weights are loaded takes a request it cannot serve, and the caller sees a timeout rather than a cold start.

Testing before you deploy

The container path has an advantage the code path does not: it is just your image, so you can run the whole thing locally with Docker before Runware ever sees it.

docker build -t my-app .
docker run --rm -p 8080:8080 my-app

curl http://localhost:8080/ready
curl -X POST http://localhost:8080/generate \
  -H 'Content-Type: application/json' \
  -d '{"prompt": "a red bicycle"}'

Local testing leaves one gap, the envelope. Runware wraps your payload as {taskId, payload} and hands your server the payload member alone, so a local curl sends the inner object directly. See Invoking an app.

What you give up

Endpoint hardware is still app-level. Every endpoint your container declares shares one queue and one worker pool, and a hardware or scaling key inside an endpoint entry is a 422. Two endpoints that genuinely need different GPUs are two apps.

And the schema is yours to maintain. A code app's schema is generated from its handler, so the two always agree. A container.yaml can drift, and nothing will tell you until a request is rejected that should have worked.