Volumes

Keep model weights out of the checkpointed filesystem so a cold start does not fetch them again.

Introduction

A volume is a directory that survives. Everything else your app writes is part of the worker's checkpointed state, and that is a more expensive place to put a large file than it looks.

runware serverless deploy ./model.py --id my-app --gpu-type l40s \
  --volume /root/.cache/huggingface

Why this exists

Your app runs in a sandbox whose filesystem is part of what gets checkpointed. That is what lets a worker resume from a snapshot, and it is why the platform starts quickly.

It also means an unmounted download is charged twice. It is copied into every checkpoint, making each one larger and slower, and it is fetched again on every cold start that begins from scratch.

A volume takes the directory out of both. The weights live on node-local storage beside the worker, outside the checkpointed root, and a worker that finds them already there skips the download entirely.

This is the single biggest lever on cold-start time for a model with large weights. Everything else you tune is smaller than skipping a fifteen-gigabyte download.

The mount path is the whole thing

A volume has no name. Its mount path is its identity, both to your code and to the node-local directory it is keyed by.

{
  "volumes": [
    { "mountPath": "/root/.cache/huggingface" },
    { "mountPath": "/models" }
  ]
}

Nothing to create beforehand, nothing to look up, and nothing to reference by handle. You name the path your code already opens and the platform mounts it there.

What the platform will reject

  • Absolute paths only, matching ^/[A-Za-z0-9._/+@-]+$, up to 2048 characters. Each component between the slashes is at most 255 bytes, which is the filesystem limit the mirrored directory has to satisfy.
  • Mounting / is rejected. A volume is a directory inside the tree.
  • /data/liquid-components is reserved by the platform. That path, anything under it and anything containing it are refused. The refusal comes when the app rolls out, so a create can be accepted and the deploy then fail.
  • No duplicates and no overlaps. /models and /models/flux together are a mistake rather than a merge, and the platform treats them as one.
  • At most 30 per app.

They are fixed when the app is created

This is the constraint that catches people. --volume is a create-time flag. The update route accepts a name, a configuration, a source and secrets, and it does not accept volumes.

An app's volume set is fixed at creation, so there is no way to add, remove or move one afterwards. Passing --volume to a deploy on a live app is an error, not an update. Changing your mount layout means creating a new app.

The set is also frozen into each version, so a rollback restores the volume layout that version shipped with.

Decide the paths before the first deploy. The cost of getting it wrong lands after a build you already waited for.

What belongs on one

Anything your app downloads or generates at runtime and would rather do once. A Hugging Face cache, weights pulled on first load, a compiled kernel cache, an index built at startup.

Anything you can bake into the image at build time should stay there. A file that is already in the layer is already on the node, and mounting it over adds a moving part for nothing.

@serve
class FluxDev:
    def load(self) -> None:
        # HF_HOME points at the mounted volume, so the second worker
        # on this node finds the weights already on disk
        self.pipe = FluxPipeline.from_pretrained("black-forest-labs/FLUX.1-dev")

Point your library's cache at the mount path with an environment variable and it will do the rest. Most model libraries already respect one.