# Modal

> Use when running Python or GPU workloads serverlessly on Modal — modal.App, inline container Images, gpu= on @app.function, Volumes for weight caching, Cron schedules, ASGI endpoints, modal run vs serve vs deploy. NOT managed prediction APIs with no container of your own (that is replicate); NOT SSH-able GPU boxes rented by the hour (that is runpod).

- Skill: `ericrisco/modal` (Agent Skill, multi-file: 6 files)
- Install (CLI): `npx skillmds@latest add ericrisco/modal`
- Raw SKILL.md: https://api.skillmd.com/api/skills/ericrisco/modal/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: DevOps & Infra
- Author: ericrisco (https://skillmd.com/u/ericrisco)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/ericrisco/modal

---


# Modal — serverless Python & GPU as decorators

Modal runs your Python on remote containers without you ever writing a Dockerfile or a YAML
file. The mental model: **infrastructure is declared inline as Python decorators.** A
`modal.App` is the deployable unit; each `@app.function` runs in its own container built from
a `modal.Image` you describe in code; you attach a GPU, a Volume, or a Secret as keyword
arguments and the platform provisions, scales to zero, and tears down for you. There is no
control plane to babysit — the source file *is* the infra.

Pinned stack: **modal 1.4.3** (released 2026-05-18), Python **3.10–3.14** (`>=3.10,<3.15`).
Install with `pip install modal` then `modal setup` to authenticate. Everything below uses
the Modal 1.0+ API; several pre-1.0 forms were removed and are called out as Bad→Good.

## Not this skill

Modal owns the serverless-container-as-decorators surface and its CLI lifecycle; the *contents*
of your function belong elsewhere.

| The job | Goes to |
|---|---|
| Calling a **managed prediction API** with no container of your own | [`replicate`](../replicate/SKILL.md) / [`together-fireworks`](../together-fireworks/SKILL.md) / [`fal`](../fal/SKILL.md) |
| Renting a **persistent, SSH-able GPU box** by the hour/week | [`runpod`](../runpod/SKILL.md) |
| **FastAPI design** (routing, Pydantic, deps) independent of host | [`fastapi`](../fastapi/SKILL.md) |
| Writing a **Dockerfile** for a registry / k8s / Compose | [`docker`](../docker/SKILL.md) |
| General **Python language/runtime** questions | [`python`](../python/SKILL.md) |
| **RAG / LLM pipeline** orchestration logic itself | [`llm-pipeline`](../llm-pipeline/SKILL.md) |

## Decision: which entrypoint?

| You want… | Use | Persists after exit? |
|---|---|---|
| Run a function once and exit (script, batch) | `modal run app.py` + `@app.local_entrypoint()` | No (ephemeral) |
| Hot-reload dev loop for a web endpoint | `modal serve app.py` | No (dies on Ctrl-C) |
| A persistent named deployment (prod, schedules, endpoints) | `modal deploy app.py` | Yes |
| Fan out work across many containers | `.map()` / `.starmap()` / `.spawn()` inside an entrypoint | n/a |

Rule: **schedules and live web endpoints require `modal deploy`.** `modal run` exits when the
entrypoint returns, so a Cron defined under `modal run` never fires. `modal serve` is for the
dev loop only — it watches your files and redeploys on save, but the app vanishes when you
stop it.

## The minimal app skeleton

```python
import modal

app = modal.App("hello-modal")

# The image is the container spec. Build it once, reuse across functions.
image = modal.Image.debian_slim(python_version="3.12").uv_pip_install("requests")


@app.function(image=image)
def fetch(url: str) -> int:
    import requests  # imported INSIDE the function: it lives in the remote image, not locally

    return len(requests.get(url).content)


@app.local_entrypoint()
def main() -> None:
    # Runs on your laptop; .remote() ships the call to a Modal container.
    print(fetch.remote("https://modal.com"))
```

Run it: `modal run app.py`. **Bad** = wiring infra with argparse + a bash launcher + a
hand-rolled Dockerfile. **Good** = the decorators above; the app, image, and scaling are all
declared in the one file. Note the in-function import: dependencies you `uv_pip_install` exist
in the *remote* image, so import them inside the function (or guard top-level imports), not at
module top where your laptop would need them too.

## Images — pin, layer, cache

Build images by chaining methods on `modal.Image`. Rules, each with its why:

1. **Prefer `.uv_pip_install(...)` over `.pip_install(...)`** — it resolves and installs with
   `uv`, materially faster image builds.
2. **Pin versions** — `.uv_pip_install("torch==2.5.1", "transformers==4.46.0")`. Unpinned
   deps make builds non-reproducible and silently drift on rebuild.
3. **Order layers stable→volatile** — system packages and big wheels first, your fast-changing
   code last. Modal caches each layer; a change busts that layer and everything after it.
4. **Add your own code with `.add_local_dir(...)` / `.add_local_python_source(...)`**, not by
   pip-installing your repo. These are applied last so editing your source doesn't rebuild torch.
5. **`.from_registry("...")`** when you need a specific base image; **`.apt_install("ffmpeg")`**
   for system binaries; **`.run_commands(...)`** for arbitrary build steps.

```python
image = (
    modal.Image.debian_slim(python_version="3.12")
    .apt_install("ffmpeg")                                   # stable: rarely changes
    .uv_pip_install("torch==2.5.1", "transformers==4.46.0")  # heavy wheels, pinned
    .add_local_python_source("my_pkg")                       # volatile: your code, applied last
)
```

→ [`references/images-gpu-cookbook.md`](references/images-gpu-cookbook.md) for vLLM / torch+CUDA
/ diffusers recipes and the download-once weight-cache pattern.

## GPU — it's a string now

In Modal 1.0+ the GPU is a **string** on the decorator. The old `modal.gpu.H100()` objects
were removed.

- Single GPU: `gpu="H100"`.
- Count via colon: `gpu="A100:2"` (two A100s in one container).
- Memory variant: `gpu="A100-80GB"` (also `A100-40GB`).
- Fallback list (first available wins): `gpu=["H100", "A100", "any"]`.
- Supported types: `T4, L4, A10, L40S, A100(-40GB/-80GB), RTX-PRO-6000, H100, H200, B200`.

```python
# Bad — removed API, raises at import.
# @app.function(gpu=modal.gpu.A100())

# Good — string form.
@app.function(image=image, gpu="A100-80GB", timeout=600)
def embed(texts: list[str]) -> list[list[float]]: ...
```

Pick the smallest GPU that fits: **T4/L4** for cheap inference and small models, **A10/L40S**
mid-range, **A100/H100** for training and large-model serving, **H200/B200** for frontier-scale.
GPU time is billed per second a container is alive — never attach a GPU to a CPU-only job, and
keep `scaledown_window` tight so idle GPU containers don't burn money.

## Scaling & lifecycle

Tune these keyword args on `@app.function`, each with its why:

| Param | Effect | Why |
|---|---|---|
| `min_containers=N` | Keep N warm instances always running | Kills cold starts for latency-sensitive endpoints (costs idle compute) |
| `buffer_containers=N` | Pre-warm N extra beyond current load | Smooths bursty traffic |
| `scaledown_window=300` | Seconds an idle container lingers before shutdown | Reuse hot containers across nearby calls; lower = cheaper, higher = warmer |
| `timeout=600` | Max seconds a single call may run | Caps runaway jobs |
| `retries=3` | Auto-retry failed inputs | Survives transient failures in `.map()` fan-outs |

Migration note: `keep_warm` → **`min_containers`** and `container_idle_timeout` →
**`scaledown_window`** in the 1.0 migration. The old names are gone.

Concurrency within a container is now its own decorator: **`@modal.concurrent(max_inputs=N)`**
stacked under `@app.function` (it replaces the old `allow_concurrent_inputs=` argument). Use it
so one container handles N simultaneous requests instead of one-per-container.

## Volumes & Secrets

A `Volume` is a distributed filesystem you mount into containers to persist data across runs —
the canonical use is caching downloaded model weights so cold starts skip the re-download.

```python
weights = modal.Volume.from_name("hf-cache", create_if_missing=True)


@app.function(image=image, gpu="H100", volumes={"/cache": weights})
def serve_model():
    # Reader: refresh the view so you see writes from other containers.
    weights.reload()
    # ... load model from /cache ...


@app.function(image=image, volumes={"/cache": weights})
def download_weights():
    # ... write files into /cache ...
    weights.commit()  # WITHOUT this, writes are NOT durable across containers
```

**Gotcha:** writers must call `vol.commit()` to persist; readers call `vol.reload()` to see
another container's committed writes. Forgetting `commit()` is the #1 "my cache is empty"
bug — the files existed in that container and vanished with it.

Secrets land as **environment variables** in the container:

```python
@app.function(image=image, secrets=[modal.Secret.from_name("hf-token")])
def pull():
    import os

    token = os.environ["HF_TOKEN"]  # value injected from the named Modal Secret
```

Never bake a token into the image (`.run_commands("export TOKEN=...")`) — it's recorded in
layer history. Use a Secret. The cookbook above also carries the HF/OpenAI secret patterns.

## Web endpoints

Stack a web decorator **under** `@app.function`. Pick by surface:

| Decorator | Use for | Needs |
|---|---|---|
| `@modal.fastapi_endpoint()` | A single GET/POST function-as-URL | `fastapi[standard]` in image |
| `@modal.asgi_app()` | A full FastAPI/Starlette app you return | `fastapi[standard]` |
| `@modal.wsgi_app()` | A Flask/Django WSGI app | the framework |
| `@modal.web_server(port=8000)` | Your own server process (e.g. vLLM) on a port | the server |

**Decorator stack order matters:** `@app.function` is outermost (top), then optional
`@modal.concurrent`, then the web decorator innermost (bottom, closest to `def`).

```python
@app.function(image=image, gpu="H100", min_containers=1, scaledown_window=300)
@modal.concurrent(max_inputs=10)   # middle
@modal.asgi_app()                  # innermost
def web():
    from fastapi import FastAPI

    api = FastAPI()

    @api.get("/health")
    def health():
        return {"ok": True}

    return api
```

Develop with `modal serve app.py` (hot-reload); ship with `modal deploy app.py` (stable URL).
For custom domains, proxy-auth tokens, batching (`@modal.batched`), and concurrency tuning →
[`references/web-and-scaling.md`](references/web-and-scaling.md). For the FastAPI app's *own*
design (routes, Pydantic, deps), that's [`fastapi`](../fastapi/SKILL.md) — this skill only mounts it.

## Scheduled jobs

```python
# Fixed wall-clock time, with timezone — survives redeploys at the same clock time.
@app.function(schedule=modal.Cron("0 6 * * *", timezone="America/New_York"))
def nightly_report(): ...


# Interval relative to deploy time.
@app.function(schedule=modal.Period(hours=5))
def every_five_hours(): ...
```

**Gotcha:** `Period` is measured from deploy time and **resets on every redeploy** — redeploy
at 4:59 and your "every 5 hours" clock restarts. `Cron` is wall-clock stable; prefer it for
"run at 6am" semantics. Either way you must **`modal deploy`** (not `modal run`) for the
schedule to live on the platform.

## Parallelism

Fan a function out across containers without managing a pool:

```python
@app.local_entrypoint()
def main():
    urls = ["https://a.com", "https://b.com", "https://c.com"]
    # .map: one arg per call, results in input order.
    sizes = list(fetch.map(urls))
    # .starmap: each item is an argument tuple. .spawn: fire-and-forget -> handle.get() later.
    handle = fetch.spawn("https://slow.com")
    print(sizes, handle.get())
```

`.map(iterable)` returns results in order by default; pass `order_outputs=False` to yield as
they complete (faster when latencies vary). Combine with `retries=` on the function so a
single bad input doesn't sink the batch.

## Anti-patterns

| Anti-pattern | Do instead |
|---|---|
| "I'll use `gpu=modal.gpu.A100()` like the old docs" | Removed in 1.0. Use the string `gpu="A100-80GB"`. |
| "Attach a GPU, it might speed up this CPU job" | GPU is billed per second alive. CPU-only job → no `gpu=`. |
| "My files are written, the Volume will keep them" | Not without `vol.commit()` (writer) / `vol.reload()` (reader). |
| "Pin later — `uv_pip_install('torch')` is fine for now" | Unpinned deps drift; builds aren't reproducible. Pin every version. |
| "`modal run` it, the endpoint/schedule will stay up" | `run` is ephemeral; it exits. Use `modal deploy` for anything persistent. |
| "Order the decorators however — Modal figures it out" | `@app.function` outermost, web decorator innermost. Wrong order errors. |
| "Bake the HF token into the image with `run_commands`" | Leaks into layer history. Use `modal.Secret.from_name(...)`. |
| "Just call the model via a managed API through Modal" | If you write no container, that's a managed-API job → [`replicate`](../replicate/SKILL.md). |
| "I need a box to SSH into for a week" | That's a persistent rental → [`runpod`](../runpod/SKILL.md), not Modal's scale-to-zero. |
| "Set `min_containers` high so it's always fast" | Idle warm containers cost money 24/7. Tune `scaledown_window` first. |

## Verify

[`scripts/verify.sh`](scripts/verify.sh) `[TARGET]` statically lints the nearest emitted Modal
`*.py`: it requires a `modal.App(`, **fails** if the removed `modal.gpu.` object form appears,
checks that any web decorator sits under an `@app.function`, and that any `Volume` uses
`from_name(..., create_if_missing=...)`. It runs `python -c "import modal"` only if modal is
installed (skip-pass otherwise), needs **no Modal credentials**, and exits 0 on an empty target.

## Project grounding (02-DOCS)

In a project with a `02-DOCS/` layer (the [`harness`](../harness/SKILL.md) wiki), read
`02-DOCS/wiki/stack/modal.md` first, then record this app's real Modal choices there — GPU types,
image base, Volume names, schedule, endpoint shape — and index it in `02-DOCS/wiki/index.md`. No
`02-DOCS/`? Skip silently.

