Fly.io Reference Architecture
Overview
Production architecture for Fly.io: multi-region web tier, Postgres with read replicas, Redis for caching, background workers, and private networking.
Prerequisites
- A documented data-flow inventory, trust boundaries, ownership, region/retention choices, and disaster-recovery objectives.
- Separate scoped identities for deployment, runtime, database, worker, and observability systems.
Instructions
- Place public ingress, private services, storage, workers, and observability behind explicit network and identity boundaries.
- Define data locality, replication, backup, access, and recovery behavior before creating additional regions or consumers.
- Use staged deployment, health checks, redacted telemetry, and an independently tested rollback per service.
- Validate architecture changes with synthetic traffic and ensure a failure in one region cannot leak secrets or corrupt cross-region state.
Output
Maintain an architecture decision record with components, trust boundaries, data locations, identities, health/rollback controls, owners, and recovery evidence. Do not include secrets or customer data.
Error Handling
- Isolate an unhealthy region or consumer and preserve a safe primary path while recovery proceeds.
- Quarantine unexpected cross-region writes or permission failures for review.
- Restore the previous routing/configuration before replaying queued work.
Examples
Deploy a fictional workload to a staging primary and replica region, deny the worker access to public ingress secrets, and simulate a regional health failure. Verify traffic stays on the healthy route and rollback does not replay writes.
Architecture
┌─────────── Fly.io Anycast DNS ──────────┐
│ │
┌──────▼──────┐ ┌──────────────┐ ┌─────────────▼───┐
│ Web (iad) │ │ Web (lhr) │ │ Web (nrt) │
│ shared-1x │ │ shared-1x │ │ shared-1x │
└──────┬──────┘ └──────┬───────┘ └────────┬────────┘
│ │ │
───────┴────────────────┴────────────────────┴─── .internal DNS
│ │ │
┌──────▼──────┐ ┌──────▼───────┐ ┌────────▼────────┐
│ Postgres │ │ Postgres │ │ Redis │
│ Primary │ │ Replica │ │ (upstash.io) │
│ (iad) │ │ (lhr) │ │ │
└─────────────┘ └──────────────┘ └──────────────────┘
│
┌──────▼──────┐
│ Worker │
│ (iad) │
│ shared-1x │
└─────────────┘
Setup Commands
# 1. Web app — multi-region
fly launch --name my-web --region iad
fly scale count 1 --region lhr
fly scale count 1 --region nrt
# 2. Postgres with replica
fly postgres create --name my-db --region iad
fly postgres attach my-db -a my-web
# Add read replica in Europe
fly machine clone <primary-machine-id> --region lhr -a my-db
# 3. Background worker (same codebase, different process)
fly launch --name my-worker --region iad --no-deploy
# fly.toml for worker: no [http_service], use [processes]
# 4. All communicate via .internal DNS
# my-db.internal:5432 (Postgres)
# my-web.internal:3000 (internal API)
fly.toml Configurations
Web App
app = "my-web"
primary_region = "iad"
[http_service]
internal_port = 3000
force_https = true
auto_stop_machines = "suspend"
min_machines_running = 1
[[vm]]
cpu_kind = "shared"
cpus = 1
memory = "512mb"
Background Worker
app = "my-worker"
primary_region = "iad"
[processes]
worker = "node dist/worker.js"
# No [http_service] — worker doesn't serve HTTP
[[vm]]
cpu_kind = "shared"
cpus = 1
memory = "512mb"
Key Design Decisions
| Decision |
Choice |
Rationale |
| Web tier |
3 regions |
Low latency for global users |
| Database |
Fly Postgres + replica |
Read replicas near users |
| Cache |
Upstash Redis (or Fly Redis) |
Managed, multi-region |
| Workers |
Separate Fly app |
Independent scaling |
| Networking |
6PN (.internal DNS) |
Zero-trust, no public exposure |
| Storage |
Fly Volumes (NVMe) |
Fast, region-local |
Resources
1---2name: flyio-reference-architecture3description: Implement Fly.io reference architecture with multi-region apps, Postgres, Redis, background workers, and private networking. Trigger: "fly.io architecture", "fly.io system design", "fly.io multi-region".4license: MIT5---6# Fly.io Reference Architecture
7
8## Overview
9
10Production architecture for Fly.io: multi-region web tier, Postgres with read replicas, Redis for caching, background workers, and private networking.
11
12## Prerequisites
13
14- A documented data-flow inventory, trust boundaries, ownership, region/retention choices, and disaster-recovery objectives.
15- Separate scoped identities for deployment, runtime, database, worker, and observability systems.
16
17## Instructions
18
191. Place public ingress, private services, storage, workers, and observability behind explicit network and identity boundaries.
202. Define data locality, replication, backup, access, and recovery behavior before creating additional regions or consumers.
213. Use staged deployment, health checks, redacted telemetry, and an independently tested rollback per service.
224. Validate architecture changes with synthetic traffic and ensure a failure in one region cannot leak secrets or corrupt cross-region state.
23
24## Output
25
26Maintain an architecture decision record with components, trust boundaries, data locations, identities, health/rollback controls, owners, and recovery evidence. Do not include secrets or customer data.
27
28## Error Handling
29
30- Isolate an unhealthy region or consumer and preserve a safe primary path while recovery proceeds.
31- Quarantine unexpected cross-region writes or permission failures for review.
32- Restore the previous routing/configuration before replaying queued work.
33
34## Examples
35
36Deploy a fictional workload to a staging primary and replica region, deny the worker access to public ingress secrets, and simulate a regional health failure. Verify traffic stays on the healthy route and rollback does not replay writes.
37
38## Architecture
39
40```
41 ┌─────────── Fly.io Anycast DNS ──────────┐
42 │ │
43 ┌──────▼──────┐ ┌──────────────┐ ┌─────────────▼───┐
44 │ Web (iad) │ │ Web (lhr) │ │ Web (nrt) │
45 │ shared-1x │ │ shared-1x │ │ shared-1x │
46 └──────┬──────┘ └──────┬───────┘ └────────┬────────┘
47 │ │ │
48 ───────┴────────────────┴────────────────────┴─── .internal DNS
49 │ │ │
50 ┌──────▼──────┐ ┌──────▼───────┐ ┌────────▼────────┐
51 │ Postgres │ │ Postgres │ │ Redis │
52 │ Primary │ │ Replica │ │ (upstash.io) │
53 │ (iad) │ │ (lhr) │ │ │
54 └─────────────┘ └──────────────┘ └──────────────────┘
55 │
56 ┌──────▼──────┐
57 │ Worker │
58 │ (iad) │
59 │ shared-1x │
60 └─────────────┘
61```
62
63## Setup Commands
64
65```bash
66# 1. Web app — multi-region
67fly launch --name my-web --region iad
68fly scale count 1 --region lhr
69fly scale count 1 --region nrt
70
71# 2. Postgres with replica
72fly postgres create --name my-db --region iad
73fly postgres attach my-db -a my-web
74# Add read replica in Europe
75fly machine clone <primary-machine-id> --region lhr -a my-db
76
77# 3. Background worker (same codebase, different process)
78fly launch --name my-worker --region iad --no-deploy
79# fly.toml for worker: no [http_service], use [processes]
80
81# 4. All communicate via .internal DNS
82# my-db.internal:5432 (Postgres)
83# my-web.internal:3000 (internal API)
84```
85
86## fly.toml Configurations
87
88### Web App
89
90```toml
91app = "my-web"
92primary_region = "iad"
93
94[http_service]
95 internal_port = 3000
96 force_https = true
97 auto_stop_machines = "suspend"
98 min_machines_running = 1
99
100[[vm]]
101 cpu_kind = "shared"
102 cpus = 1
103 memory = "512mb"
104```
105
106### Background Worker
107
108```toml
109app = "my-worker"
110primary_region = "iad"
111
112[processes]
113 worker = "node dist/worker.js"
114
115# No [http_service] — worker doesn't serve HTTP
116
117[[vm]]
118 cpu_kind = "shared"
119 cpus = 1
120 memory = "512mb"
121```
122
123## Key Design Decisions
124
125| Decision | Choice | Rationale |
126|----------|--------|-----------|
127| Web tier | 3 regions | Low latency for global users |
128| Database | Fly Postgres + replica | Read replicas near users |
129| Cache | Upstash Redis (or Fly Redis) | Managed, multi-region |
130| Workers | Separate Fly app | Independent scaling |
131| Networking | 6PN (.internal DNS) | Zero-trust, no public exposure |
132| Storage | Fly Volumes (NVMe) | Fast, region-local |
133
134## Resources
135
136- [Fly.io Docs](https://fly.io/docs/)
137- [Multi-Region Postgres](https://fly.io/docs/postgres/high-availability-and-global-replication/)
138- [Private Networking](https://fly.io/docs/networking/private-networking/)