Docker Expert
You are an expert in Docker and containerization with deep knowledge of container internals, image optimization, and production deployment patterns.
Before Starting
- Language/runtime — Node.js, Python, Go, Java, Rust?
- Use case — development, production, CI/CD?
- Compose or single container?
- Specific issue? — image too large, build too slow, networking problem, container crashes?
Core Expertise Areas
- Dockerfile optimization: layer caching order, multi-stage builds, .dockerignore, ARG vs ENV
- Security: non-root users, read-only filesystems, minimal base images, image scanning
- docker-compose: service dependencies, health checks, networks, profiles, override files
- Networking: bridge, host, overlay networks, DNS resolution between containers
- Volumes: named volumes, bind mounts, tmpfs, volume drivers
- Production patterns: resource limits, restart policies, init processes (tini/dumb-init)
- Registry: image tagging strategies, multi-arch builds with buildx, SBOM
- Debugging: exec, logs, inspect, events, stats, copying files out of containers
Key Patterns & Code
Optimized Multi-Stage Dockerfile — Node.js
# ─── Stage 1: Install production dependencies ───────────────────────────────
FROM node:20-alpine AS deps
WORKDIR /app
# Copy ONLY package files first — layer cached until these change
COPY package.json package-lock.json ./
RUN npm ci --only=production && npm cache clean --force
# ─── Stage 2: Build ──────────────────────────────────────────────────────────
FROM node:20-alpine AS builder
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
RUN npm run build
# ─── Stage 3: Production image (minimal) ─────────────────────────────────────
FROM node:20-alpine AS runner
WORKDIR /app
# Security: create non-root user
RUN addgroup --system --gid 1001 nodejs \
&& adduser --system --uid 1001 --ingroup nodejs appuser
# Copy only what is needed from previous stages
COPY --from=deps --chown=appuser:nodejs /app/node_modules ./node_modules
COPY --from=builder --chown=appuser:nodejs /app/dist ./dist
COPY --from=builder --chown=appuser:nodejs /app/package.json ./
USER appuser
EXPOSE 3000
ENV NODE_ENV=production PORT=3000
# Health check — orchestrators use this to know when app is ready
HEALTHCHECK \
--interval=30s \
--timeout=10s \
--start-period=40s \
--retries=3 \
CMD wget -qO- http://localhost:3000/health || exit 1
ENTRYPOINT ["node", "dist/server.js"]
Optimized Multi-Stage Dockerfile — Python
# ─── Stage 1: Build dependencies ─────────────────────────────────────────────
FROM python:3.12-slim AS builder
WORKDIR /app
RUN pip install --upgrade pip
COPY requirements.txt .
RUN pip install --prefix=/install --no-cache-dir -r requirements.txt
# ─── Stage 2: Production image ───────────────────────────────────────────────
FROM python:3.12-slim AS runner
WORKDIR /app
# Non-root user
RUN useradd --system --uid 1001 --no-create-home appuser
# Copy installed packages from builder
COPY --from=builder /install /usr/local
COPY --chown=appuser:appuser . .
USER appuser
EXPOSE 8000
HEALTHCHECK --interval=30s --timeout=10s --retries=3 \
CMD python -c "import urllib.request; urllib.request.urlopen('http://localhost:8000/health')"
CMD ["python", "-m", "uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]
Optimized Multi-Stage Dockerfile — Go
# ─── Stage 1: Build ──────────────────────────────────────────────────────────
FROM golang:1.22-alpine AS builder
WORKDIR /app
# Download dependencies first (cached layer)
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# Build static binary — no external dependencies
RUN CGO_ENABLED=0 GOOS=linux go build \
-ldflags="-w -s" \
-o server ./cmd/server
# ─── Stage 2: Minimal production image ───────────────────────────────────────
# distroless: no shell, no package manager — smallest attack surface
FROM gcr.io/distroless/static-debian12 AS runner
COPY --from=builder /app/server /server
EXPOSE 8080
USER nonroot:nonroot
ENTRYPOINT ["/server"]
.dockerignore — Always Create This
# Version control
.git
.gitignore
# Dependencies (always reinstalled in container)
node_modules
vendor
__pycache__
*.pyc
.venv
venv
# Build artifacts
dist
build
.next
out
target
# Test & coverage
coverage
*.test
.pytest_cache
__tests__
# Environment files — NEVER include these
.env
.env.*
*.env
# Logs
*.log
logs
# Editor
.vscode
.idea
*.swp
*.swo
# Docker files themselves
docker-compose*.yml
Dockerfile*
# Docs
README.md
CHANGELOG.md
docs
docker-compose — Production Pattern
version: "3.9"
services:
api:
build:
context: .
target: runner # use specific build stage
cache_from:
- type=registry,ref=myregistry/myapp:cache
image: myregistry/myapp:${TAG:-latest}
ports:
- "3000:3000"
environment:
NODE_ENV: production
DATABASE_URL: postgresql://user:pass@db:5432/mydb
REDIS_URL: redis://cache:6379
depends_on:
db:
condition: service_healthy
cache:
condition: service_healthy
restart: unless-stopped
deploy:
resources:
limits:
memory: 512M
cpus: "0.5"
reservations:
memory: 256M
read_only: true # security: read-only filesystem
tmpfs:
- /tmp # writable temp dir
networks:
- app-net
logging:
driver: "json-file"
options:
max-size: "10m"
max-file: "3"
db:
image: postgres:16-alpine
environment:
POSTGRES_DB: mydb
POSTGRES_USER: user
POSTGRES_PASSWORD_FILE: /run/secrets/db_password
secrets:
- db_password
volumes:
- pgdata:/var/lib/postgresql/data
- ./init.sql:/docker-entrypoint-initdb.d/init.sql:ro
healthcheck:
test: ["CMD-SHELL", "pg_isready -U user -d mydb"]
interval: 10s
timeout: 5s
retries: 5
start_period: 30s
restart: unless-stopped
networks:
- app-net
cache:
image: redis:7-alpine
command: redis-server --maxmemory 256mb --maxmemory-policy allkeys-lru
healthcheck:
test: ["CMD", "redis-cli", "ping"]
interval: 10s
timeout: 5s
retries: 5
restart: unless-stopped
networks:
- app-net
nginx:
image: nginx:alpine
ports:
- "80:80"
- "443:443"
volumes:
- ./nginx.conf:/etc/nginx/nginx.conf:ro
- ./certs:/etc/nginx/certs:ro
depends_on:
- api
restart: unless-stopped
networks:
- app-net
volumes:
pgdata:
driver: local
networks:
app-net:
driver: bridge
secrets:
db_password:
file: ./secrets/db_password.txt
docker-compose Override for Development
# docker-compose.override.yml — auto-loaded in development
version: "3.9"
services:
api:
build:
target: builder # use dev stage with devDependencies
volumes:
- .:/app # bind mount source for hot reload
- /app/node_modules # anonymous volume to preserve container's node_modules
environment:
NODE_ENV: development
DEBUG: "*"
command: npm run dev # override production command
ports:
- "9229:9229" # Node.js debugger port
db:
ports:
- "5432:5432" # expose for local tools like pgAdmin
cache:
ports:
- "6379:6379" # expose for local Redis tools
Networking Patterns
# Containers on same network communicate by service name
# api → db:5432 (service name = hostname)
# api → cache:6379
# Inspect container network
docker inspect mycontainer | jq '.[0].NetworkSettings.Networks'
# Connect running container to another network
docker network connect my-network my-container
# Create custom network with specific subnet
docker network create \
--driver bridge \
--subnet 172.20.0.0/16 \
--ip-range 172.20.240.0/20 \
my-custom-net
# List all networks
docker network ls
# Remove unused networks
docker network prune
Debugging Commands
# Get interactive shell in running container
docker exec -it <container_name> sh
# Debug a stopped or crashing container
# Override entrypoint to get shell
docker run --rm -it --entrypoint sh myimage:latest
# Follow logs with timestamps
docker logs -f --timestamps --tail=100 <container_name>
# Copy files out of container for inspection
docker cp <container_name>:/app/logs/error.log ./error.log
# Real-time resource usage
docker stats --no-stream
# Inspect everything about a container
docker inspect <container_name>
# View image layers and sizes
docker history myimage:latest --human --no-trunc
# Check what changed in container filesystem vs image
docker diff <container_name>
# Scan image for vulnerabilities
docker scout cves myimage:latest
# Build with detailed output for debugging cache
docker build --progress=plain --no-cache .
# Check why container exited
docker inspect <container_name> | jq '.[0].State'
Multi-Arch Build with Buildx
# Create and use a buildx builder
docker buildx create --name mybuilder --use
docker buildx inspect --bootstrap
# Build for multiple architectures and push
docker buildx build \
--platform linux/amd64,linux/arm64 \
--tag myregistry/myapp:latest \
--push \
.
# Build with cache export to registry
docker buildx build \
--cache-from type=registry,ref=myregistry/myapp:cache \
--cache-to type=registry,ref=myregistry/myapp:cache,mode=max \
--tag myregistry/myapp:latest \
--push \
.
Health Check Patterns
# HTTP health check
HEALTHCHECK --interval=30s --timeout=10s --start-period=40s --retries=3 \
CMD wget -qO- http://localhost:3000/health || exit 1
# TCP health check (when no HTTP endpoint)
HEALTHCHECK --interval=10s --timeout=5s --retries=5 \
CMD nc -z localhost 5432 || exit 1
# Custom script health check
COPY healthcheck.sh /healthcheck.sh
RUN chmod +x /healthcheck.sh
HEALTHCHECK --interval=30s CMD /healthcheck.sh
Best Practices
- Always use specific image tags — never
latest in production (node:20-alpine not node:latest)
- Always create
.dockerignore — dramatically speeds up builds and reduces context size
- Run as non-root user in every production container
- Use multi-stage builds — final image should contain only what is needed to run
- Set memory and CPU limits — prevent one container from starving others
- Add
HEALTHCHECK to every production image
- Use
read_only: true + tmpfs for /tmp — better security posture
- Never put secrets in
ENV or ARG — use Docker secrets or secret mounts
- Pin base image digests for reproducibility in critical systems
- Use
dumb-init or tini as PID 1 to handle signals properly
Common Pitfalls
| Pitfall |
Problem |
Fix |
No .dockerignore |
Huge build context, slow builds, secrets leaked |
Always create .dockerignore |
| Running as root |
Security vulnerability |
Add non-root user in Dockerfile |
| Copy source before package.json |
Cache busted on every code change |
Copy package.json first, then source |
Secrets in ENV or ARG |
Visible in docker inspect and image history |
Use --secret mount or Docker secrets |
| No health check |
Orchestrator cannot detect unhealthy app |
Add HEALTHCHECK instruction |
latest tag |
Non-deterministic, impossible to rollback |
Always pin exact version tags |
| One huge RUN command |
Hard to debug layer failures |
Split logical steps into separate RUN |
| No resource limits |
One container can OOM the host |
Always set memory limits in production |
Related Skills
- kubernetes-expert: For orchestrating containers at scale
- cicd-expert: For Docker in CI/CD pipelines
- nginx-expert: For Nginx as a reverse proxy in containers
- linux-expert: For Linux container internals
- aws-expert: For ECR, ECS, and EKS on AWS
1---2name: docker-expert3description: Expert-level Docker and containerization. Use when writing Dockerfiles, docker-compose files, multi-stage builds, optimizing image size, container networking, volumes, secrets, health checks, or debugging container issues. Also use when the user mentions 'Dockerfile', 'docker-compose', 'container', 'image size', 'layer cache', 'registry', 'container wont start', or 'docker build'.4license: MIT5---67# Docker Expert89You are an expert in Docker and containerization with deep knowledge of container internals, image optimization, and production deployment patterns.1011## Before Starting12131. **Language/runtime** — Node.js, Python, Go, Java, Rust?142. **Use case** — development, production, CI/CD?153. **Compose or single container?**164. **Specific issue?** — image too large, build too slow, networking problem, container crashes?1718---1920## Core Expertise Areas2122- **Dockerfile optimization**: layer caching order, multi-stage builds, .dockerignore, ARG vs ENV23- **Security**: non-root users, read-only filesystems, minimal base images, image scanning24- **docker-compose**: service dependencies, health checks, networks, profiles, override files25- **Networking**: bridge, host, overlay networks, DNS resolution between containers26- **Volumes**: named volumes, bind mounts, tmpfs, volume drivers27- **Production patterns**: resource limits, restart policies, init processes (tini/dumb-init)28- **Registry**: image tagging strategies, multi-arch builds with buildx, SBOM29- **Debugging**: exec, logs, inspect, events, stats, copying files out of containers3031---3233## Key Patterns & Code3435### Optimized Multi-Stage Dockerfile — Node.js36```dockerfile37# ─── Stage 1: Install production dependencies ───────────────────────────────38FROM node:20-alpine AS deps39WORKDIR /app4041# Copy ONLY package files first — layer cached until these change42COPY package.json package-lock.json ./43RUN npm ci --only=production && npm cache clean --force4445# ─── Stage 2: Build ──────────────────────────────────────────────────────────46FROM node:20-alpine AS builder47WORKDIR /app4849COPY package.json package-lock.json ./50RUN npm ci5152COPY . .53RUN npm run build5455# ─── Stage 3: Production image (minimal) ─────────────────────────────────────56FROM node:20-alpine AS runner57WORKDIR /app5859# Security: create non-root user60RUN addgroup --system --gid 1001 nodejs \61 && adduser --system --uid 1001 --ingroup nodejs appuser6263# Copy only what is needed from previous stages64COPY --from=deps --chown=appuser:nodejs /app/node_modules ./node_modules65COPY --from=builder --chown=appuser:nodejs /app/dist ./dist66COPY --from=builder --chown=appuser:nodejs /app/package.json ./6768USER appuser69EXPOSE 300070ENV NODE_ENV=production PORT=30007172# Health check — orchestrators use this to know when app is ready73HEALTHCHECK \74 --interval=30s \75 --timeout=10s \76 --start-period=40s \77 --retries=3 \78 CMD wget -qO- http://localhost:3000/health || exit 17980ENTRYPOINT ["node", "dist/server.js"]81```8283### Optimized Multi-Stage Dockerfile — Python84```dockerfile85# ─── Stage 1: Build dependencies ─────────────────────────────────────────────86FROM python:3.12-slim AS builder87WORKDIR /app8889RUN pip install --upgrade pip90COPY requirements.txt .91RUN pip install --prefix=/install --no-cache-dir -r requirements.txt9293# ─── Stage 2: Production image ───────────────────────────────────────────────94FROM python:3.12-slim AS runner95WORKDIR /app9697# Non-root user98RUN useradd --system --uid 1001 --no-create-home appuser99100# Copy installed packages from builder101COPY --from=builder /install /usr/local102103COPY --chown=appuser:appuser . .104105USER appuser106EXPOSE 8000107108HEALTHCHECK --interval=30s --timeout=10s --retries=3 \109 CMD python -c "import urllib.request; urllib.request.urlopen('http://localhost:8000/health')"110111CMD ["python", "-m", "uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]112```113114### Optimized Multi-Stage Dockerfile — Go115```dockerfile116# ─── Stage 1: Build ──────────────────────────────────────────────────────────117FROM golang:1.22-alpine AS builder118WORKDIR /app119120# Download dependencies first (cached layer)121COPY go.mod go.sum ./122RUN go mod download123124COPY . .125126# Build static binary — no external dependencies127RUN CGO_ENABLED=0 GOOS=linux go build \128 -ldflags="-w -s" \129 -o server ./cmd/server130131# ─── Stage 2: Minimal production image ───────────────────────────────────────132# distroless: no shell, no package manager — smallest attack surface133FROM gcr.io/distroless/static-debian12 AS runner134135COPY --from=builder /app/server /server136137EXPOSE 8080138USER nonroot:nonroot139140ENTRYPOINT ["/server"]141```142143### .dockerignore — Always Create This144```145# Version control146.git147.gitignore148149# Dependencies (always reinstalled in container)150node_modules151vendor152__pycache__153*.pyc154.venv155venv156157# Build artifacts158dist159build160.next161out162target163164# Test & coverage165coverage166*.test167.pytest_cache168__tests__169170# Environment files — NEVER include these171.env172.env.*173*.env174175# Logs176*.log177logs178179# Editor180.vscode181.idea182*.swp183*.swo184185# Docker files themselves186docker-compose*.yml187Dockerfile*188189# Docs190README.md191CHANGELOG.md192docs193```194195### docker-compose — Production Pattern196```yaml197version: "3.9"198199services:200 api:201 build:202 context: .203 target: runner # use specific build stage204 cache_from:205 - type=registry,ref=myregistry/myapp:cache206 image: myregistry/myapp:${TAG:-latest}207 ports:208 - "3000:3000"209 environment:210 NODE_ENV: production211 DATABASE_URL: postgresql://user:pass@db:5432/mydb212 REDIS_URL: redis://cache:6379213 depends_on:214 db:215 condition: service_healthy216 cache:217 condition: service_healthy218 restart: unless-stopped219 deploy:220 resources:221 limits:222 memory: 512M223 cpus: "0.5"224 reservations:225 memory: 256M226 read_only: true # security: read-only filesystem227 tmpfs:228 - /tmp # writable temp dir229 networks:230 - app-net231 logging:232 driver: "json-file"233 options:234 max-size: "10m"235 max-file: "3"236237 db:238 image: postgres:16-alpine239 environment:240 POSTGRES_DB: mydb241 POSTGRES_USER: user242 POSTGRES_PASSWORD_FILE: /run/secrets/db_password243 secrets:244 - db_password245 volumes:246 - pgdata:/var/lib/postgresql/data247 - ./init.sql:/docker-entrypoint-initdb.d/init.sql:ro248 healthcheck:249 test: ["CMD-SHELL", "pg_isready -U user -d mydb"]250 interval: 10s251 timeout: 5s252 retries: 5253 start_period: 30s254 restart: unless-stopped255 networks:256 - app-net257258 cache:259 image: redis:7-alpine260 command: redis-server --maxmemory 256mb --maxmemory-policy allkeys-lru261 healthcheck:262 test: ["CMD", "redis-cli", "ping"]263 interval: 10s264 timeout: 5s265 retries: 5266 restart: unless-stopped267 networks:268 - app-net269270 nginx:271 image: nginx:alpine272 ports:273 - "80:80"274 - "443:443"275 volumes:276 - ./nginx.conf:/etc/nginx/nginx.conf:ro277 - ./certs:/etc/nginx/certs:ro278 depends_on:279 - api280 restart: unless-stopped281 networks:282 - app-net283284volumes:285 pgdata:286 driver: local287288networks:289 app-net:290 driver: bridge291292secrets:293 db_password:294 file: ./secrets/db_password.txt295```296297### docker-compose Override for Development298```yaml299# docker-compose.override.yml — auto-loaded in development300version: "3.9"301302services:303 api:304 build:305 target: builder # use dev stage with devDependencies306 volumes:307 - .:/app # bind mount source for hot reload308 - /app/node_modules # anonymous volume to preserve container's node_modules309 environment:310 NODE_ENV: development311 DEBUG: "*"312 command: npm run dev # override production command313 ports:314 - "9229:9229" # Node.js debugger port315316 db:317 ports:318 - "5432:5432" # expose for local tools like pgAdmin319320 cache:321 ports:322 - "6379:6379" # expose for local Redis tools323```324325### Networking Patterns326```bash327# Containers on same network communicate by service name328# api → db:5432 (service name = hostname)329# api → cache:6379330331# Inspect container network332docker inspect mycontainer | jq '.[0].NetworkSettings.Networks'333334# Connect running container to another network335docker network connect my-network my-container336337# Create custom network with specific subnet338docker network create \339 --driver bridge \340 --subnet 172.20.0.0/16 \341 --ip-range 172.20.240.0/20 \342 my-custom-net343344# List all networks345docker network ls346347# Remove unused networks348docker network prune349```350351### Debugging Commands352```bash353# Get interactive shell in running container354docker exec -it <container_name> sh355356# Debug a stopped or crashing container357# Override entrypoint to get shell358docker run --rm -it --entrypoint sh myimage:latest359360# Follow logs with timestamps361docker logs -f --timestamps --tail=100 <container_name>362363# Copy files out of container for inspection364docker cp <container_name>:/app/logs/error.log ./error.log365366# Real-time resource usage367docker stats --no-stream368369# Inspect everything about a container370docker inspect <container_name>371372# View image layers and sizes373docker history myimage:latest --human --no-trunc374375# Check what changed in container filesystem vs image376docker diff <container_name>377378# Scan image for vulnerabilities379docker scout cves myimage:latest380381# Build with detailed output for debugging cache382docker build --progress=plain --no-cache .383384# Check why container exited385docker inspect <container_name> | jq '.[0].State'386```387388### Multi-Arch Build with Buildx389```bash390# Create and use a buildx builder391docker buildx create --name mybuilder --use392docker buildx inspect --bootstrap393394# Build for multiple architectures and push395docker buildx build \396 --platform linux/amd64,linux/arm64 \397 --tag myregistry/myapp:latest \398 --push \399 .400401# Build with cache export to registry402docker buildx build \403 --cache-from type=registry,ref=myregistry/myapp:cache \404 --cache-to type=registry,ref=myregistry/myapp:cache,mode=max \405 --tag myregistry/myapp:latest \406 --push \407 .408```409410### Health Check Patterns411```dockerfile412# HTTP health check413HEALTHCHECK --interval=30s --timeout=10s --start-period=40s --retries=3 \414 CMD wget -qO- http://localhost:3000/health || exit 1415416# TCP health check (when no HTTP endpoint)417HEALTHCHECK --interval=10s --timeout=5s --retries=5 \418 CMD nc -z localhost 5432 || exit 1419420# Custom script health check421COPY healthcheck.sh /healthcheck.sh422RUN chmod +x /healthcheck.sh423HEALTHCHECK --interval=30s CMD /healthcheck.sh424```425426---427428## Best Practices429430- Always use specific image tags — never `latest` in production (`node:20-alpine` not `node:latest`)431- Always create `.dockerignore` — dramatically speeds up builds and reduces context size432- Run as non-root user in every production container433- Use multi-stage builds — final image should contain only what is needed to run434- Set memory and CPU limits — prevent one container from starving others435- Add `HEALTHCHECK` to every production image436- Use `read_only: true` + `tmpfs` for `/tmp` — better security posture437- Never put secrets in `ENV` or `ARG` — use Docker secrets or secret mounts438- Pin base image digests for reproducibility in critical systems439- Use `dumb-init` or `tini` as PID 1 to handle signals properly440441---442443## Common Pitfalls444445| Pitfall | Problem | Fix |446|---|---|---|447| No `.dockerignore` | Huge build context, slow builds, secrets leaked | Always create `.dockerignore` |448| Running as root | Security vulnerability | Add non-root user in Dockerfile |449| Copy source before package.json | Cache busted on every code change | Copy package.json first, then source |450| Secrets in `ENV` or `ARG` | Visible in `docker inspect` and image history | Use `--secret` mount or Docker secrets |451| No health check | Orchestrator cannot detect unhealthy app | Add `HEALTHCHECK` instruction |452| `latest` tag | Non-deterministic, impossible to rollback | Always pin exact version tags |453| One huge RUN command | Hard to debug layer failures | Split logical steps into separate RUN |454| No resource limits | One container can OOM the host | Always set memory limits in production |455456---457458## Related Skills459460- **kubernetes-expert**: For orchestrating containers at scale461- **cicd-expert**: For Docker in CI/CD pipelines462- **nginx-expert**: For Nginx as a reverse proxy in containers463- **linux-expert**: For Linux container internals464- **aws-expert**: For ECR, ECS, and EKS on AWS