RunPod Serverless GPU Inference
Deploy and manage GPU inference endpoints on RunPod Serverless using their REST API. Handles endpoint creation, cold start optimization, request queuing, and auto-scaling configuration for image generation models.
Installation
Use the upstream install or setup path that matches your environment:
- Make API requests
Requirements and caveats from upstream:
- The container instances that execute your code when requests arrive at your endpoint. Each worker runs your custom Docker container with your application code and dependencies. Runpod automatically manages worker life...
- Build and push the worker image to Docker Hub (or another container registry).
Basic usage or getting-started notes:
Concepts
Manage API keys
Agent skills NEW