# RunPod Serverless GPU Inference

> Deploy and manage GPU inference endpoints on RunPod Serverless using their REST API. Handles endpoint creation, cold start optimization, request queuing, and auto-scaling configuration for image generation models.

- Skill: `agentskillexchange/runpod-serverless-gpu-inference` (Agent Skill)
- Install (CLI): `npx skillmds@latest add agentskillexchange/runpod-serverless-gpu-inference`
- Raw SKILL.md: https://api.skillmd.com/api/skills/agentskillexchange/runpod-serverless-gpu-inference/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: DevOps & Infra
- Author: agentskillexchange (https://skillmd.com/u/agentskillexchange)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/agentskillexchange/runpod-serverless-gpu-inference

---


# RunPod Serverless GPU Inference

Deploy and manage GPU inference endpoints on RunPod Serverless using their REST API. Handles endpoint creation, cold start optimization, request queuing, and auto-scaling configuration for image generation models.

## Installation

Use the upstream install or setup path that matches your environment:
- Make API requests

Requirements and caveats from upstream:
- The container instances that execute your code when requests arrive at your endpoint. Each worker runs your custom Docker container with your application code and dependencies. Runpod automatically manages worker life...
- Build and push the worker image to Docker Hub (or another container registry).

Basic usage or getting-started notes:
- Concepts
- Manage API keys
- Agent skills NEW

- Source: https://docs.runpod.io/serverless/overview

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/runpod-serverless-gpu-inference/)

