RunPod Serverless GPU Inference

Deploy and manage GPU inference endpoints on RunPod Serverless using their REST API. Handles endpoint creation, cold start optimization, request queuing, and auto-scaling configuration for image generation models.

agentskillexchange Updated 28 repo stars

File contents

RunPod Serverless GPU Inference

Deploy and manage GPU inference endpoints on RunPod Serverless using their REST API. Handles endpoint creation, cold start optimization, request queuing, and auto-scaling configuration for image generation models.

Installation

Use the upstream install or setup path that matches your environment:

  • Make API requests

Requirements and caveats from upstream:

  • The container instances that execute your code when requests arrive at your endpoint. Each worker runs your custom Docker container with your application code and dependencies. Runpod automatically manages worker life...
  • Build and push the worker image to Docker Hub (or another container registry).

Basic usage or getting-started notes:

Source

agentskillexchange/skills/tree/main/skills/runpod-serverless-gpu-inference commit 8d97086706

Frequently asked questions

npx skillmds@latest add agentskillexchange/runpod-serverless-gpu-inference