Together Dedicated Model Inference

Deploy and operate models on dedicated GPUs with Together AI's Dedicated Model Inference (DMI, the v2 dedicated endpoints API): beta endpoints, deployments, deployment profiles and hardware configs, autoscaling, traffic splitting, A/B tests, shadow experiments, Prometheus metrics, and custom model or LoRA adapter uploads. Reach for it whenever the user mentions together beta endpoints or tg beta commands, client.beta.endpoints, DMI resources like ep_/dep_/cr_/ml_ IDs, or wants production model serving with traffic management on Together AI. This is the current dedicated-hosting API and also covers migrating off the retired legacy v1 endpoints API (non-beta client.endpoints / together endpoints), whose create and restart now return HTTP 403.

togethercomputer 85a168a 8 files · 90.7 KB Updated

File contents

togethercomputer/skills/tree/main/skills/together-dedicated-model-inference commit 85a168abd7

Frequently asked questions

npx skillmds@latest add togethercomputer/together-dedicated-model-inference