LLM API Cost Optimization

Use when you need to measure and reduce LLM/API spend with concrete tactics — per-call token instrumentation, prompt caching, model tiering (cheap vs. premium), request batching, context trimming, and a cost-per-request dashboard. Focuses on measurement plus implementable changes.

FluxonLab 6725abc 11.2 KB Updated

File contents

FluxonLab/Skillry/tree/main/plugins/performance-and-cost/skills/325-llm-api-cost-optimization commit 6725abc1fe

Frequently asked questions

npx skillmds@latest add fluxonlab/llm-api-cost-optimization