Packs
5 packscurated
Optimize GKE Costs
Analyzes current usage, recommends cost-saving measures, and applies optimizations to GKE workloads.
3 skills · pack
curated
GKE Batch & Inference
For teams running batch/HPC and AI/ML inference workloads on GKE with specialized hardware.
2 skills · pack
curated
Secure Google Cloud Workload
Assesses security requirements, identifies risks, and provides actionable recommendations for IAM, network, and data protection.
4 skills · pack
curated
Deploy AI Inference on GKE
Deploy and optimize AI/ML inference workloads on GKE using GPUs, TPUs, and model servers.
3 skills · pack
curated
Google Cloud Well-Architected
For architects evaluating Google Cloud workloads against the Well-Architected Framework pillars: reliability, cost optimization, and operational excellence.
6 skills · pack
Results for “workload”
5 skillsenterprise-agent-ops
Operate long-lived agent workloads with observability, security boundaries, and lifecycle management.
226k
graphsignal-profiler
Set up GPU profiling, tracing, and monitoring for inference workloads using vLLM, SGLang, PyTorch, and dstack services via the Graphsignal Profiler sidecar.
242 · bundle
More results
delegation
Offload sub-tasks to other agents, specialist models, or background jobs, with guidance on choosing the right method and verifying results.
2
tokenwise
Auto-routes Claude Code subtasks to the cheapest capable model (Haiku/Sonnet/Opus), logs token costs, and A/B tests tiers to validate savings against real workloads.
42.4k
google-cloud-solution-agentic-ai-bidirectional-streaming
Designs and implements a Google Cloud solution for live, bidirectional multimodal streaming workloads with AI agents, covering requirements discovery, architecture design, and deployment planning.
14.4k