Results for “venly”
12 skillsvss-deploy-profile
Selects, configures, deploys, verifies, debugs, or tears down a VSS profile (base, search, lvs, warehouse, edge) for NVIDIA's video search and summarization stack.
2.2k · bundle
managing-ably
Manages and analyzes Ably real-time messaging resources, covering channels, presence, connections, usage analytics, and account health via REST and Control APIs.
7
vss-deploy-dense-captioning
Deploy a standalone RT-VLM dense-captioning microservice and exercise its REST API endpoints for file upload, caption generation, streaming, chat completions, and Kafka integration.
2.2k · bundle
vercel-cli-with-tokens
Deploy and manage projects on Vercel using token-based authentication, without interactive login.
28.7k
jetson-llm-serve
Serve LLMs and VLMs on NVIDIA Jetson devices using vLLM or SGLang with optimized Docker containers and quantization presets.
2.2k · bundle
vercel-deploy
Deploy applications and websites to Vercel as preview or production deployments.
23.3k · bundle
vss-deploy-detection-tracking-2d
Deploy, debug, and operate the RTVI-CV 2D detection/tracking microservice and call its REST API for stream management, health checks, and metrics.
2.2k · bundle
vss-setup-video-analytics-api
Deploys the vss-video-analytics-api REST service standalone with config, data-log bind, and optional Elasticsearch/Kafka connectivity.
2.2k · bundle
vercel
Deploy applications and manage projects with complete CLI reference. Commands for deployments, projects, domains, environment variables, and live documentation access.
2 · bundle
vercel-deploy
Operate Vercel deployments including preview deploys, production deploys, staged promote flows, alias/domain management, environment variable sync, and rollback response.
42 · bundle
vast-gpu
Rent, manage, and destroy GPU instances on vast.ai. Use when user says "rent gpu", "vast.ai", "rent a server", "cloud gpu", or needs on-demand GPU without owning hardware.
1k
vllm
You are an expert in vLLM, the high-throughput LLM serving engine. You help developers deploy open-source models (Llama, Mistral, Qwen, Phi, Gemma) with PagedAttention for efficient memory management, continuous batching, tensor parallelism for multi-GPU, OpenAI-compatible API, and quantization support — achieving 2-24x higher throughput than HuggingFace Transformers for production LLM serving.
0