Plugins
2 pluginsResults for “inference”
332 skillsLlama Cpp
llama.cpp local GGUF inference + HF Hub model discovery.
1 · bundle
Hf Mem
Estimates the memory required to load Safetensors or GGUF model weights for inference from the Hugging Face Hub, using HTTP Range requests without downloading weights.
10.8k
Speculative Decoding
Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques for 1.5-3.6× speedup without quality loss.
10.4k · bundle
Hf Mem
Estimates GPU memory required to load Safetensors or GGUF model weights for inference from the Hugging Face Hub using HTTP Range requests, without downloading weights locally.
42.4k
Customer Persona
Create data-backed customer personas with market research, demographics, psychographics, jobs-to-be-done, journey mapping, and avatar generation.
584
Prompt Engineering
Learn and apply prompt engineering techniques for LLMs, image generators, and video models using the inference.sh CLI.
584
Recombinator
Simulates meiotic recombination to produce offspring genomes from parent pairs, modeling Mendelian segregation, de novo mutation, sex determination, trait inference, and clinical evaluation against a disease registry.
17 · bundle
Tao Train Dino
Train, evaluate, export, distill, quantize, or run inference for a TAO DINO 2D object detector using transformer-based detection with denoising training and multi-scale features.
2.2k · bundle
Tao Launch Workflow
Collects launch inputs and runs preflight checks before executing TAO workflows such as AutoML, training, evaluation, inference, export, TensorRT engine generation, or DEFT jobs on supported platforms.
2.2k · bundle
Dialogue Audio
Create realistic multi-speaker dialogue audio using Dia TTS via the inference.sh CLI, with control over speaker tags, emotion, pacing, and conversation structure.
584
Gguf Quantization
Convert and quantize models to GGUF format for efficient CPU/GPU inference with llama.cpp, supporting 2-8 bit quantization and Apple Silicon acceleration.
10.4k · bundle
Inference Sh CLI
Runs 150+ AI apps in the cloud via the infsh CLI, covering image generation, video creation, search, 3D, and social automation without needing a GPU.
2
Colibri
Assist with Colibri: pure-C LLM inference engine for running GLM-5.2 (744B MoE) on consumer machines with ~25 GB RAM. Use when setting up, building, converting models, running inference, configuring expert streaming and caching, optimizing speculative decoding (MTP), GPU integration, and integrating Colibri into production pipelines. Includes build setup, model download & conversion, chat/inference modes, performance tuning, and API integration patterns.
42 · bundle
AI Podcast Creation
Create AI-powered podcasts and audio content using text-to-speech, music generation, and audio editing via the inference.sh CLI.
584
Agent UI
Add a batteries-included agent component to React/Next.js apps with runtime, streaming, human-in-the-loop approvals, and client-side tools.
584
Elevenlabs Stt
Transcribe audio with high accuracy using ElevenLabs Scribe models, supporting speaker diarization, audio event tagging, forced alignment, and subtitle generation via the inference.sh CLI.
584
Content Repurposing
Repurpose long-form content into multiple formats including social media posts, newsletters, videos, and quote cards using the inference.sh CLI.
584
Social Media Carousel
Design high-engagement multi-slide carousels for Instagram, LinkedIn, and Twitter/X with layout rules, text hierarchy, swipe psychology, and platform-specific specs.
584
Youtube Thumbnail Design
Design high-CTR YouTube thumbnails with AI image generation, covering dimensions, safe zones, color strategy, text rules, face expression psychology, and A/B testing.
584
AI Infra
Operates AI infrastructure as a production dependency: manages GPU utilization, MCP servers, LLM gateways, inference pipelines, token costs, semantic caching, and model observability.
2
Agent Tools
Run 250+ AI apps from the command line: generate images and videos, call LLMs, search the web, create 3D models, and automate Twitter posts.
584 · bundle
AI Marketing Videos
Create professional marketing videos for ads, promos, product launches, and brand content using AI video generation models and voiceover tools via the inference.sh CLI.
584
Nim Job Status
Check the status and result of an NVIDIA NIM inference job
118 · bundle
Transformers
Load pre-trained models from Hugging Face Hub, run pipeline inference, generate text, and fine-tune models on NLP, vision, audio, and multimodal tasks using the Transformers library.
30.2k · bundle
Storyboard Creation
Generate visual storyboards with AI image generation, covering shot types, camera angles, movement, continuity rules, and panel layout for video planning and pre-production.
584
Press Release Writing
Write professional press releases in AP style with inverted pyramid structure, including formatting, datelines, quotes, boilerplates, and fact-checking via the inference.sh CLI.
584
Hf CLI
Manage Hugging Face Hub resources via the `hf` CLI: download and upload models, datasets, and spaces; manage buckets, cache, collections, discussions, and inference endpoints; run SQL queries on datasets.
2 · bundle
Jetson Inference Mem Tune
Recommends an inference runtime and memory-related launch flags for LLM/VLM workloads on NVIDIA Jetson devices, based on a live memory audit snapshot.
2.2k · bundle
Berry Juicer
Deposit ERC-20 token supply into a Berry Juicer vault on Base to earn trading fees as USDC, then spend the yield as AI inference across 140+ models.
1.2k · bundle
Nim Job Submit
Submit an inference job to NVIDIA NIM and return a job ID
118 · bundle
Serving Llms Vllm
Deploy and serve LLMs with high throughput using vLLM's PagedAttention and continuous batching. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism for production inference.
10.4k · bundle
Ray Data
Process large ML datasets in parallel across CPU or GPU clusters, with streaming execution, multi-format I/O, and integration with Ray Train, PyTorch, and TensorFlow for batch inference and preprocessing pipelines.
3 · bundle
Prime Radiant
Mathematical AI interpretability with sheaf cohomology, spectral analysis, causal inference, and hallucination prevention
0
Tensorrt LLM
Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.
1 · bundle
Tensorrt LLM
Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.
0 · bundle
Agent Llama Cpp V2
Expert en inference llama.cpp avancé (GGUF, quantization, local models, HTTP server, hardware)
6