Results for “multimodal”

45 skills
More results
lingxling
Daily
Reference for building real-time voice and multimodal AI applications with Pipecat, covering pipelines, speech services, LLM integration, and transports.
253
antigravity
Daily
Build real-time voice and multimodal AI applications using Pipecat and Daily, covering pipeline architecture, AI service integration, and transport options.
42.4k
oyi77
Gemini API Dev
Build applications with the Google Gemini API, covering chat completions, multimodal inputs, function calling, streaming, and grounding with Google Search.
10
phoroth
Daily
Reference for building real-time voice and multimodal AI applications with Daily and Pipecat, covering pipeline architecture, AI service integrations, transports, and client SDKs.
3
lucaspmarie-a11y
Daily
Reference for building real-time voice and multimodal AI applications with Pipecat, covering pipelines, speech services, LLMs, transports, and deployment.
5
orchestra-research
Llama Factory
Provides expert guidance for fine-tuning LLMs with LLaMA-Factory, covering WebUI no-code, 100+ models, 2/3/4/5/6/8-bit QLoRA, and multimodal support.
10.4k · bundle
jorcan
Daily
Provides a reference for building real-time voice and multimodal AI agents with Pipecat, covering pipeline architecture, speech services, LLM integration, transports, and deployment.
0 · bundle
vikingokft
Gemini API Dev
Build applications with Gemini API hosted models, including multimodal content, function calling, and structured outputs, using the latest SDKs and model specifications.
0
joshuashepherd
Gemini API
Builds or debugs Google Gemini features using @google/generative-ai, covering generateContent, function calling, grounding, multimodal input, and streaming, with guidance for editing src/lib/ai/clients/google.ts.
1
google
Google Cloud Solution Agentic AI Bidirectional Streaming
Designs and implements a Google Cloud solution for live, bidirectional multimodal streaming workloads with AI agents, covering requirements discovery, architecture design, and deployment planning.
14.4k
google-gemini
Gemini API Dev
Build applications with Gemini API hosted models, including Gemini and Gemma 4, using multimodal content, function calling, structured outputs, and current SDKs for Python, JavaScript, Go, and Java.
3.8k
google-gemini
Gemini Interactions API
Call the Gemini API for text generation, chat, multimodal understanding, image/video/audio generation, streaming, function calling, structured output, and managed agents using the Interactions API in Python and TypeScript.
3.8k · bundle
google
Gemini API
Guides usage of the Gemini API on Agent Platform with the Google Gen AI SDK, covering SDK usage (Python, JS/TS, Go, Java, C#), capabilities like multimodal inputs, tools, media generation, caching, batch prediction, and Live API.
14.4k · bundle
inehemiasm
Firebase AI Logic Basics
Official skill for integrating Firebase AI Logic (Gemini API) into web applications. Covers setup, multimodal inference, structured output, and security.
0 · bundle
jiachen-t-wang
Mmbench Is Your Multi Modal Model An All Around Player Arxiv
MMBench: Is Your Multi-modal Model an All-around Player?
6
manu14357
AI Engineer
Build production-ready LLM applications, advanced RAG systems, and intelligent agents. Implements vector search, multimodal AI, agent orchestration, and enterprise AI integrations.
16
bouclem
AI Engineer
Build production-ready LLM applications, advanced RAG systems, and intelligent agents. Implements vector search, multimodal AI, agent orchestration, and enterprise AI integrations.
7
nvidia
Nemo Mbridge Perf Moe Vlm Training
Provides practical guidance for training Mixture-of-Experts Vision-Language Models in Megatron Bridge, comparing FSDP and 3D-parallel approaches with lessons from recent multimodal experiments.
2.2k · bundle
levalencia
Pathml
Full-featured computational pathology toolkit. Use for advanced WSI analysis including multiplexed immunofluorescence (CODEX, Vectra), nucleus segmentation, tissue graph construction, and ML model training on pathology data. Supports 160+ slide formats. For simple tile extraction from H&E slides, histolab may be simpler.
3 · bundle
lingxling
Pathml
Loads and processes whole-slide pathology images, builds spatial graphs, trains deep learning models, and analyzes multiplexed immunofluorescence data across 160+ slide formats.
253 · bundle
k-dense-ai
Modal
Deploy and serve AI/ML models on Modal's serverless cloud platform with on-demand GPUs, autoscaling containers, persistent storage, and scheduled jobs.
30.2k · bundle
demerzels-lab
Moa
Orchestrates three frontier models to debate a question and synthesizes their best insights into a single superior answer.
10 · bundle
prime-skills
Seedance V2
Generate cinematic short-form video with ByteDance Seedance 2.0 Pro on RunComfy. Documents Seedance 2.0 Pro's strengths (multi-modal references — up to 9 images, 3 videos, 3 audio — synchronized in-pass audio with natural lip-sync, cinematic motion refinement), the 4–15s duration schema, and when to route to HappyHorse 1.0 / Wan 2.7 / Kling instead. Calls `runcomfy run bytedance/seedance-v2/pro` through the local RunComfy CLI. Triggers on "seedance", "seedance 2", "seedance v2", "seedance pro", "bytedance video", or any explicit ask to generate video with this model.
33
composiohq
Openai Automation
Automate OpenAI API operations: generate text and multimodal responses with structured output, create embeddings, generate images, and list models via the Composio MCP integration.
66.9k
vvieira010-pixel
Multi Perspective Decision Wheel
Structure a decision or design challenge through multiple perspectives before committing to action. Use as a synthesis step after scoping, mapping, and dilemma navigation when a group needs a wiser next step.
0
muratcankoylan
Multi Agent Patterns
Design multi-agent systems with context isolation, supervisor or swarm coordination, explicit handoffs, parallel execution, and decision frameworks for when multiple agents are justified.
16.9k · bundle
arustydev
Convert Clojure Roc
Bidirectional conversion between Clojure and Roc. Use when migrating projects between these languages in either direction. Extends meta-convert-dev with Clojure↔Roc specific patterns. Use when migrating Clojure applications to Roc's platform model, translating dynamic functional code to static functional style, or refactoring REPL-driven code to compile-time verified patterns. Extends meta-convert-dev with Clojure-to-Roc specific patterns.
8
nous-hermeshub
AI Engineer
Build production-ready LLM applications, advanced RAG systems, and intelligent agents. Implements vector search, multimodal AI, agent orchestration, and enterprise AI integrations.
1
orchestra-research
Axolotl
Provides expert guidance for fine-tuning LLMs with Axolotl, covering YAML configs, LoRA/QLoRA, DPO/KTO/ORPO/GRPO, and multimodal support.
10.4k · bundle
nimoqup046-collab
Daily
Reference for building real-time voice and multimodal AI agents with Pipecat, covering pipelines, speech services, LLMs, transports, and deployment.
2