Results for “image-text”

112 skills
orchestra-research
clip
Enables zero-shot image classification, image-text matching, and cross-modal retrieval using OpenAI's CLIP model.
10.4k · bundle
oyi77
geminigen-ai
Unified multimedia generation API for images, videos, and text-to-speech, replacing separate providers for a single workflow.
10
orchestra-research
blip-2-vision-language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
phoroth
mmx-cli
Generates text, images, video, speech, and music, and performs web searches via the MiniMax AI platform using the mmx terminal CLI.
3
qhjqhj00
vpeval
Evaluates text-to-image generation models by decomposing assessment into five specialized skills (object presence, count, spatial relations, scale, and text rendering) and open-ended prompts, producing interpretable binary scores with visual and textual explanations.
3
seaworld008
sketch
Generating AI image-generation code using the Gemini API. Handles text-to-image generation, image editing, and prompt optimization. Use when image generation code is needed.
65 · bundle
upayanghosh
synapse-image-describe
Provides detailed, structured image descriptions covering objects, people, colors, text, and scene context, with an overview and interpretation.
14
microsoft
azure-ai-contentsafety-py
Detect harmful user-generated and AI-generated content in text and images using Azure AI Content Safety SDK for Python.
2.7k
antigravity
mmx-cli
Generate text, images, video, speech, and music via the MiniMax AI platform using the mmx CLI.
42.4k
lord1egypt
clip
Enables zero-shot image classification, image-text matching, and cross-modal retrieval using OpenAI's CLIP model, with code for semantic search, content moderation, and vector database integration.
2
github
generate-image
Generate images using AI from OpenAI or Google Gemini, with support for textures, icons, sprites, and visual assets.
36.2k
inference-sh
ai-video-generation
Generate videos from text, images, or references using 40+ AI models via the inference.sh CLI.
584
phoroth
imagen
Generates images from text prompts using Google Gemini's image generation model, saving them as PNG files for use in UI, documentation, and design projects.
3
jimliu
baoyu-cover-image
Generates article cover images with customizable type, palette, rendering, text, and mood dimensions, supporting multiple aspect ratios and backends.
23.1k · bundle
nimoqup046-collab
imagen
Generates images from text prompts using Google Gemini's image generation model, saving them as PNG files for use in UI, documentation, and design assets.
2
qhjqhj00
visor
Evaluates text-to-image models on spatial relationship accuracy using the VISOR metric, separating object detection from spatial correctness to reveal biases like object priority and merging.
3
lingxling
imagen
Generates images from text prompts using Google Gemini's image generation model, saving them as PNG files for use in UI, documentation, and design assets.
253
jorcan
imagen
Generates images from text prompts using Google Gemini's image generation model, saving them as PNG files for use in UI, documentation, and design assets.
0 · bundle
jimliu
baoyu-image-gen
Generates images using multiple AI providers including OpenAI, Google, Azure, and others. Supports text-to-image, reference images, aspect ratios, and batch generation from prompt files.
23.1k · bundle
mhassan0000
fal-3d
Generates 3D models from text or images via fal.ai, useful for game assets, AR previews, product mockups, and concept sculpting.
1
lucaspmarie-a11y
imagen
Generates images from text prompts using Google Gemini's image generation model, saving them as PNG files for use in UI, documentation, and design projects.
5
samuraigpt
muapi-one-shot-video
Generate a single continuous cinematic shot video with no cuts, using text or image prompts and configurable style, duration, and aspect ratio.
3.7k
oyi77
video-gen
Generates short AI videos from text or images across Runway, Kling, Sora, Pika, Seedance, Grok Imagine, and Veo, with provider-specific APIs and prompt guidance.
10
samuraigpt
muapi-keyboard-art-maker
Generate artistic top-down photos of keyboard keycaps arranged to spell out custom text messages.
3.7k
nvidia
tao-finetune-clip
Fine-tune and deploy CLIP vision-language models for zero-shot classification, image-text retrieval, and embedding extraction with ONNX and TensorRT support.
2.2k · bundle
demerzels-lab
fal
Search, explore, and run fal.ai generative AI models for image, video, audio, and 3D generation directly from the command line.
10 · bundle
oyi77
image-gen
Generates images from text prompts using diffusion models, covering prompt engineering, inpainting/outpainting, ControlNet, and API integration for production workflows.
10
johnalbertini14-glitch
fal
Search, explore, and run fal.ai generative AI models for image, video, audio, and 3D generation, including schema lookup, job submission, status polling, result retrieval, and file uploads.
1 · bundle
inference-sh
happyhorse
Generate and edit videos using Alibaba HappyHorse 1.0 models via the inference.sh CLI, supporting text-to-video, image-to-video, reference-to-video, and video editing with natural language.
584
microsoft
azure-ai-vision-imageanalysis-py
Analyze images using Azure AI Vision SDK: generate captions, tags, detect objects, extract text (OCR), detect people, and suggest smart crops.
2.7k
composiohq
openai-automation
Automate OpenAI API operations: generate text and multimodal responses with structured output, create embeddings, generate images, and list models via the Composio MCP integration.
66.9k
inference-sh
p-video
Generate videos using Pruna's optimized P-Video and WAN models via the inference.sh CLI, supporting text-to-video, image-to-video, audio input, and multiple resolutions.
584
inference-sh
seedance
Generate videos with synchronized audio using ByteDance Seedance 2.0 via the inference.sh CLI, supporting text-to-video, image-to-video, and reference-to-video modes up to 1080p.
584
vikingokft
gemini-interactions-api
Writes Python and TypeScript code that calls the Gemini Interactions API for text generation, chat, multimodal understanding, image generation, streaming, research, function calling, and structured output, including migration from the legacy generateContent API.
0 · bundle
google-gemini
gemini-omni-flash-api
Generate and edit videos using the Gemini Omni Flash model: text-to-video, image-to-video, video editing, and turn-by-turn refinement via the official google-genai SDK.
3.8k · bundle
nvidia
tao-train-ocrnet
Trains, evaluates, exports, prunes, quantizes, retrains, and runs inference for TAO OCRNet models for scene text recognition from cropped text-region images, supporting CTC and attention-based decoders.
2.2k · bundle