Results for “image-captioning”

20 skills
More results
orchestra-research
llava
Enables visual instruction tuning and image-based conversations using open-source vision-language models. Supports multi-turn image chat, visual question answering, and image understanding tasks.
10.4k · bundle
upayanghosh
synapse-image-describe
Provides detailed, structured image descriptions covering objects, people, colors, text, and scene context, with an overview and interpretation.
14
jiachen-t-wang
sbu-captions-dataset-crossref-nips-2011-sbu
SBU Captions Dataset
6
heygen
captions-overlay
Defines the caption model (drop/rail/embed) and overlay law for compositing captions on top of video, never reserving a bottom band.
nvidia
tao-generate-image-grounding
Generates phrase-grounded bounding box annotations from image-caption pairs using a VLM, producing cleaned captions, referring expressions, and pixel-space bounding boxes.
2.2k · bundle
orchestra-research
blip-2-vision-language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
jiachen-t-wang
flamingo-a-visual-language-model-for-few-shot-learning-arxiv
Flamingo: A Visual Language Model for Few-Shot Learning
6
aniruddhaadak80
llava
Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.
0 · bundle
peteedoo
llava
Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.
0 · bundle
phoroth
imagen
Generates images from text prompts using Google Gemini's image generation model, saving them as PNG files for use in UI, documentation, and design projects.
3
peteedoo
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
antigravity
imagen
Generates images using Google Gemini's image generation model for UI placeholders, documentation, and design assets.
42.4k
orchestra-research
clip
Enables zero-shot image classification, image-text matching, and cross-modal retrieval using OpenAI's CLIP model.
10.4k · bundle
qhjqhj00
spice
Evaluates image captions by converting them into scene graphs and computing an F-score over semantic propositions, measuring how well a generated caption captures the meaning of an image compared to human references.
3
kairyou
at-vision
Inspect screenshots, photos, diagrams, image paths, and image URLs when the task depends on visible content. Use when the prompt lacks actual image content, native inspection fails, or the user requests inspect_image; prefer the MCP tool, then the installed CLI.
167
microsoft
azure-ai-vision-imageanalysis-py
Analyze images using Azure AI Vision SDK: generate captions, tags, detect objects, extract text (OCR), detect people, and suggest smart crops.
2.7k
qhjqhj00
polos
Scores generated image captions against reference captions and source images using the Polos metric, which is trained to align with human judgments and probes hallucination robustness and open-vocabulary evaluation.
3