Results for “captioning”

22 skills
More results
microsoft
azure-ai-vision-imageanalysis-py
Analyze images using Azure AI Vision SDK: generate captions, tags, detect objects, extract text (OCR), detect people, and suggest smart crops.
2.7k
heygen
captions-overlay
Defines the caption model (drop/rail/embed) and overlay law for compositing captions on top of video, never reserving a bottom band.
qhjqhj00
polos
Scores generated image captions against reference captions and source images using the Polos metric, which is trained to align with human judgments and probes hallucination robustness and open-vocabulary evaluation.
3
herdiansah
social-caption-writer
Write platform-specific social media captions that drive engagement and conversions. Use when the user needs compelling written content for social posts.
23
jiachen-t-wang
sbu-captions-dataset-crossref-nips-2011-sbu
SBU Captions Dataset
6
qhjqhj00
spice
Evaluates image captions by converting them into scene graphs and computing an F-score over semantic propositions, measuring how well a generated caption captures the meaning of an image compared to human references.
3
nvidia
tao-generate-image-grounding
Generates phrase-grounded bounding box annotations from image-caption pairs using a VLM, producing cleaned captions, referring expressions, and pixel-space bounding boxes.
2.2k · bundle
upayanghosh
synapse-image-describe
Provides detailed, structured image descriptions covering objects, people, colors, text, and scene context, with an overview and interpretation.
14
jiachen-t-wang
flamingo-a-visual-language-model-for-few-shot-learning-arxiv
Flamingo: A Visual Language Model for Few-Shot Learning
6
orchestra-research
blip-2-vision-language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
aniruddhaadak80
llava
Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.
0 · bundle
muratcankoylan
context-optimization
Extends effective context capacity through strategic compression, masking, caching, and partitioning techniques.
16.9k · bundle
samuraigpt
muapi-keyboard-art-maker
Generate artistic top-down photos of keyboard keycaps arranged to spell out custom text messages.
3.7k
herdiansah
social-media-content-repurposer
Transform content across platforms with platform-specific optimization. Use when converting blog posts, videos, or articles into threads, captions, and posts for various social networks.
23
tools-only
151-copy-e139d00c
封装阶段3的Copy Spec为可批量执行的提示词包,生成JSONL请求并调用APIMart API出图。
7 · bundle
orchestra-research
llava
Enables visual instruction tuning and image-based conversations using open-source vision-language models. Supports multi-turn image chat, visual question answering, and image understanding tasks.
10.4k · bundle
kbarbel640-del
vgl
Generates structured VGL JSON prompts for Bria's FIBO image generation models, covering text-to-image, editing, inpainting, outpainting, and captioning with a deterministic schema.
1 · bundle