Results for “image-description”
18 skillssynapse-image-describe
Provides detailed, structured image descriptions covering objects, people, colors, text, and scene context, with an overview and interpretation.
14
cider
Computes CIDEr and related metrics to score how well generated image descriptions align with human consensus, using reference sentences and triplet annotations.
3
More results
posh
Evaluates automated metrics and vision-language models on identifying granular errors in detailed image descriptions and ranking paired descriptions against human judgments, using macro F1, pairwise accuracy, Spearman rank ρ, and Kendall's τ.
3
nano-banana-2
Generate images using Google Gemini 3.1 Flash Image Preview via the inference.sh CLI, with support for text-to-image, image editing, multi-image input, and Google Search grounding.
584
blip-2-vision-language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
vision-analyze
画像を理解(被写体・テキストOCR・構図・色・UI構造の分析)し、結果を構造化して返すスキル。CC CLI は GLM-5.3 等の vision 非対応モデルで稼働中のため画像を直接視認できず、主ルート Gemini 2.5 Flash(scripts/api/gemini_vision.py・無料枠)と副ルート 4_5v MCP(analyze_image・Readが返すCDN URL)の2経路で分析し、CCは結果の構造化・比較・保存に専任する。 ユーザーが「画像見て」「この画像何が写ってる」「画像比較して」「スクショ見て」「画像分析して」「画像理解」「vision-analyze」と言った時、または /vision-analyze を呼んだ時にトリガー。 ※画像生成(image generation)は対象外(make-song / video-prompt-spec / demo-site-sales参照)。ピクセル修正(花鈿除去等)は remove-huadian の役割。楽曲分析は analyze-song / reverse-engineer-song。
0
muapi-keyboard-art-maker
Generate artistic top-down photos of keyboard keycaps arranged to spell out custom text messages.
3.7k
image-enhancer
Improves the quality of images, especially screenshots, by enhancing resolution, sharpness, and clarity. Perfect for preparing images for presentations, documentation, or social media posts.
3
rednote
将内容拆解为系列卡通风格信息图提示词,通过 Chrome DevTools MCP 自动化调用 Gemini web app 生成图片,并自动收集到本地临时目录。
2
image-upscaling
Upscale and enhance images using Real-ESRGAN, Thera, FLUX Upscaler, and Topaz via the inference.sh CLI.
584
clip
Enables zero-shot image classification, image-text matching, and cross-modal retrieval using OpenAI's CLIP model.
10.4k · bundle
vision-analysis
Analyze, describe, and extract information from images using the MiniMax vision MCP tool, with modes for general description, OCR, UI review, chart data extraction, and object detection.
12.9k
muapi-multi-angle-reshoot
Re-renders a subject or scene from multiple dramatic camera angles while maintaining consistent identity and detail.
3.7k
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
coyo-700m-image-text-pair-dataset-github-kakaobrain-coyo-700
COYO-700M: Image-Text Pair Dataset
6
spice
Evaluates image captions by converting them into scene graphs and computing an F-score over semantic propositions, measuring how well a generated caption captures the meaning of an image compared to human references.
3
sbu-captions-dataset-crossref-nips-2011-sbu
SBU Captions Dataset
6
character-design-sheet
Create consistent characters across AI-generated images using reference sheets, detailed descriptions, and LoRA techniques.
584