Results for “visual-question-answering”

18 skills
More results
nvidia
Vss Ask Video
Ask visual questions about video clips using a VSS agent's video_understanding tool, requiring a fresh look at frames rather than prior metadata or search results.
2.2k · bundle
nvidia
Vss Query Analytics
Queries video analytics incidents, alerts, metrics, and sensor data from Elasticsearch via the VA-MCP server.
2.2k · bundle
orchestra-research
Blip 2 Vision Language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
jiachen-t-wang
Snli Ve Visual Entailment Dataset Arxiv 1901 06706v1
SNLI-VE: Visual Entailment Dataset
6
peteedoo
Llava
Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.
0 · bundle
aniruddhaadak80
Llava
Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.
0 · bundle
qhjqhj00
Visor
Evaluates text-to-image models on spatial relationship accuracy using the VISOR metric, separating object detection from spatial correctness to reveal biases like object priority and merging.
3
jiachen-t-wang
Flamingo A Visual Language Model For Few Shot Learning Arxiv
Flamingo: A Visual Language Model for Few-Shot Learning
6
jimliu
Baoyu Article Illustrator
Analyzes article structure, identifies positions requiring visual aids, and generates illustrations with consistent type, style, and palette.
23.1k · bundle
gabrielmoreira
Vox Explainer
Produces a complete narrated, subtitled, scored explainer video from a single topic prompt using a six-stage pipeline with script, voiceover, keyframes, animation, music, and local assembly.
17 · bundle
fukukei23
Vision Analyze
画像を理解(被写体・テキストOCR・構図・色・UI構造の分析)し、結果を構造化して返すスキル。CC CLI は GLM-5.3 等の vision 非対応モデルで稼働中のため画像を直接視認できず、主ルート Gemini 2.5 Flash(scripts/api/gemini_vision.py・無料枠)と副ルート 4_5v MCP(analyze_image・Readが返すCDN URL)の2経路で分析し、CCは結果の構造化・比較・保存に専任する。 ユーザーが「画像見て」「この画像何が写ってる」「画像比較して」「スクショ見て」「画像分析して」「画像理解」「vision-analyze」と言った時、または /vision-analyze を呼んだ時にトリガー。 ※画像生成(image generation)は対象外(make-song / video-prompt-spec / demo-site-sales参照)。ピクセル修正(花鈿除去等)は remove-huadian の役割。楽曲分析は analyze-song / reverse-engineer-song。
0
jiachen-t-wang
Visual Prompt Tuning Arxiv 2203 12119v2
Visual Prompt Tuning
6
qhjqhj00
Vpeval
Evaluates text-to-image generation models by decomposing assessment into five specialized skills (object presence, count, spatial relations, scale, and text rendering) and open-ended prompts, producing interpretable binary scores with visual and textual explanations.
3
jiachen-t-wang
Docvqa A Dataset For Vqa On Document Images Arxiv 2007 00398
DocVQA: A Dataset for VQA on Document Images
6
alunadev
Animation Vocabulary
Reverse-lookup glossary that turns a vague description of a web animation or motion effect into its exact term ("the bouncy thing when a popover opens" → Pop in; "the iOS rubber-band scroll" → Rubber-banding). Use when the user asks "what's it called when…", or describes a motion effect without knowing its name and wants the right word to prompt an AI or designer with. For naming an effect, not designing or building one. Source: github.com/emilkowalski/skills.
3
vvieira010-pixel
Dual Coding Designer
Design a visual complement to verbal content using dual coding principles for stronger encoding. Use when creating slides, diagrams, posters, or visual explanations of complex concepts.
0