Results for “visual-differentiation”

30 skills
More results
lionelndong
visuals-adversarial
Skeptical pushback on visual placement — both density and quality. Reads the annotated outline plus the visuals manifest and asks (a) whether the article hits the density target from editorial-principles-visuals.md, (b) whether each [VISUAL:...] earns its place, (c) whether sections without one would benefit. One revision pass on FAIL (BLOG_AGENT_VISUALS_REVISION_BUDGET, default 1).
0
jiachen-t-wang
coco-microsoft-coco-common-objects-in-context-arxiv-1405-031
COCO: Microsoft COCO: Common Objects in Context
6
jiachen-t-wang
masked-autoencoders-are-scalable-vision-learners-arxiv-2111-
Masked Autoencoders Are Scalable Vision Learners
6
tangchunwu
visual-verdict
Structured visual QA verdict for screenshot-to-reference comparisons
1
jiachen-t-wang
nlvr2-a-visual-reasoning-benchmark-for-natural-language-arxi
NLVR2: A Visual Reasoning Benchmark for Natural Language
6
jiachen-t-wang
towards-open-world-segmentation-of-parts-arxiv-2305-06914v3
Towards Open-World Segmentation of Parts
6
jiachen-t-wang
vila-on-pre-training-for-visual-language-models-arxiv-2312-0
VILA: On Pre-training for Visual Language Models
6
lionelndong
visual-package
Build a visual sequence that proves, explains, and supports decisions.
0
jiachen-t-wang
open-vocabulary-object-detection-using-captions-arxiv-2011-1
Open-Vocabulary Object Detection Using Captions
6
jiachen-t-wang
mosaic-augmentation-for-detection-and-segmentation-arxiv-yol
Mosaic Augmentation for Detection and Segmentation
6
qcmuu
blip-2-vision-language
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
0 · bundle
jiachen-t-wang
nocaps-novel-object-captioning-at-scale-arxiv-1812-08658v2
Nocaps: Novel Object Captioning at Scale
6
jiachen-t-wang
multimodal-few-shot-learning-with-frozen-language-models-arx
Multimodal Few-Shot Learning with Frozen Language Models
6
tianhao909
blip-2-vision-language
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
1 · bundle
q2805187159
chart-visualization
This skill should be used when the user wants to visualize data. It intelligently selects the most suitable chart type from 26 available options, extracts parameters based on detailed specifications, and generates a chart image using a JavaScript script.
3 · bundle
omer-metin
vfx-realtime
Expert real-time VFX artist specializing in particle systems, shader effects, and the invisible craft that makes games feel satisfying. Masters Niagara, VFX Graph, Godot GPU particles, and understands the AAA principles that make effects read clearly at 60fps. Use when "particle system, visual effects, vfx, particles, niagara, vfx graph, flipbook, sprite sheet, explosion effect, magic effect, trail effect, beam effect, dissolve, distortion, force field, hit effect, muzzle flash, impact effect, smoke particles, fire effect, soft particles, game juice, screen shake, particle overdraw, effect optimization, vfx, particles, effects, niagara, vfx-graph, game-juice, visual-effects, shaders, flipbook, trails, beams, explosions, optimization, gpu-particles" mentioned.
128 · bundle
jiachen-t-wang
visual-instruction-tuning-arxiv-2304-08485v2
Visual Instruction Tuning
6
jiachen-t-wang
longva-long-context-transfer-from-language-to-vision-arxiv-2
LongVA: Long Context Transfer from Language to Vision
6
orchestra-research
blip-2-vision-language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
dromlakhani
bppv-differential-diagnosis
BPPV Differential Diagnosis Checker
10
salacoste
visual-verdict
Structured visual QA verdict for screenshot-to-reference comparisons
1
owl-listener
data-visualization
Design clear, accessible data visualizations with appropriate chart selection and styling.
1.7k
jiachen-t-wang
mantis-interleaved-multi-image-instruction-tuning-arxiv-2405
Mantis: Interleaved Multi-Image Instruction Tuning
6
timlai666
video-processing
This skill provides guidance for video analysis and processing tasks using computer vision techniques. It should be used when analyzing video frames, detecting motion or events, tracking objects, extracting temporal data (e.g., identifying specific frames like takeoff/landing moments), or performing frame-by-frame processing with OpenCV or similar libraries.
1
jiachen-t-wang
svit-scaling-up-visual-instruction-tuning-arxiv-2307-04087v2
SVIT: Scaling up Visual Instruction Tuning
6
jiachen-t-wang
llava-next-improved-reasoning-ocr-and-world-knowledge-arxiv-
LLaVA-NeXT: Improved Reasoning, OCR, and World Knowledge
6
jiachen-t-wang
glip-grounded-language-image-pre-training-arxiv-2112-03857v2
GLIP: Grounded Language-Image Pre-training
6
jiachen-t-wang
synthtext-synthetic-data-for-text-detection-arxiv-1604-06646
SynthText: Synthetic Data for Text Detection
6
ekatasingh1107
visual-brand
Extract brand visual system from a website or brand assets into a structured style guide
2 · bundle