Results for “robotic-vision”
51 skillsMore results
senior-computer-vision
Computer vision engineering skill for object detection, image segmentation, and visual AI systems. Covers CNN and Vision Transformer architectures, YOLO/Faster R-CNN/DETR detection, Mask R-CNN/SAM segmentation, and production deployment with ONNX/TensorRT. Includes PyTorch, torchvision, Ultralytics, Detectron2, and MMDetection frameworks. Use when building detection pipelines, training custom models, optimizing inference, or deploying vision systems.
1 · bundle
huggingface-vision-trainer
Trains and fine-tunes vision models for object detection, image classification, and segmentation using Hugging Face Transformers on cloud GPUs, with automatic dataset validation and Hub persistence.
10.8k · bundle
deepstream-import-vision-model
Import object detection models from HuggingFace or NVIDIA NGC into a DeepStream pipeline with automated ONNX download, TensorRT engine build, custom parser, multi-stream benchmark, and PDF report generation.
2.2k · bundle
vss-deploy-detection-tracking-3d
Deploy and operate the RTVI-CV-3D microservice for multi-camera 3D detection and tracking, supporting sample datasets, custom videos, and RTSP streams.
2.2k · bundle
coco-microsoft-coco-common-objects-in-context-arxiv-1405-031
COCO: Microsoft COCO: Common Objects in Context
6
visor
Evaluates text-to-image models on spatial relationship accuracy using the VISOR metric, separating object detection from spatial correctness to reveal biases like object priority and merging.
3
robotics-v3-ia
Expert en robotique avancée (ROS, SLAM, manipulation, autonomous navigation, DZ context)
6
blip-2-vision-language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
blip-2-vision-language
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
0 · bundle
blip-2-vision-language
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
1 · bundle
video-processing
This skill provides guidance for video analysis and processing tasks using computer vision techniques. It should be used when analyzing video frames, detecting motion or events, tracking objects, extracting temporal data (e.g., identifying specific frames like takeoff/landing moments), or performing frame-by-frame processing with OpenCV or similar libraries.
1
llava-onevision-easy-visual-task-transfer-arxiv-2408-03326v2
LLaVA-OneVision: Easy Visual Task Transfer
6
vision-analyze
画像を理解(被写体・テキストOCR・構図・色・UI構造の分析)し、結果を構造化して返すスキル。CC CLI は GLM-5.3 等の vision 非対応モデルで稼働中のため画像を直接視認できず、主ルート Gemini 2.5 Flash(scripts/api/gemini_vision.py・無料枠)と副ルート 4_5v MCP(analyze_image・Readが返すCDN URL)の2経路で分析し、CCは結果の構造化・比較・保存に専任する。 ユーザーが「画像見て」「この画像何が写ってる」「画像比較して」「スクショ見て」「画像分析して」「画像理解」「vision-analyze」と言った時、または /vision-analyze を呼んだ時にトリガー。 ※画像生成(image generation)は対象外(make-song / video-prompt-spec / demo-site-sales参照)。ピクセル修正(花鈿除去等)は remove-huadian の役割。楽曲分析は analyze-song / reverse-engineer-song。
0
ros-robotics
Develop, migrate, and debug ROS 1 and ROS 2 robotics projects with support for build systems, navigation, control, simulation, and embedded integration.
54 · bundle
vue-debug-guides
Diagnose and fix Vue 3 runtime errors, warnings, async failures, and SSR/hydration issues with targeted debugging guides.
2.7k · bundle
tao-train-bevfusion
Trains, evaluates, and runs inference for BEVFusion multi-sensor 3D object detection models that fuse LiDAR and camera data in bird's-eye-view space for autonomous driving.
2.2k · bundle
visual-plan
Transform text plans into interactive visual documents with diagrams, code snippets, and review surfaces for coding agents.
3.4k · bundle
nlvr2-a-visual-reasoning-benchmark-for-natural-language-arxi
NLVR2: A Visual Reasoning Benchmark for Natural Language
6
grit-general-robust-image-task-benchmark-arxiv-2306-14818v2
Grit: General Robust Image Task Benchmark
6
virtualized-lists
Render long lists (1k+ rows) efficiently in Salesforce LWC using lightning-datatable infinite loading or an IntersectionObserver sentinel, with keyset pagination and accessibility guidance.
15 · bundle
visual-qa-agent
Agent profile for inspect screenshots, viewports, layout overlaps, visual regressions, spacing, typography, and interaction states. Use when Codex needs a specialist agent perspective for planning, implementation, review, debugging, validation, or handoff in this domain.
1 · bundle
nocaps-novel-object-captioning-at-scale-arxiv-1812-08658v2
Nocaps: Novel Object Captioning at Scale
6
open-vocabulary-object-detection-using-captions-arxiv-2011-1
Open-Vocabulary Object Detection Using Captions
6
multimodal-few-shot-learning-with-frozen-language-models-arx
Multimodal Few-Shot Learning with Frozen Language Models
6
scaling-vision-transformers-to-22-billion-parameters-arxiv-2
Scaling Vision Transformers to 22 Billion Parameters
6
muapi-drone-style-video
Generate aerial drone-perspective footage with sweeping bird's-eye views, orbit shots, and flyover sequences for landscapes, architecture, and events.
3.7k
mosaic-augmentation-for-detection-and-segmentation-arxiv-yol
Mosaic Augmentation for Detection and Segmentation
6
videodb
See, Understand, Act on video and audio. See- ingest from local files, URLs, RTSP/live feeds, or live record desktop; return realtime context and playable stream links. Understand- extract frames, build visual/semantic/temporal indexes, and search moments with timestamps and auto-clips. Act- transcode and normalize (codec, fps, resolution, aspect ratio), perform timeline edits (subtitles, text/image overlays, branding, audio overlays, dubbing, translation), generate media assets (image, audio, video), and create real time alerts for events from live streams or desktop capture.
0 · bundle
longva-long-context-transfer-from-language-to-vision-arxiv-2
LongVA: Long Context Transfer from Language to Vision
6
video-watching
Watch and understand video files by converting them to viewable image storyboards for enjoyment, analysis, species identification, behavior tracking, and comprehension
1
llava
Enables visual instruction tuning and image-based conversations using open-source vision-language models. Supports multi-turn image chat, visual question answering, and image understanding tasks.
10.4k · bundle
llava
Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.
0 · bundle
data-visualization
Create clear, effective charts and dashboards from structured data using matplotlib, seaborn, and plotly. Use when the user requests data visualization or provides relevant inputs for this workflow.
159
scaling-vision-with-sparse-mixture-of-experts-arxiv-2106-059
Scaling Vision with Sparse Mixture of Experts
6
llava
Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.
0 · bundle