Results for “vision-doc”

12 skills
More results
nvidia
Tao Train Foundation Stereo
Trains, evaluates, exports, and runs inference on FoundationStereo models for stereo depth estimation and 3D reconstruction from stereo image pairs.
2.2k · bundle
huggingface
Huggingface Vision Trainer
Trains and fine-tunes vision models for object detection, image classification, and segmentation using Hugging Face Transformers on cloud GPUs, with automatic dataset validation and Hub persistence.
10.8k · bundle
orchestra-research
Blip 2 Vision Language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
auto-skiller
Videodb
Ingest, index, search, edit, and generate video and audio assets from files, URLs, RTSP feeds, or desktop capture, with real-time alerts and stream links.
1 · bundle
sakamoto-family-smile
Videodb
Ingests video and audio from files, URLs, RTSP feeds, or desktop capture; indexes and searches moments with timestamps; transcodes, edits timelines, generates media assets, and creates real-time alerts for live streams.
0 · bundle
mhassan0000
Videodb
Ingests video and audio from files, URLs, live feeds, or desktop capture; indexes and searches moments with timestamps; transcodes, edits timelines, generates media assets, and emits real-time alerts.
1 · bundle
zhixuli0406
DOCX
Read, create, and convert Microsoft Word (.docx) documents — extract text and tables, build reports from markdown/JSON, and export to PDF.
45 · bundle
jiachen-t-wang
Donut Document Understanding Transformer Without Ocr Arxiv 2
Donut: Document Understanding Transformer without OCR
6
nvidia
Tao Train Depth Anything V2
Train, evaluate, export, and run inference for monocular depth estimation models using Metric Depth Anything v2 or Relative Depth Anything architectures via the TAO toolkit.
2.2k · bundle
qhjqhj00
Posh
Evaluates automated metrics and vision-language models on identifying granular errors in detailed image descriptions and ranking paired descriptions against human judgments, using macro F1, pairwise accuracy, Spearman rank ρ, and Kendall's τ.
3