Results for “vision-doc”

59 skills
intense-visions
Docs Craft
Docs Craft
18 · bundle
omer-metin
Document AI
Comprehensive patterns for AI-powered document understanding including PDF parsing, OCR, invoice/receipt extraction, table extraction, multimodal RAG with vision models, and structured data output. Use when "document parsing, PDF extraction, OCR, invoice processing, receipt extraction, document understanding, LlamaParse, Unstructured, vision document, table extraction, structured output from PDF, " mentioned.
128 · bundle
atc-net
Azure AI Vision
Expert knowledge for Azure AI Vision development including decision making, limits & quotas, configuration, integrations & coding patterns, and deployment. Use when using Image Analysis, Read OCR containers, smart-crop thumbnails, background removal, or video frame analysis, and other Azure AI Vision related development tasks. Not for Azure AI Custom Vision (use azure-custom-vision), Azure AI Video Indexer (use azure-video-indexer), Azure AI Document Intelligence (use azure-document-intelligence), Azure AI Immersive Reader (use azure-immersive-reader).
3
nvidia
Tao Port Huggingface Model
Integrate a HuggingFace computer vision model into the NVIDIA TAO Toolkit ecosystem, covering the full pipeline from prerequisites to container testing.
2.2k · bundle
peteedoo
Brain To Docs
Extract project vision and decisions from a user through iterative Q&A, then convert them into clear repo documentation (README, ADRs, lessons). Use when user asks to document ideas in their head.
0
More results
affaan-m
Visa Doc Translate
Translate visa application documents from images to English and generate a bilingual PDF with the original and translation side by side.
226k · bundle
jantoniofc
Doc
Use when the task involves reading, creating, or editing `.docx` documents, especially when formatting or layout fidelity matters; prefer `python-docx` plus the bundled `scripts/render_docx.py` for visual checks.
6 · bundle
johnalbertini14-glitch
Opentangl
Configures a self-driving development loop for JavaScript/TypeScript projects, generating configuration files and preparing OpenTangl to run autonomously.
1 · bundle
jarbitechture
Doc
Use when the task involves reading, creating, or editing `.docx` documents, especially when formatting or layout fidelity matters; prefer `python-docx` plus the bundled `scripts/render_docx.py` for visual checks.
0 · bundle
modbender
Doc
Use when the task involves reading, creating, or editing `.docx` documents, especially when formatting or layout fidelity matters; prefer `python-docx` plus the bundled `scripts/render_docx.py` for visual checks.
12 · bundle
jackychenlu
Doc
Use when the task involves reading, creating, or editing `.docx` documents, especially when formatting or layout fidelity matters; prefer `python-docx` plus the bundled `scripts/render_docx.py` for visual checks.
0 · bundle
builderio
Visual Plan
Transform text plans into interactive visual documents with diagrams, code snippets, and review surfaces for coding agents.
3.4k · bundle
stribus
Reversa Docs
Orquestrador do Time Reversa Docs. Gera um mini-site HTML autocontido em .reversa/documentation/ com arquitetura 3D, dashboards, glossário, deck e páginas por feature, a partir do conhecimento já extraído pelo core do Reversa. Ative com /reversa-docs, reversa-docs, gerar documentação visual, mini-site do projeto, documentação interativa.
1 · bundle
heath-gtm
Docs Ship Check
The pre-share gate for a Superhuman Docs (Coda) document. Before a doc reaches a reader, run the QA checklist, choose the interaction mode, set the publish frame, and confirm the share. Catches the broken formula, the wall of text, the orphaned bullet, the missing cover, the wrong audience. Trigger on "is this doc ready to share", "publish this doc", "QA this doc", "ship-check", "final pass before I send", "which interaction mode", or any moment a Superhuman Doc is about to leave the building.
0
jeffallan
Vue Expert JS
Builds Vue 3 components, composables, and Vite projects using JavaScript with JSDoc type annotations instead of TypeScript.
10.4k · bundle
joshuashepherd
Doc Organizer
Scans repositories for markdown and HTML documentation, classifies files as repo-local or centralizable content, generates an inventory report, and optionally consolidates content into a centralized docs repository.
1
jiachen-t-wang
Coco Microsoft Coco Common Objects In Context Arxiv 1405 031
COCO: Microsoft COCO: Common Objects in Context
6
qcmuu
Blip 2 Vision Language
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
0 · bundle
huggingface
Huggingface Vision Trainer
Trains and fine-tunes vision models for object detection, image classification, and segmentation using Hugging Face Transformers on cloud GPUs, with automatic dataset validation and Hub persistence.
10.8k · bundle
nvidia
Tao Train Foundation Stereo
Trains, evaluates, exports, and runs inference on FoundationStereo models for stereo depth estimation and 3D reconstruction from stereo image pairs.
2.2k · bundle
dataroaring
Docs Architect
Apply world-class developer documentation principles (Stripe, Snowflake, Databricks, TiDB Cloud) to structure, write, review, or refactor technical documentation. Use this skill whenever the user mentions documentation, docs, sidebar or navigation, information architecture, restructuring a section, writing or editing a guide, reviewing docs, where content belongs, English doc prose, headings, code comments, link text, docs home pages, section landing pages, long-form guides mixing content types, cross-referencing, or making docs readable for AI agents and LLMs. Covers VeloDB Cloud docs work (Monitoring restructure, sidebar, EN/中文 alignment, writing style, landing pages, LLM-friendly docs) and any SaaS or database documentation task. Trigger broadly: if the conversation touches doc organization, page structure, doc quality, doc sentences, landing pages, or AI-readable docs, consult this skill rather than answering from intuition.
0 · bundle
tangchunwu
Videodb
See, Understand, Act on video and audio. See- ingest from local files, URLs, RTSP/live feeds, or live record desktop; return realtime context and playable stream links. Understand- extract frames, build visual/semantic/temporal indexes, and search moments with timestamps and auto-clips. Act- transcode and normalize (codec, fps, resolution, aspect ratio), perform timeline edits (subtitles, text/image overlays, branding, audio overlays, dubbing, translation), generate media assets (image, audio, video), and create real time alerts for events from live streams or desktop capture.
1 · bundle
tianhao909
Blip 2 Vision Language
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
1 · bundle
moonladderstudios
Document Author
Author a new docs-native canonical or working document by choosing the correct location, filename, viewpoint template, metadata header, stable claims, and embedded rationale without creating spec.md.
12
srednoff888-art
Visual QA Agent
Agent profile for inspect screenshots, viewports, layout overlaps, visual regressions, spacing, typography, and interaction states. Use when Codex needs a specialist agent perspective for planning, implementation, review, debugging, validation, or handoff in this domain.
1 · bundle
jarbitechture
Videodb
See, Understand, Act on video and audio. See- ingest from local files, URLs, RTSP/live feeds, or live record desktop; return realtime context and playable stream links. Understand- extract frames, build visual/semantic/temporal indexes, and search moments with timestamps and auto-clips. Act- transcode and normalize (codec, fps, resolution, aspect ratio), perform timeline edits (subtitles, text/image overlays, branding, audio overlays, dubbing, translation), generate media assets (image, audio, video), and create real time alerts for events from live streams or desktop capture.
0 · bundle
orchestra-research
Blip 2 Vision Language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
minimax-ai
Minimax DOCX
Creates, edits, and formats DOCX documents using OpenXML SDK (.NET) with CLI tools and C# scripts, supporting creation, content editing, and template application.
12.9k · bundle
auto-skiller
Videodb
Ingest, index, search, edit, and generate video and audio assets from files, URLs, RTSP feeds, or desktop capture, with real-time alerts and stream links.
1 · bundle
brycewang-stanford
Data Doc
Document datasets, variables, sources, and merge keys for replication
1k
sakamoto-family-smile
Videodb
Ingests video and audio from files, URLs, RTSP feeds, or desktop capture; indexes and searches moments with timestamps; transcodes, edits timelines, generates media assets, and creates real-time alerts for live streams.
0 · bundle
mhassan0000
Videodb
Ingests video and audio from files, URLs, live feeds, or desktop capture; indexes and searches moments with timestamps; transcodes, edits timelines, generates media assets, and emits real-time alerts.
1 · bundle
joshuashepherd
Docs Setup
Bootstraps a two-part _docs directory (_build/ for technical docs, _public/ for research and proposals) with a CONSTITUTION.md governing the split, and can audit or migrate an existing structure.
1
heath-gtm
Meeting To Doc
Meeting to doc
0
jiachen-t-wang
Longva Long Context Transfer From Language To Vision Arxiv 2
LongVA: Long Context Transfer from Language to Vision
6
om-scogo
Videodb
See, Understand, Act on video and audio. See- ingest from local files, URLs, RTSP/live feeds, or live record desktop; return realtime context and playable stream links. Understand- extract frames, build visual/semantic/temporal indexes, and search moments with timestamps and auto-clips. Act- transcode and normalize (codec, fps, resolution, aspect ratio), perform timeline edits (subtitles, text/image overlays, branding, audio overlays, dubbing, translation), generate media assets (image, audio, video), and create real time alerts for events from live streams or desktop capture.
0 · bundle