Packs
1 packResults for “caption”
68 skillsremotion-captions
Provides a Caption type definition and links to skills for transcribing, displaying, and importing captions in Remotion videos.
3.9k · bundle
captions-overlay
Defines the caption model (drop/rail/embed) and overlay law for compositing captions on top of video, never reserving a bottom band.
sbu-captions-dataset-crossref-nips-2011-sbu
SBU Captions Dataset
6
embedded-captions
Add captions or subtitles to a single-subject talking-head video without editing the footage, using a catalog of visual identities.
· bundle
nocaps-novel-object-captioning-at-scale-arxiv-1812-08658v2
Nocaps: Novel Object Captioning at Scale
6
open-vocabulary-object-detection-using-captions-arxiv-2011-1
Open-Vocabulary Object Detection Using Captions
6
More results
dreamlip-language-image-pre-training-with-long-captions-arxi
DreamLIP: Language-Image Pre-training with Long Captions
6
polos
Scores generated image captions against reference captions and source images using the Polos metric, which is trained to align with human judgments and probes hallucination robustness and open-vocabulary evaluation.
3
castingwords-automation
Automate Castingwords transcription and captioning tasks through Composio's toolkit via Rube MCP.
66.9k
vss-deploy-dense-captioning
Deploy a standalone RT-VLM dense-captioning microservice and exercise its REST API endpoints for file upload, caption generation, streaming, chat completions, and Kafka integration.
2.2k · bundle
amara-automation
Automate Amara subtitle and captioning operations through Composio's Amara toolkit via Rube MCP.
66.9k
dall-e-3-improving-image-generation-with-better-captions-arx
DALL-E 3: Improving Image Generation with Better Captions
6
youtube-transcribe
Fetches the transcript of a YouTube video as text or JSON using the video's caption track.
2 · bundle
social-caption-writer
Write platform-specific social media captions that drive engagement and conversions. Use when the user needs compelling written content for social posts.
23
spice
Evaluates image captions by converting them into scene graphs and computing an F-score over semantic propositions, measuring how well a generated caption captures the meaning of an image compared to human references.
3
changelog-video
Turn a weekly changelog markdown into a branded square video with animated UI mockups, voiceover, and captions.
· bundle
embedded-captions
Add captions or subtitles to an existing single-subject talking-head video without editing the footage. Use for plain verbatim captions, cinematic captions embedded behind the subject, VFX captions, “炸/特效/酷炫字幕,” or a named identity from the 35-style catalog. Route by visual identity, not by backend engine. The quiet `anchor` rail is the default; embed every word only when the user explicitly wants a fully cinematic treatment. The workflow runs locally end to end, including transcription and subject matting; split multi-shot footage before applying it.
0 · bundle
video-production
Provides domain-specific knowledge for working with Remotion, covering 3D, animations, assets, audio, captions, and more.
1 · bundle
llava
Vision-language chat: VQA, captioning, image dialogue.
2 · bundle
export-handoff
Exports approved Figma assets and assembles the delivery package with named files, captions, alt text, and publishing instructions.
2
capsfusion-rethinking-image-text-data-at-scale-arxiv-2310-20
CapsFusion: Rethinking Image-Text Data at Scale
6
remotion-best-practices
Create and edit video compositions in React using Remotion, with guidance on layout, animation, assets, captions, effects, and rendering.
3.9k · bundle
coco-microsoft-coco-common-objects-in-context-arxiv-1405-031
COCO: Microsoft COCO: Common Objects in Context
6
muapi-instagram-post
Generates a polished Instagram post with a hero image, caption, and hashtags based on a brief and brand style.
3.7k
azure-ai-vision-imageanalysis-java
Analyze images using Azure AI Vision SDK for Java, enabling captioning, OCR, object detection, tagging, and smart cropping.
2.7k · bundle
talking-head-recut
Packages an existing talking-head, interview, or podcast video with timed, designed graphic overlay cards—kinetic titles, lower-thirds, data callouts, quotes, side panels, picture-in-picture—synced to the transcript, on a 16:9, 9:16, or 4:5 canvas.
· bundle
tao-generate-image-grounding
Generates phrase-grounded bounding box annotations from image-caption pairs using a VLM, producing cleaned captions, referring expressions, and pixel-space bounding boxes.
2.2k · bundle
muapi-social-pack
Re-render a hero image into aspect ratios for Instagram, TikTok, YouTube Shorts, and Twitter/X.
3.7k
laclip-improving-clip-training-with-language-rewrites-arxiv-
LaCLIP: Improving CLIP Training with Language Rewrites
6
social-reddit-card
Renders a story, question, or meme as a realistic Reddit post card with vote rail, comments, and awards, suitable for video overlays or social media sharing.
· bundle
social-x-post-card
Renders tweet content as a realistic X (Twitter) post card image for video overlays or image sharing, with interactive metrics and customizable themes.
· bundle
card-twitter
Generates a Twitter share card with a hero quote, author attribution, category tag, and subtle texture, ready for screenshot and posting.
· bundle
synapse-image-describe
Provides detailed, structured image descriptions covering objects, people, colors, text, and scene context, with an overview and interpretation.
14
frame-light-leak-cinema
Generates a single-frame HTML template with a cinematic film-leak aesthetic: warm light leaks, 35mm grain, 2.39:1 letterbox, and serif typography for opening titles or chapter cards.
· bundle
instagram-profile-posts
Scrapes posts from an Instagram user's profile feed including captions, media URLs, like/comment counts, timestamps, and location tags using cursor-based pagination.
3.7k · bundle
sigmoid-loss-for-language-image-pre-training-arxiv-2303-1534
Sigmoid Loss for Language Image Pre-Training
6