Packs

1 pack

Results for “caption”

68 skills
More results
jiachen-t-wang
dreamlip-language-image-pre-training-with-long-captions-arxi
DreamLIP: Language-Image Pre-training with Long Captions
6
qhjqhj00
polos
Scores generated image captions against reference captions and source images using the Polos metric, which is trained to align with human judgments and probes hallucination robustness and open-vocabulary evaluation.
3
composiohq
castingwords-automation
Automate Castingwords transcription and captioning tasks through Composio's toolkit via Rube MCP.
66.9k
nvidia
vss-deploy-dense-captioning
Deploy a standalone RT-VLM dense-captioning microservice and exercise its REST API endpoints for file upload, caption generation, streaming, chat completions, and Kafka integration.
2.2k · bundle
composiohq
amara-automation
Automate Amara subtitle and captioning operations through Composio's Amara toolkit via Rube MCP.
66.9k
jiachen-t-wang
dall-e-3-improving-image-generation-with-better-captions-arx
DALL-E 3: Improving Image Generation with Better Captions
6
infinition
youtube-transcribe
Fetches the transcript of a YouTube video as text or JSON using the video's caption track.
2 · bundle
herdiansah
social-caption-writer
Write platform-specific social media captions that drive engagement and conversions. Use when the user needs compelling written content for social posts.
23
qhjqhj00
spice
Evaluates image captions by converting them into scene graphs and computing an F-score over semantic propositions, measuring how well a generated caption captures the meaning of an image compared to human references.
3
heygen
changelog-video
Turn a weekly changelog markdown into a branded square video with animated UI mockups, voiceover, and captions.
· bundle
fukukei23
embedded-captions
Add captions or subtitles to an existing single-subject talking-head video without editing the footage. Use for plain verbatim captions, cinematic captions embedded behind the subject, VFX captions, “炸/特效/酷炫字幕,” or a named identity from the 35-style catalog. Route by visual identity, not by backend engine. The quiet `anchor` rail is the default; embed every word only when the user explicitly wants a fully cinematic treatment. The workflow runs locally end to end, including transcription and subject matting; split multi-shot footage before applying it.
0 · bundle
auto-skiller
video-production
Provides domain-specific knowledge for working with Remotion, covering 3D, animations, assets, audio, captions, and more.
1 · bundle
thedixitjain
llava
Vision-language chat: VQA, captioning, image dialogue.
2 · bundle
pwdev-solucoes
export-handoff
Exports approved Figma assets and assembles the delivery package with named files, captions, alt text, and publishing instructions.
2
jiachen-t-wang
capsfusion-rethinking-image-text-data-at-scale-arxiv-2310-20
CapsFusion: Rethinking Image-Text Data at Scale
6
remotion-dev
remotion-best-practices
Create and edit video compositions in React using Remotion, with guidance on layout, animation, assets, captions, effects, and rendering.
3.9k · bundle
jiachen-t-wang
coco-microsoft-coco-common-objects-in-context-arxiv-1405-031
COCO: Microsoft COCO: Common Objects in Context
6
samuraigpt
muapi-instagram-post
Generates a polished Instagram post with a hero image, caption, and hashtags based on a brief and brand style.
3.7k
microsoft
azure-ai-vision-imageanalysis-java
Analyze images using Azure AI Vision SDK for Java, enabling captioning, OCR, object detection, tagging, and smart cropping.
2.7k · bundle
heygen
talking-head-recut
Packages an existing talking-head, interview, or podcast video with timed, designed graphic overlay cards—kinetic titles, lower-thirds, data callouts, quotes, side panels, picture-in-picture—synced to the transcript, on a 16:9, 9:16, or 4:5 canvas.
· bundle
nvidia
tao-generate-image-grounding
Generates phrase-grounded bounding box annotations from image-caption pairs using a VLM, producing cleaned captions, referring expressions, and pixel-space bounding boxes.
2.2k · bundle
samuraigpt
muapi-social-pack
Re-render a hero image into aspect ratios for Instagram, TikTok, YouTube Shorts, and Twitter/X.
3.7k
jiachen-t-wang
laclip-improving-clip-training-with-language-rewrites-arxiv-
LaCLIP: Improving CLIP Training with Language Rewrites
6
nexu-io
social-reddit-card
Renders a story, question, or meme as a realistic Reddit post card with vote rail, comments, and awards, suitable for video overlays or social media sharing.
· bundle
nexu-io
social-x-post-card
Renders tweet content as a realistic X (Twitter) post card image for video overlays or image sharing, with interactive metrics and customizable themes.
· bundle
nexu-io
card-twitter
Generates a Twitter share card with a hero quote, author attribution, category tag, and subtle texture, ready for screenshot and posting.
· bundle
upayanghosh
synapse-image-describe
Provides detailed, structured image descriptions covering objects, people, colors, text, and scene context, with an overview and interpretation.
14
nexu-io
frame-light-leak-cinema
Generates a single-frame HTML template with a cinematic film-leak aesthetic: warm light leaks, 35mm grain, 2.39:1 letterbox, and serif typography for opening titles or chapter cards.
· bundle
browser-act
instagram-profile-posts
Scrapes posts from an Instagram user's profile feed including captions, media URLs, like/comment counts, timestamps, and location tags using cursor-based pagination.
3.7k · bundle
jiachen-t-wang
sigmoid-loss-for-language-image-pre-training-arxiv-2303-1534
Sigmoid Loss for Language Image Pre-Training
6