Results for “captioning”

32 skills
More results
tianhao909
blip-2-vision-language
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
1 · bundle
remotion-dev
remotion-captions
Provides a Caption type definition and links to skills for transcribing, displaying, and importing captions in Remotion videos.
3.9k · bundle
jiachen-t-wang
open-vocabulary-object-detection-using-captions-arxiv-2011-1
Open-Vocabulary Object Detection Using Captions
6
jiachen-t-wang
capsfusion-rethinking-image-text-data-at-scale-arxiv-2310-20
CapsFusion: Rethinking Image-Text Data at Scale
6
jiachen-t-wang
dreamlip-language-image-pre-training-with-long-captions-arxi
DreamLIP: Language-Image Pre-training with Long Captions
6
jiachen-t-wang
laclip-improving-clip-training-with-language-rewrites-arxiv-
LaCLIP: Improving CLIP Training with Language Rewrites
6
jiachen-t-wang
coco-microsoft-coco-common-objects-in-context-arxiv-1405-031
COCO: Microsoft COCO: Common Objects in Context
6
matrixx0070
i18n-subtitle
Create subtitles or captions with correct timing, line breaks, and reading-speed limits.
0
jiachen-t-wang
dall-e-3-improving-image-generation-with-better-captions-arx
DALL-E 3: Improving Image Generation with Better Captions
6
heygen
talking-head-recut
Packages an existing talking-head, interview, or podcast video with timed, designed graphic overlay cards—kinetic titles, lower-thirds, data callouts, quotes, side panels, picture-in-picture—synced to the transcript, on a 16:9, 9:16, or 4:5 canvas.
· bundle
qcmuu
blip-2-vision-language
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
0 · bundle
jiachen-t-wang
visual-instruction-tuning-arxiv-2304-08485v2
Visual Instruction Tuning
6
samuraigpt
muapi-social-pack
Re-render a hero image into aspect ratios for Instagram, TikTok, YouTube Shorts, and Twitter/X.
3.7k
nexu-io
frame-light-leak-cinema
Generates a single-frame HTML template with a cinematic film-leak aesthetic: warm light leaks, 35mm grain, 2.39:1 letterbox, and serif typography for opening titles or chapter cards.
· bundle
samuraigpt
muapi-youtube-thumbnail
Generate high-CTR YouTube thumbnails with striking imagery, bold text placement, and emotional subjects using AI image generation.
3.7k
lucian55
wujing-skill
吴京(演员 / 导演)认知与表达框架(压缩蒸馏):主旋律英雄壳、直男效能叙事 触发:战狼、流浪地球 等。非煽动仇恨
9 · bundle
orchestra-research
blip-2-vision-language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
mocchalera
finish-interview
Applies dialogue MA, portrait reframing, and loudness/sync QA to an interview or talking-head rough cut using canonical timeline metadata and a shared renderer.
3 · bundle
theycallmeholla
connotation-cop
Police the project's vocabulary — bust vague terms, keep the CONTEXT.md glossary sharp, and lock in decisions worth remembering as ADRs. Use when the user debates naming, says "what should we call this", asks to pin down terminology, wants a decision recorded, or when another skill (hot-seat, whiteboard) surfaces a decision that clears the ADR bar. Just reading the glossary for vocabulary is NOT this skill — trigger only when the words or decisions are being changed.
0 · bundle
muratcankoylan
context-optimization
Extends effective context capacity through strategic compression, masking, caching, and partitioning techniques.
16.9k · bundle
samuraigpt
muapi-keyboard-art-maker
Generate artistic top-down photos of keyboard keycaps arranged to spell out custom text messages.
3.7k
azusagasaku
manim-video
构建可复用的Manim解释器,用于技术概念、图表、系统图和产品演示,并在需要时移交给更广泛的ECC视频栈。当用户希望获得清晰的动画解释而非通用的人物讲解脚本时使用。
0 · bundle
lionelndong
visual-package
Build a visual sequence that proves, explains, and supports decisions.
0
dracounion
perspective-framing
当需要在写作中与读者建立情感连接、放大问题严重性,并引出后续解决方案时
11 · bundle
jiachen-t-wang
glip-grounded-language-image-pre-training-arxiv-2112-03857v2
GLIP: Grounded Language-Image Pre-training
6
dracounion
perspective-reframing
当需要从现有内容中创造新意,或避免内容同质化时
11 · bundle
aiweline
context-compression
上下文压缩省 Token。对话变长或开新任务时输出/使用 6 块压缩结构,总长 400~800 tokens。必加载。
1
orchestra-research
llava
Enables visual instruction tuning and image-based conversations using open-source vision-language models. Supports multi-turn image chat, visual question answering, and image understanding tasks.
10.4k · bundle
kbarbel640-del
vgl
Generates structured VGL JSON prompts for Bria's FIBO image generation models, covering text-to-image, editing, inpainting, outpainting, and captioning with a deterministic schema.
1 · bundle