Results for “image-text”
48 skillstextvqa-towards-reasoning-about-text-in-images-arxiv-1904-08
TextVQA: Towards Reasoning about Text in Images
6
capsfusion-rethinking-image-text-data-at-scale-arxiv-2310-20
CapsFusion: Rethinking Image-Text Data at Scale
6
dall-e-zero-shot-text-to-image-generation-arxiv-2102-12092v2
DALL-E: Zero-Shot Text-to-Image Generation
6
vgl
Generates structured VGL JSON prompts for Bria's FIBO image generation models, covering text-to-image, editing, inpainting, outpainting, and captioning with a deterministic schema.
1 · bundle
blip-2-vision-language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
threejs-textures
Three.js textures - texture types, UV mapping, environment maps, texture settings. Use when working with images, UV coordinates, cubemaps, HDR environments, or texture optimization.
0
More results
ltx2
AI video generation with LTX-2.3 22B — text-to-video, image-to-video clips for video production. Use when generating video clips, animating images, creating b-roll, animated backgrounds, or motion content. Triggers include video generation, animate image, b-roll, motion, video clip, text-to-video, image-to-video.
2
muapi-youtube-thumbnail
Generate high-CTR YouTube thumbnails with striking imagery, bold text placement, and emotional subjects using AI image generation.
3.7k
watermark-an-image
Apply a text watermark to a photo using layer-based image composition for brand protection and copyright.
2
ivx-om-ltx2
AI video generation with LTX-2.3 22B — text-to-video, image-to-video clips for video production. Use when generating video clips, animating images, creating b-roll, animated backgrounds, or motion content. Triggers include video generation, animate image, b-roll, motion, video clip, text-to-video, image-to-video.
0 · bundle
baoyu-cover-image
Generates article cover images with customizable type, palette, rendering, text, and mood dimensions, supporting multiple aspect ratios and backends.
23.1k · bundle
baoyu-image-gen
AI image generation with OpenAI, Google and DashScope APIs. Supports text-to-image, reference images, aspect ratios. Sequential by default; parallel generation available on request. Use when user asks to generate, create, or draw images.
1 · bundle
threejs-loaders
Three.js asset loading - GLTF, textures, images, models, async patterns. Use when loading 3D models, textures, HDR environments, or managing loading progress.
0
muapi-one-shot-video
Generate a single continuous cinematic shot video with no cuts, using text or image prompts and configurable style, duration, and aspect ratio.
3.7k
muapi-keyboard-art-maker
Generate artistic top-down photos of keyboard keycaps arranged to spell out custom text messages.
3.7k
gradio
Creates ML demo interfaces with Gradio, supporting images, text, audio, and video inputs/outputs.
2 · bundle
textme
Text Claude from your phone — set up the njerschow/textme daemon so inbound iMessages drive a Claude Code session on your laptop, with voice notes, image input, code execution, and a phone-number whitelist.
11
gpt-image2
Use when the user asks Codex to directly generate images with gpt-image-2 using inherited OpenAI/Codex-compatible environment credentials or local GPT_IMAGE2_* overrides, including text-to-image, reference-image guided generation, ratios, resolution, quality, variants, and saved local image files; run the bundled Node CLI and keep URL/sk configuration private.
65 · bundle
og-image-design
Design Open Graph and social sharing images with platform-specific specs, text placement, and branding guidelines. Generate images via HTML-to-image or AI, and configure OG meta tags for Facebook, Twitter, LinkedIn, and more.
584
synthtext-synthetic-data-for-text-detection-arxiv-1604-06646
SynthText: Synthetic Data for Text Detection
6
image-manipulation-image-magick
Process and manipulate images using ImageMagick: resize, convert formats, batch process, and retrieve metadata.
36.2k
generate-front-book-cover
Generate a front cover image with custom artwork, title text, and author attribution.
2
convert-image-format
Convert an image between PNG, JPEG, and WebP formats with quality control for web optimization.
2
vfx-text-cursor
Creates a video intro frame with a typing cursor effect that reveals text character by character, accompanied by chromatic aberration trails and directional light leaks.
· bundle
generate-image
Generate Image (gemini / openai backends)
580 · bundle
imagenet-a-large-scale-hierarchical-image-database-crossref-
ImageNet: A Large-Scale Hierarchical Image Database
6
blip-2-vision-language
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
1 · bundle
imagen
|
6
sigmoid-loss-for-language-image-pre-training-arxiv-2303-1534
Sigmoid Loss for Language Image Pre-Training
6
chameleon-mixed-modal-early-fusion-foundation-models-arxiv-2
Chameleon: Mixed-Modal Early-Fusion Foundation Models
6
social-x-post-card
Renders tweet content as a realistic X (Twitter) post card image for video overlays or image sharing, with interactive metrics and customizable themes.
· bundle
blip-2-vision-language
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
0 · bundle
visual-package
Build a visual sequence that proves, explains, and supports decisions.
0
imagine
Generate or edit images with Codex. Use this skill whenever the user says "imagine ...", asks to create an image from a text description, transform or restyle an existing image, produce artwork / illustrations / logos / concept art, make image variations, or asks for any kind of AI image generation or image-to-image editing. All outputs are saved inside the current project's `./images/` folder by default.
13 · bundle
ascii-art
Renders text or images as ASCII art for terminal-friendly output, including banners, frames, QR codes, and weather.
2
gemini-interactions-api
Writes Python and TypeScript code that calls the Gemini Interactions API for text generation, chat, multimodal understanding, image generation, streaming, research, function calling, and structured output, including migration from the legacy generateContent API.
0 · bundle