Results for “image-text-retrieval”
64 skillsclip
Enables zero-shot image classification, image-text matching, and cross-modal retrieval using OpenAI's CLIP model.
10.4k · bundle
blip-2-vision-language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
tao-finetune-clip
Fine-tune and deploy CLIP vision-language models for zero-shot classification, image-text retrieval, and embedding extraction with ONNX and TensorRT support.
2.2k · bundle
fal
Search, explore, and run fal.ai generative AI models for image, video, audio, and 3D generation, including schema lookup, job submission, status polling, result retrieval, and file uploads.
32 · bundle
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
3 · bundle
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
1 · bundle
More results
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
context-retrieval
Retrieve relevant information from a knowledge base using semantic, keyword, or hybrid search to ground a query. Use when the task starts with a corpus or index that must be searched; use context-ranking when candidate chunks already exist and only need ordering.
159
embeddings
Explains dense vector embeddings, their key concepts, common use cases, and best practices for semantic search and RAG applications.
1
synthtext-synthetic-data-for-text-detection-arxiv-1604-06646
SynthText: Synthetic Data for Text Detection
6
nemo-retriever
Index folders of PDFs and other documents into LanceDB for vector search, then query them with semantic search, page filters, verbatim quotes, and cross-document aggregation.
2.2k · bundle
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
gifgrep
Search GIF providers with CLI/TUI, download results, and extract stills/sheets.
228
dall-e-zero-shot-text-to-image-generation-arxiv-2102-12092v2
DALL-E: Zero-Shot Text-to-Image Generation
6
textvqa-towards-reasoning-about-text-in-images-arxiv-1904-08
TextVQA: Towards Reasoning about Text in Images
6
gifgrep
Search GIF providers with CLI/TUI, download results, and extract stills/sheets.
0
gifgrep
Search GIF providers with CLI/TUI, download results, and extract stills/sheets.
0
ivx-om-dashscope
DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans). Use when generating images via Qwen-Image, narrating via Qwen-TTS, or transcribing with word-level timestamps via Qwen-ASR.
0 · bundle
flamingo-a-visual-language-model-for-few-shot-learning-arxiv
Flamingo: A Visual Language Model for Few-Shot Learning
6
ocr-and-documents
Extract text from PDFs/scans (pymupdf, marker-pdf).
0 · bundle
ocr-and-documents
Extracts text from PDFs and scanned documents, choosing the cheapest method that works, from direct file reads to full OCR.
2 · bundle
gifgrep
Search GIF providers with CLI/TUI, download results, and extract stills/sheets.
0
capsfusion-rethinking-image-text-data-at-scale-arxiv-2310-20
CapsFusion: Rethinking Image-Text Data at Scale
6
google-image-api-skill
Extracts structured image metadata from Google Images via the BrowserAct API, including thumbnails, titles, source logos, and click-through URLs.
3.7k · bundle
nano-banana-2
Generate images using Google Gemini 3.1 Flash Image Preview via the inference.sh CLI, with support for text-to-image, image editing, multi-image input, and Google Search grounding.
584
rag
Build and debug Retrieval-Augmented Generation pipelines — chunking, embedding, retrieval, reranking
1 · bundle
baoyu-danger-gemini-web
Generates images and text via reverse-engineered Gemini Web API, supporting reference images and multi-turn conversations.
23.1k · bundle
ivx-imagegen
Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts. Use when Codex should create a brand-new image, transform an existing image, or derive visual variants from references, and the output should be a bitmap asset rather than repo-native code or vector. Do not use when the task is better handled by editing existing SVG/vector/code-native assets, extending an established icon or logo system, or building the visual directly in HTML/CSS/canvas.
0 · bundle
meme-maker
Search meme templates, suggest formats, and generate local or hosted image memes.
0 · bundle
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
image-upscaling
Upscale and enhance images using Real-ESRGAN, Thera, FLUX Upscaler, and Topaz via the inference.sh CLI.
584
sigmoid-loss-for-language-image-pre-training-arxiv-2303-1534
Sigmoid Loss for Language Image Pre-Training
6
imagen
|
6
image-seo
Image Seo — Skill especializada para otimizar imagens para mecanismos de busca, melhorando a visibilidade e performance.
55