Results for “image-text-retrieval”

64 skills
More results
qcmuu
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
jackychenlu
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
bog5d
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
seb1n
context-retrieval
Retrieve relevant information from a knowledge base using semantic, keyword, or hybrid search to ground a query. Use when the task starts with a corpus or index that must be searched; use context-ranking when candidate chunks already exist and only need ordering.
159
neuralblitz
embeddings
Explains dense vector embeddings, their key concepts, common use cases, and best practices for semantic search and RAG applications.
1
jiachen-t-wang
synthtext-synthetic-data-for-text-detection-arxiv-1604-06646
SynthText: Synthetic Data for Text Detection
6
nvidia
nemo-retriever
Index folders of PDFs and other documents into LanceDB for vector search, then query them with semantic search, page filters, verbatim quotes, and cross-document aggregation.
2.2k · bundle
aniruddhaadak80
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
infometa
gifgrep
Search GIF providers with CLI/TUI, download results, and extract stills/sheets.
228
jiachen-t-wang
dall-e-zero-shot-text-to-image-generation-arxiv-2102-12092v2
DALL-E: Zero-Shot Text-to-Image Generation
6
jiachen-t-wang
textvqa-towards-reasoning-about-text-in-images-arxiv-1904-08
TextVQA: Towards Reasoning about Text in Images
6
jrennie99-glitch
gifgrep
Search GIF providers with CLI/TUI, download results, and extract stills/sheets.
0
aniruddhaadak80
gifgrep
Search GIF providers with CLI/TUI, download results, and extract stills/sheets.
0
intelli-verse-x
ivx-om-dashscope
DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans). Use when generating images via Qwen-Image, narrating via Qwen-TTS, or transcribing with word-level timestamps via Qwen-ASR.
0 · bundle
jiachen-t-wang
flamingo-a-visual-language-model-for-few-shot-learning-arxiv
Flamingo: A Visual Language Model for Few-Shot Learning
6
flyfiref
ocr-and-documents
Extract text from PDFs/scans (pymupdf, marker-pdf).
0 · bundle
infinition
ocr-and-documents
Extracts text from PDFs and scanned documents, choosing the cheapest method that works, from direct file reads to full OCR.
2 · bundle
solizardking
gifgrep
Search GIF providers with CLI/TUI, download results, and extract stills/sheets.
0
jiachen-t-wang
capsfusion-rethinking-image-text-data-at-scale-arxiv-2310-20
CapsFusion: Rethinking Image-Text Data at Scale
6
browser-act
google-image-api-skill
Extracts structured image metadata from Google Images via the BrowserAct API, including thumbnails, titles, source logos, and click-through URLs.
3.7k · bundle
inference-sh
nano-banana-2
Generate images using Google Gemini 3.1 Flash Image Preview via the inference.sh CLI, with support for text-to-image, image editing, multi-image input, and Google Search grounding.
584
lucassantana-dev
rag
Build and debug Retrieval-Augmented Generation pipelines — chunking, embedding, retrieval, reranking
1 · bundle
jimliu
baoyu-danger-gemini-web
Generates images and text via reverse-engineered Gemini Web API, supporting reference images and multi-turn conversations.
23.1k · bundle
intelli-verse-x
ivx-imagegen
Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts. Use when Codex should create a brand-new image, transform an existing image, or derive visual variants from references, and the output should be a bitmap asset rather than repo-native code or vector. Do not use when the task is better handled by editing existing SVG/vector/code-native assets, extending an established icon or logo system, or building the visual directly in HTML/CSS/canvas.
0 · bundle
aniruddhaadak80
meme-maker
Search meme templates, suggest formats, and generate local or hosted image memes.
0 · bundle
peteedoo
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
inference-sh
image-upscaling
Upscale and enhance images using Real-ESRGAN, Thera, FLUX Upscaler, and Topaz via the inference.sh CLI.
584
jiachen-t-wang
sigmoid-loss-for-language-image-pre-training-arxiv-2303-1534
Sigmoid Loss for Language Image Pre-Training
6
rootcastleco
imagen
|
6
kursku
image-seo
Image Seo — Skill especializada para otimizar imagens para mecanismos de busca, melhorando a visibilidade e performance.
55