blip-2-vision-language

orchestra-research/blip-2-vision-language · Agent Skill (multi-file)

by Orchestra Research · bundle

Published · Last updated


Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.

SKILL.md

Files

This skill is a package of 3 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁references
  • 📄advanced-usage.md 18.4 KB
  • 📄troubleshooting.md 11.7 KB

Related

  1. llava · lord1egypt
    Runs the open-source LLaVA vision-language model for image understanding, captioning, visual question answering, and multi-turn image conversations, including setup, inference, and training guidance.
    2
    repo stars
  2. llava · orchestra-research bundle
    Enables visual instruction tuning and image-based conversations using open-source vision-language models. Supports multi-turn image chat, visual question answering, and image understanding tasks.
    10.4k
    repo stars
  3. nv-reason-cxr · nvidia bundle
    Runs chest X-ray reasoning smoke tests using the NV-Reason-CXR-3B model via local inference or a public Hugging Face Space API.
    2.2k
    repo stars
  4. gemini-interactions-api · google-gemini bundle
    Call the Gemini API for text generation, chat, multimodal understanding, image/video/audio generation, streaming, function calling, structured output, and managed agents using the Interactions API in Python and TypeScript.
    3.8k
    repo stars
  5. nv-generate-ct-rflow · nvidia bundle
    Generates synthetic CT volumes and masks using NVIDIA's rectified-flow pipeline for medical imaging research.
    2.2k
    repo stars
  6. tao-finetune-clip · nvidia bundle
    Fine-tune and deploy CLIP vision-language models for zero-shot classification, image-text retrieval, and embedding extraction with ONNX and TensorRT support.
    2.2k
    repo stars

Frequently asked questions

How do I install the blip-2-vision-language skill?

Run npx skillmds add orchestra-research/blip-2-vision-language in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the blip-2-vision-language skill do?

Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs. It is listed under AI & ML, Coding & Dev Tools, Research & Search, Image & Video Generation, RAG & Embeddings on SkillMD.

Is blip-2-vision-language safe to use?

SkillMD's automated safety review verdict for this skill is PASS. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: makes network calls, reads secrets. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with blip-2-vision-language?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is blip-2-vision-language free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under MIT.

Who published blip-2-vision-language?

Orchestra Research (@orchestra-research) published this skill. Their other Agent Skills are listed on their SkillMD profile.