Results for “imagen-4”
26 skillsMore results
Gemini API Dev
Build applications with Gemini API hosted models, including multimodal content, function calling, and structured outputs, using the latest SDKs and model specifications.
0
Image Upscaling
Upscale and enhance images using Real-ESRGAN, Thera, FLUX Upscaler, and Topaz via the inference.sh CLI.
584
Qwen Image 2
Generate and edit images using Alibaba Qwen-Image-2.0 models via the inference.sh CLI, with support for text-to-image, multi-image editing, and text rendering.
584
Vision Analyze
画像を理解(被写体・テキストOCR・構図・色・UI構造の分析)し、結果を構造化して返すスキル。CC CLI は GLM-5.3 等の vision 非対応モデルで稼働中のため画像を直接視認できず、主ルート Gemini 2.5 Flash(scripts/api/gemini_vision.py・無料枠)と副ルート 4_5v MCP(analyze_image・Readが返すCDN URL)の2経路で分析し、CCは結果の構造化・比較・保存に専任する。 ユーザーが「画像見て」「この画像何が写ってる」「画像比較して」「スクショ見て」「画像分析して」「画像理解」「vision-analyze」と言った時、または /vision-analyze を呼んだ時にトリガー。 ※画像生成(image generation)は対象外(make-song / video-prompt-spec / demo-site-sales参照)。ピクセル修正(花鈿除去等)は remove-huadian の役割。楽曲分析は analyze-song / reverse-engineer-song。
0
151 Copy E139d00c
封装阶段3的Copy Spec为可批量执行的提示词包,生成JSONL请求并调用APIMart API出图。
7 · bundle
Gemini API Dev
Build applications with Gemini API hosted models, including Gemini and Gemma 4, using multimodal content, function calling, structured outputs, and current SDKs for Python, JavaScript, Go, and Java.
3.8k
Tao Mine Aoi Images
Embeds target and source image parquets, then mines nearest-neighbour source images for augmentation in VCN AOI workflows.
2.2k · bundle
Huggingface Accelerate
Simplest distributed training API. 4 lines to add distributed support to any PyTorch script. Unified API for DeepSpeed/FSDP/Megatron/DDP. Automatic device placement, mixed precision (FP16/BF16/FP8). Interactive config, single launch command. HuggingFace ecosystem standard.
0 · bundle
Gif Sticker Maker
Convert user photos into 4 animated GIF stickers in Funko Pop / Pop Mart style with customizable captions.
12.9k · bundle
Nano Banana 2
Generate images using Google Gemini 3.1 Flash Image Preview via the inference.sh CLI, with support for text-to-image, image editing, multi-image input, and Google Search grounding.
584
Blip 2 Vision Language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
Flux 2 Klein
Generate images with Flux 2 Klein (Black Forest Labs' distilled fast variant of Flux 2) on RunComfy — bundled with the model's documented prompting patterns so the skill gets sharper output than naive prompting against the same model. Documents Flux 2 Klein's strengths (sub-second latency, multi-reference brand styling, declarative subject-first prompts), the step-count strategy (4–8 for fast iteration, ~25 for polish), the 9B vs 4B variant trade-off, and when to route to Flux 2 Pro / Seedream 5 / GPT Image 2 instead. Calls `runcomfy run blackforestlabs/flux-2-klein/9b/text-to-image` (or `/4b/`) through the local RunComfy CLI. Triggers on "flux 2 klein", "flux-2-klein", "flux klein", "BFL flux 2", or any explicit ask to generate with this model.
33
Flamingo A Visual Language Model For Few Shot Learning Arxiv
Flamingo: A Visual Language Model for Few-Shot Learning
6
Gpt Image
Generate and edit images using OpenAI's GPT-Image-2 model via the inference.sh CLI, supporting text-to-image, image editing, inpainting, and batch generation.
584
Flux Image
Generate images using FLUX models via the inference.sh CLI, supporting text-to-image, image-to-image, and LoRA fine-tuning.
584
Fal AI
Generates and edits images and videos via fal.ai's queue-based API, supporting models like Flux, Gemini image, and Kling video-to-video, with automatic polling.
1 · bundle
Gpt Image 2
Generates and edits images using GPT Image 2 across three modes: direct generation via OpenAI-compatible API, prompt engineering for host-native image tools, or pure prompt advisory. Includes 80+ structured templates for posters, UI mockups, product visuals, maps, slides, and more.
9.2k · bundle
Imagen
Generates images using Google Gemini's image generation model for UI placeholders, documentation, and design assets.
42.4k
Pixtral 12b A Frontier Multimodal Model Arxiv Pixtral 2024
Pixtral 12B: A Frontier Multimodal Model
6
Stable Diffusion Image Generation
Generate images from text prompts, perform image-to-image translation, inpainting, and build custom diffusion pipelines using Stable Diffusion models via HuggingFace Diffusers.
10.4k · bundle
Surrealdb
Expert guidance for architecting, developing, and operating SurrealDB 3, covering SurrealQL, multi-model data modeling, vector search, security, deployment, performance tuning, SDK integration, and ecosystem tools.
34 · bundle
Imagegen
Generate or edit images through ClawRouter's local API, with automatic payment via x402 and support for multiple models.
17
Image Enhancer
Improves the quality of images, especially screenshots, by enhancing resolution, sharpness, and clarity. Perfect for preparing images for presentations, documentation, or social media posts.
3
Huggingface Accelerate
Simplest distributed training API. 4 lines to add distributed support to any PyTorch script. Unified API for DeepSpeed/FSDP/Megatron/DDP. Automatic device placement, mixed precision (FP16/BF16/FP8). Interactive config, single launch command. HuggingFace ecosystem standard.
1 · bundle
Ivx Om Gemini Omni
Generate and conversationally edit short videos with Google Gemini Omni Flash (`gemini-omni-flash-preview`). Use when: (1) iterating on a clip with natural-language edits instead of regenerating ("make the phone invisible, keep everything else the same"), (2) generating 3-10s 720p clips with synthesized audio, rendered on-screen text, or timecoded beats, (3) binding reference images to roles with <FIRST_FRAME>/<IMAGE_REF_N> prompt tags, (4) editing an existing uploaded video. Accessed via the `gemini_omni_video` tool using the project's GEMINI_API_KEY/GOOGLE_API_KEY — the same key as Imagen and Google TTS.
0 · bundle