All Skills
25,836 skillsTextvqa Towards Reasoning About Text In Images Arxiv 1904 08
TextVQA: Towards Reasoning about Text in Images
6
Improved Baselines With Visual Instruction Tuning Arxiv 2310
Improved Baselines with Visual Instruction Tuning
6
Lora Low Rank Adaptation Of Large Language Models Arxiv 2106
LoRA: Low-Rank Adaptation of Large Language Models
6
Mosaic Augmentation For Detection And Segmentation Arxiv Yol
Mosaic Augmentation for Detection and Segmentation
6
Kinetics 400 A Large Video Understanding Dataset Arxiv 1705
Kinetics-400: A Large Video Understanding Dataset
6
Cogvlm Visual Expert For Pretrained Language Models Arxiv 23
CogVLM: Visual Expert for Pretrained Language Models
6
Chameleon Mixed Modal Early Fusion Foundation Models Arxiv 2
Chameleon: Mixed-Modal Early-Fusion Foundation Models
6
Longva Long Context Transfer From Language To Vision Arxiv 2
LongVA: Long Context Transfer from Language to Vision
6
Scaling Vision Transformers To 22 Billion Parameters Arxiv 2
Scaling Vision Transformers to 22 Billion Parameters
6
Refinedweb A Falcons Recipe For High Quality Web Data Arxiv
RefinedWeb: A Falcon's Recipe for High-Quality Web Data
6
Donut Document Understanding Transformer Without Ocr Arxiv 2
Donut: Document Understanding Transformer without OCR
6
Drivelm Driving With Graph Visual Question Answering Arxiv 2
DriveLM: Driving with Graph Visual Question Answering
6
Laclip Improving Clip Training With Language Rewrites Arxiv
LaCLIP: Improving CLIP Training with Language Rewrites
6
Nuscenes A Multimodal Dataset For Autonomous Driving Arxiv 1
nuScenes: A Multimodal Dataset for Autonomous Driving
6
Eva Clip Improved Training Techniques For Clip At Scale Arxi
EVA-CLIP: Improved Training Techniques for CLIP at Scale
6
Deduplicating Training Data Makes Language Models Better Arx
Deduplicating Training Data Makes Language Models Better
6
Dreamlip Language Image Pre Training With Long Captions Arxi
DreamLIP: Language-Image Pre-training with Long Captions
6
Flamingo A Visual Language Model For Few Shot Learning Arxiv
Flamingo: A Visual Language Model for Few-Shot Learning
6
No Robots A Dataset Of Personally Written Instructions Arxiv
No Robots: A Dataset of Personally Written Instructions
6
Emu2 Generative Multimodal Models Are In Context Learners Ar
Emu2: Generative Multimodal Models are In-Context Learners
6
3d LLM Injecting The 3d World Into Large Language Models Arx
3D-LLM: Injecting the 3D World into Large Language Models
6
Multimodal Few Shot Learning With Frozen Language Models Arx
Multimodal Few-Shot Learning with Frozen Language Models
6
Nlvr2 A Visual Reasoning Benchmark For Natural Language Arxi
NLVR2: A Visual Reasoning Benchmark for Natural Language
6
Qmd
Search local markdown knowledge bases, notes, docs, and wikis with QMD. Use when users ask to find notes, retrieve documents, inspect a wiki, answer from indexed markdown, or set up QMD access.
0 · bundle
P9
P9 Tech Lead mode — write Task Prompts, manage P8 agent teams, never write code yourself. Use when user says 'P9模式', 'tech-lead', '帮我管理这个项目', '任务拆解', or when coordinating 3+ parallel agents. Produces: Task Prompts (六要素) + P8 team delivery.
0 · bundle
Hunt
Finds root cause before applying fixes for errors, crashes, regressions, failing tests, broken behavior, and screenshot-reported defects. Use when users report in any language errors, crashes, broken behavior, regressions, failing tests, screenshot evidence, or something that used to work and now fails. Not for code review or new features.
0 · bundle
Read
Reads URLs and PDFs by fetching source content, defaulting to concise summaries for plain read requests and clean Markdown when asked to convert, save, quote, cite, or feed downstream work. Use when users ask in any language to read, fetch, check, summarize, quote, cite, convert, or save a URL or PDF. Not for local text files already in the repo.
0 · bundle
P10
P10 CTO mode — define strategic direction, design org topology, manage P9 teams. Use when user says 'CTO模式', 'P10', '战略规划', '架构委员会', or when facing cross-team architectural decisions. Produces: strategic input templates + org design.
0 · bundle
Pro
PUA Pro extensions: self-evolution notes, compaction state continuity, KPI-style summaries, flavor switching, and feedback tools.
0 · bundle
Pua
Use for PUA/try-harder productivity coaching when the user expresses frustration, repeated failure, quality complaint, passive behavior, says retry/change approach/don't give up, asks for evidence/completion check/test before done, or wants Ding-style workplace reminders. Triggers include: try harder, stop giving up, figure it out, again??, why still failing, change approach, no evidence, run tests, done without proof, 换个方法, 再试试, 别摆烂, 别偷懒, 为什么还不行, 又错了, 证据呢, 没跑测试别说完成, 验收, 闭环, 自嗨, 置身钉外, 无招, 老板体感. Do not use for calm first-attempt requests.
0 · bundle
Learn
Runs a six-phase research workflow that turns unfamiliar domains, source bundles, or collected material into publish-ready output. Use when users ask in any language to research, study, deep-dive, compile sources, synthesize unfamiliar material, or turn a source bundle into a coherent reference. Not for quick lookups or single-file reads.
0 · bundle
Think
Turns rough ideas into approved, decision-complete plans with validated structure before coding. Use when users ask in any language for planning, architecture, design direction, feasibility, value judgment, or whether a feature is worth doing before implementation. Not for bug fixes or small edits.
0 · bundle
Write
Rewrites and polishes prose in Chinese or English, removes AI-like wording, and reviews product localization copy while preserving intent for drafts, docs, release notes, launch copy, and social posts. Use when users ask in any language to draft, rewrite, proofread, localize, polish release notes, remove AI-like wording, or prepare launch and social copy. Not for code comments, commit messages, or inline docs.
0 · bundle
Vet
Run vet immediately after ANY logical unit of code changes. Do not batch your changes, do not wait to be asked to run vet, make sure you are proactive.
0 · bundle
Tend
Tend the Allium garden. Use when the user wants to write, edit, update, add to, improve, clarify, refine, restructure, fix or migrate Allium specs. Covers adding entities, rules, triggers, surfaces and contracts, fixing syntax or validation errors, renaming or refactoring within specs, migrating specs to a new language version, and translating requirements into well-formed specifications. Pushes back on vague requirements.
0 · bundle
Weed
Weed the Allium garden. Find where Allium specifications and implementation code have diverged, and help resolve the divergences. Use when the user wants to check spec-code alignment, compare specs against implementation, audit for spec drift or violations, sync specs with code or code with specs, or verify whether the implementation matches what the spec says.
0 · bundle