Coding & Dev Tools
9,772 skillsDall E Zero Shot Text To Image Generation Arxiv 2102 12092v2
DALL-E: Zero-Shot Text-to-Image Generation
6
Minicpm V A Gpt 4v Level Mllm On Your Phone Arxiv 2408 01800
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
6
Sigmoid Loss For Language Image Pre Training Arxiv 2303 1534
Sigmoid Loss for Language Image Pre-Training
6
Data Centric Artificial Intelligence A Survey Arxiv 2303 101
Data-Centric Artificial Intelligence: A Survey
6
Scaling Instruction Finetuned Language Models Arxiv 2210 114
Scaling Instruction-Finetuned Language Models
6
Scaling Vision With Sparse Mixture Of Experts Arxiv 2106 059
Scaling Vision with Sparse Mixture of Experts
6
Emu Generative Pretraining In Multimodality Arxiv 2307 05222
Emu: Generative Pretraining in Multimodality
6
Hard Negative Mixing For Contrastive Learning Arxiv 2010 010
Hard Negative Mixing for Contrastive Learning
6
Coco Microsoft Coco Common Objects In Context Arxiv 1405 031
COCO: Microsoft COCO: Common Objects in Context
6
Open Vocabulary Object Detection Using Captions Arxiv 2011 1
Open-Vocabulary Object Detection Using Captions
6
Capsfusion Rethinking Image Text Data At Scale Arxiv 2310 20
CapsFusion: Rethinking Image-Text Data at Scale
6
Masked Autoencoders Are Scalable Vision Learners Arxiv 2111
Masked Autoencoders Are Scalable Vision Learners
6
Textvqa Towards Reasoning About Text In Images Arxiv 1904 08
TextVQA: Towards Reasoning about Text in Images
6
Improved Baselines With Visual Instruction Tuning Arxiv 2310
Improved Baselines with Visual Instruction Tuning
6
Lora Low Rank Adaptation Of Large Language Models Arxiv 2106
LoRA: Low-Rank Adaptation of Large Language Models
6
Mosaic Augmentation For Detection And Segmentation Arxiv Yol
Mosaic Augmentation for Detection and Segmentation
6
Cogvlm Visual Expert For Pretrained Language Models Arxiv 23
CogVLM: Visual Expert for Pretrained Language Models
6
Chameleon Mixed Modal Early Fusion Foundation Models Arxiv 2
Chameleon: Mixed-Modal Early-Fusion Foundation Models
6
Longva Long Context Transfer From Language To Vision Arxiv 2
LongVA: Long Context Transfer from Language to Vision
6
Scaling Vision Transformers To 22 Billion Parameters Arxiv 2
Scaling Vision Transformers to 22 Billion Parameters
6
Refinedweb A Falcons Recipe For High Quality Web Data Arxiv
RefinedWeb: A Falcon's Recipe for High-Quality Web Data
6
Drivelm Driving With Graph Visual Question Answering Arxiv 2
DriveLM: Driving with Graph Visual Question Answering
6
Laclip Improving Clip Training With Language Rewrites Arxiv
LaCLIP: Improving CLIP Training with Language Rewrites
6
Eva Clip Improved Training Techniques For Clip At Scale Arxi
EVA-CLIP: Improved Training Techniques for CLIP at Scale
6
Deduplicating Training Data Makes Language Models Better Arx
Deduplicating Training Data Makes Language Models Better
6
Dreamlip Language Image Pre Training With Long Captions Arxi
DreamLIP: Language-Image Pre-training with Long Captions
6
Emu2 Generative Multimodal Models Are In Context Learners Ar
Emu2: Generative Multimodal Models are In-Context Learners
6
Multimodal Few Shot Learning With Frozen Language Models Arx
Multimodal Few-Shot Learning with Frozen Language Models
6
Nlvr2 A Visual Reasoning Benchmark For Natural Language Arxi
NLVR2: A Visual Reasoning Benchmark for Natural Language
6
P10
P10 CTO mode — define strategic direction, design org topology, manage P9 teams. Use when user says 'CTO模式', 'P10', '战略规划', '架构委员会', or when facing cross-team architectural decisions. Produces: strategic input templates + org design.
0 · bundle
Pro
PUA Pro extensions: self-evolution notes, compaction state continuity, KPI-style summaries, flavor switching, and feedback tools.
0 · bundle
Pua
Use for PUA/try-harder productivity coaching when the user expresses frustration, repeated failure, quality complaint, passive behavior, says retry/change approach/don't give up, asks for evidence/completion check/test before done, or wants Ding-style workplace reminders. Triggers include: try harder, stop giving up, figure it out, again??, why still failing, change approach, no evidence, run tests, done without proof, 换个方法, 再试试, 别摆烂, 别偷懒, 为什么还不行, 又错了, 证据呢, 没跑测试别说完成, 验收, 闭环, 自嗨, 置身钉外, 无招, 老板体感. Do not use for calm first-attempt requests.
0 · bundle
Think
Turns rough ideas into approved, decision-complete plans with validated structure before coding. Use when users ask in any language for planning, architecture, design direction, feasibility, value judgment, or whether a feature is worth doing before implementation. Not for bug fixes or small edits.
0 · bundle
Write
Rewrites and polishes prose in Chinese or English, removes AI-like wording, and reviews product localization copy while preserving intent for drafts, docs, release notes, launch copy, and social posts. Use when users ask in any language to draft, rewrite, proofread, localize, polish release notes, remove AI-like wording, or prepare launch and social copy. Not for code comments, commit messages, or inline docs.
0 · bundle
Vet
Run vet immediately after ANY logical unit of code changes. Do not batch your changes, do not wait to be asked to run vet, make sure you are proactive.
0 · bundle
Tend
Tend the Allium garden. Use when the user wants to write, edit, update, add to, improve, clarify, refine, restructure, fix or migrate Allium specs. Covers adding entities, rules, triggers, surfaces and contracts, fixing syntax or validation errors, renaming or refactoring within specs, migrating specs to a new language version, and translating requirements into well-formed specifications. Pushes back on vague requirements.
0 · bundle