All Skills

25,836 skills
jiachen-t-wang
Textvqa Towards Reasoning About Text In Images Arxiv 1904 08
TextVQA: Towards Reasoning about Text in Images
6
jiachen-t-wang
Improved Baselines With Visual Instruction Tuning Arxiv 2310
Improved Baselines with Visual Instruction Tuning
6
jiachen-t-wang
Lora Low Rank Adaptation Of Large Language Models Arxiv 2106
LoRA: Low-Rank Adaptation of Large Language Models
6
jiachen-t-wang
Mosaic Augmentation For Detection And Segmentation Arxiv Yol
Mosaic Augmentation for Detection and Segmentation
6
jiachen-t-wang
Kinetics 400 A Large Video Understanding Dataset Arxiv 1705
Kinetics-400: A Large Video Understanding Dataset
6
jiachen-t-wang
Cogvlm Visual Expert For Pretrained Language Models Arxiv 23
CogVLM: Visual Expert for Pretrained Language Models
6
jiachen-t-wang
Chameleon Mixed Modal Early Fusion Foundation Models Arxiv 2
Chameleon: Mixed-Modal Early-Fusion Foundation Models
6
jiachen-t-wang
Longva Long Context Transfer From Language To Vision Arxiv 2
LongVA: Long Context Transfer from Language to Vision
6
jiachen-t-wang
Scaling Vision Transformers To 22 Billion Parameters Arxiv 2
Scaling Vision Transformers to 22 Billion Parameters
6
jiachen-t-wang
Refinedweb A Falcons Recipe For High Quality Web Data Arxiv
RefinedWeb: A Falcon's Recipe for High-Quality Web Data
6
jiachen-t-wang
Donut Document Understanding Transformer Without Ocr Arxiv 2
Donut: Document Understanding Transformer without OCR
6
jiachen-t-wang
Drivelm Driving With Graph Visual Question Answering Arxiv 2
DriveLM: Driving with Graph Visual Question Answering
6
jiachen-t-wang
Laclip Improving Clip Training With Language Rewrites Arxiv
LaCLIP: Improving CLIP Training with Language Rewrites
6
jiachen-t-wang
Nuscenes A Multimodal Dataset For Autonomous Driving Arxiv 1
nuScenes: A Multimodal Dataset for Autonomous Driving
6
jiachen-t-wang
Eva Clip Improved Training Techniques For Clip At Scale Arxi
EVA-CLIP: Improved Training Techniques for CLIP at Scale
6
jiachen-t-wang
Deduplicating Training Data Makes Language Models Better Arx
Deduplicating Training Data Makes Language Models Better
6
jiachen-t-wang
Dreamlip Language Image Pre Training With Long Captions Arxi
DreamLIP: Language-Image Pre-training with Long Captions
6
jiachen-t-wang
Flamingo A Visual Language Model For Few Shot Learning Arxiv
Flamingo: A Visual Language Model for Few-Shot Learning
6
jiachen-t-wang
No Robots A Dataset Of Personally Written Instructions Arxiv
No Robots: A Dataset of Personally Written Instructions
6
jiachen-t-wang
Emu2 Generative Multimodal Models Are In Context Learners Ar
Emu2: Generative Multimodal Models are In-Context Learners
6
jiachen-t-wang
3d LLM Injecting The 3d World Into Large Language Models Arx
3D-LLM: Injecting the 3D World into Large Language Models
6
jiachen-t-wang
Multimodal Few Shot Learning With Frozen Language Models Arx
Multimodal Few-Shot Learning with Frozen Language Models
6
jiachen-t-wang
Nlvr2 A Visual Reasoning Benchmark For Natural Language Arxi
NLVR2: A Visual Reasoning Benchmark for Natural Language
6
om-scogo
Qmd
Search local markdown knowledge bases, notes, docs, and wikis with QMD. Use when users ask to find notes, retrieve documents, inspect a wiki, answer from indexed markdown, or set up QMD access.
0 · bundle
om-scogo
P9
P9 Tech Lead mode — write Task Prompts, manage P8 agent teams, never write code yourself. Use when user says 'P9模式', 'tech-lead', '帮我管理这个项目', '任务拆解', or when coordinating 3+ parallel agents. Produces: Task Prompts (六要素) + P8 team delivery.
0 · bundle
om-scogo
Hunt
Finds root cause before applying fixes for errors, crashes, regressions, failing tests, broken behavior, and screenshot-reported defects. Use when users report in any language errors, crashes, broken behavior, regressions, failing tests, screenshot evidence, or something that used to work and now fails. Not for code review or new features.
0 · bundle
om-scogo
Read
Reads URLs and PDFs by fetching source content, defaulting to concise summaries for plain read requests and clean Markdown when asked to convert, save, quote, cite, or feed downstream work. Use when users ask in any language to read, fetch, check, summarize, quote, cite, convert, or save a URL or PDF. Not for local text files already in the repo.
0 · bundle
om-scogo
P10
P10 CTO mode — define strategic direction, design org topology, manage P9 teams. Use when user says 'CTO模式', 'P10', '战略规划', '架构委员会', or when facing cross-team architectural decisions. Produces: strategic input templates + org design.
0 · bundle
om-scogo
Pro
PUA Pro extensions: self-evolution notes, compaction state continuity, KPI-style summaries, flavor switching, and feedback tools.
0 · bundle
om-scogo
Pua
Use for PUA/try-harder productivity coaching when the user expresses frustration, repeated failure, quality complaint, passive behavior, says retry/change approach/don't give up, asks for evidence/completion check/test before done, or wants Ding-style workplace reminders. Triggers include: try harder, stop giving up, figure it out, again??, why still failing, change approach, no evidence, run tests, done without proof, 换个方法, 再试试, 别摆烂, 别偷懒, 为什么还不行, 又错了, 证据呢, 没跑测试别说完成, 验收, 闭环, 自嗨, 置身钉外, 无招, 老板体感. Do not use for calm first-attempt requests.
0 · bundle
om-scogo
Learn
Runs a six-phase research workflow that turns unfamiliar domains, source bundles, or collected material into publish-ready output. Use when users ask in any language to research, study, deep-dive, compile sources, synthesize unfamiliar material, or turn a source bundle into a coherent reference. Not for quick lookups or single-file reads.
0 · bundle
om-scogo
Think
Turns rough ideas into approved, decision-complete plans with validated structure before coding. Use when users ask in any language for planning, architecture, design direction, feasibility, value judgment, or whether a feature is worth doing before implementation. Not for bug fixes or small edits.
0 · bundle
om-scogo
Write
Rewrites and polishes prose in Chinese or English, removes AI-like wording, and reviews product localization copy while preserving intent for drafts, docs, release notes, launch copy, and social posts. Use when users ask in any language to draft, rewrite, proofread, localize, polish release notes, remove AI-like wording, or prepare launch and social copy. Not for code comments, commit messages, or inline docs.
0 · bundle
om-scogo
Vet
Run vet immediately after ANY logical unit of code changes. Do not batch your changes, do not wait to be asked to run vet, make sure you are proactive.
0 · bundle
om-scogo
Tend
Tend the Allium garden. Use when the user wants to write, edit, update, add to, improve, clarify, refine, restructure, fix or migrate Allium specs. Covers adding entities, rules, triggers, surfaces and contracts, fixing syntax or validation errors, renaming or refactoring within specs, migrating specs to a new language version, and translating requirements into well-formed specifications. Pushes back on vague requirements.
0 · bundle
om-scogo
Weed
Weed the Allium garden. Find where Allium specifications and implementation code have diverged, and help resolve the divergences. Use when the user wants to check spec-code alignment, compare specs against implementation, audit for spec drift or violations, sync specs with code or code with specs, or verify whether the implementation matches what the spec says.
0 · bundle