COLM Skills
A 12-skill depth pack for COLM (Conference on Language Modeling) submissions: venue routing for LM research, the March abstract/paper deadlines, 9-page format, OpenReview rebuttal, contamination-aware evaluation, API-model reproducibility, compute disclosure, camera-ready, and conference planning. Grounded in COLM 2026 official pages checked on 2026-07-08.
Skills in this plugin
11- ▌ Colm Submission · brycewang-stanfordUse when preparing or auditing a COLM submission on OpenReview — the late-March abstract and full-paper deadlines, the strict 9-page main text, double-blind rules banning acknowledgments and identity links, the Code of Ethics acknowledgment, LLM-usage disclosure, reciprocal-reviewer nomination, and pre-upload risk triage.
- ▌ Colm Experiments · brycewang-stanfordUse when designing or auditing the empirical core of a COLM paper — contamination analysis for evaluation data, fair baselines under matched prompting and compute, pinned model versions and decoding parameters, uncertainty over runs and samples, scaling coverage, and honest reporting of API-model comparisons.
- ▌ Colm Camera Ready · brycewang-stanfordUse when converting a COLM acceptance into the final paper — the August 7 camera-ready deadline in 2026, de-anonymization and the one-page acknowledgments allowance, folding rebuttal commitments into the text, publishing on OpenReview, releasing artifacts publicly, and planning the October conference in San Francisco.
- ▌ Colm Related Work · brycewang-stanfordUse when positioning a COLM paper inside the fastest-moving literature in ML — triaging arXiv-heavy citations, handling concurrent work fairly, citing model and dataset artifacts correctly, distinguishing COLM's three-edition archive from adjacent venues' archives, and keeping self-citation double-blind safe.
- ▌ Colm Supplementary · brycewang-stanfordUse when deciding what goes into a COLM paper's appendices and supplementary material versus the strict 9-page main text — verbatim prompts, full evaluation configurations, per-task result tables, contamination analyses, human-evaluation protocols, and anonymized code/data packages that survive double-blind review.
- ▌ Colm Writing Style · brycewang-stanfordUse when drafting or revising COLM paper prose — leading with a finding about language models rather than a leaderboard delta, scoping claims to tested models and scales, naming versions in text, keeping the 9-page main text self-sufficient, and matching the measured, analysis-forward voice of COLM's award lineage.
- ▌ Colm Review Process · brycewang-stanfordUse when reasoning about how COLM reviews a paper — the OpenReview pipeline from late-March submission through the May review release, the May-June rebuttal window, July decisions, reciprocal-reviewing obligations, the LLM-use rules for reviewers, and how a three-edition-old venue's norms differ from mature conferences.
- ▌ Colm Author Response · brycewang-stanfordUse when writing a COLM rebuttal in the OpenReview discussion phase — triaging reviews released in late May, running cache-based follow-up experiments inside the roughly two-and-a-half-week window, answering contamination and baseline-fairness objections with evidence, and writing for the area chair who decides in July.
- ▌ Colm Reproducibility · brycewang-stanfordUse when hardening a COLM paper's reproducibility story — pinning open-weight checkpoints and tokenizers, handling API-model drift and deprecation honestly, versioning evaluation harnesses and prompts, disclosing compute, and writing availability statements that distinguish what is releasable from what is not.
- ▌ Colm Topic Selection · brycewang-stanfordUse when deciding whether language-model research belongs at COLM or should route to ACL/EMNLP, ICLR, NeurIPS, ICML, or a workshop — applying the object-of-study test, matching against COLM's CFP lanes (training, data, evaluation, inference, safety), and weighing the trade-offs of a young venue before writing begins.
- ▌ Colm Artifact Evaluation · brycewang-stanfordUse when packaging the artifacts of a COLM paper — model weights, training data, prompts, evaluation sets, and cached model outputs — for anonymous review and public post-acceptance release, navigating licenses, API terms-of-service limits, and the absence of a formal COLM artifact track.