scdenney
- 37 skills
- 0 followers
- 1 week ago last updated
- ▌ Paper Tex · scdenney bundleTypeset a working paper or journal submission in house-style LaTeX from any draft — Markdown, Word (.docx), TeX, ODT, RTF, or HTML. Convert with pandoc, wrap in an EB Garamond template, build the PDF with latexmk, and prepare for a specific journal (spacing, page limit, anonymization, disclosures, citation style). Use for "format/typeset/convert my paper to LaTeX", "make a working paper", "prepare this for submission to <journal>".
- ▌ Presubmit · scdenneyLauncher and setup wizard for the standalone presubmit CLI, an API-driven adversarial peer-review pipeline of 30-plus stages (Red Team finders, Blue Team defence, verification cascade, legal pass, copyedit) that writes one consolidated review report to disk. Verifies the install and the API key, settles where output lands, picks smoke, standard, or custom mode, launches the run, and reports where the report ended up. Use when the user asks to run presubmit, wants a deep unattended audit of a draft before submission, or needs a math or replication-code audit. The everyday in-session self-audit is paper-review-lite, and refereeing someone else's manuscript is journal-review.
- ▌ Figure Table Audit · scdenney bundleEnd-stage QA on a finished figure and table set — an inventory with in-text callouts and producing scripts, cross-reference and numbering checks, text-to-evidence consistency between claims and plotted or tabulated values, captions and statistical notes that stand alone, accessibility and production quality, and SI and replication-package linkage. Marks anything requiring values read off an image as needing author verification rather than inventing them. Use when the user is preparing a submission and asks to check figures, tables, captions, cross-references, or table notes, or asks whether the numbers in the text match the exhibits. Design guidance during drafting lives in figures and tables.
- ▌ Narrative Building · scdenney bundleDraft or audit scientific introductions. Use for argument logic, framing, contribution structure, and coherence across multiple studies or experiments.
- ▌ Research Wayfinder · scdenney bundlePlan a research project as a durable decision map, resolving one design decision per session until the destination is a defensible, pre-registerable design. Not for executing analyses or writing the paper. Use at the start of a project, when a design has more open decisions than one conversation holds, or when planning keeps restarting.
- ▌ Hypothesis Building · scdenney bundleTurns a theory into falsifiable, pre-registerable hypotheses — DAGs and backdoor closure, how the design resolves the FPCI, SATE versus PATE, counterfactual and directional framing, a named theoretical and empirical estimand, a justified SESOI, the choice among NHST, interval, equivalence (TOST), and minimum-effect tests, scope conditions, and primary, secondary, and exploratory tiers across multi-experiment designs. Use when the user has a research question or theory and asks to turn it into hypotheses, asks whether a prediction is falsifiable or testable, asks what the estimand is, or asks how to predict a null. Framing the paper around it goes to narrative-building, the plan to pre-registration-writing.
- ▌ Replication Package · scdenney bundleScaffold or audit a social-science replication package, and audit the manuscript and its archived research objects against the FAIR principles. Scaffold mode writes the folder structure, README, master.R, figure/table crosswalk, codebook template, LICENSE placeholder, .gitignore, and pre-release checklist. Audit mode grades an existing package against that checklist and runs the FAIR block over data, code, materials, prompts, preregistrations, DOIs, metadata, licenses, access restrictions, and availability statements. Use when setting up or repairing a replication package, checking one before submission, auditing research objects against FAIR (Findable, Accessible, Interoperable, Reusable), or drafting and verifying data-, code-, and materials-availability statements. Adapted from Yusaku Horiuchi's replication-package-guide; platform-neutral (Harvard Dataverse, OSF, Zenodo, GitHub releases, institutional archives).
- ▌ Text Classification · scdenney bundleDesigns and validates LLM-based text classification for research data — codebook construction, choice of learning regime, model selection and reproducibility, prompt construction, pilot validation against human coding with agreement statistics (kappa, F1), hybrid human-LLM workflows, and reporting model-coded data. Also carries a resumable batch pipeline with a rule-based baseline for coding large sets of repeated free-text values against a closed codebook — occupation, institution, and registry text, and open-text survey responses — with a frozen input set, incremental output, residual buckets, stratified hand validation, a regex or dictionary comparison, and a published lookup table. Use when the user asks to classify, code, or label text at scale. Discovering categories rather than applying them goes to topic-modeling.
- ▌ Conjoint Diagnostics · scdenney bundleReviews an existing conjoint study for threats to inference and returns prioritized findings across five areas — design integrity (attributes, profile restrictions, task count and satisficing, randomization, power), estimation (estimand clarity, reference levels, subgroups, clustered standard errors, multiple testing), measurement error, external validity and behavioral benchmarking, and interpretation, including the guard-rail against reading an AMCE as a majority preference. Use when the user asks whether a conjoint design or analysis holds up, has referee comments on a conjoint, or wants a second opinion on estimation and interpretation choices. Building a design from scratch goes to conjoint-design, reshaping the data to conjoint-cleaning.
- ▌ Model Council Voting · scdenney bundleRuns several language models as independent coders on the same labeling or discovery task and reads their disagreement as data — panel assembly for model diversity, keeping votes independent, consensus rules, agreement statistics (Cohen's and Fleiss kappa, Krippendorff's alpha), the correlated-errors caveat that stops agreement being mistaken for validity, human validation beyond the panel, and reporting. Use when the user asks about a council, panel, ensemble, or jury of models, asks how to combine several models' labels, or asks what kappa or alpha to report for model coders. Single-model codebook and validation work goes to text-classification, per-item confidence to llm-calibration-logprobs.
- ▌ Cross National Design · scdenney bundleDesigns multi-country survey experiments — case selection with a stated comparative logic, instrument localization and TRAPD translation, origin-country and composition choices for stimuli, per-country power and error management, and a pooled-versus-per-country analytical strategy with measurement-equivalence checks. Use when the user is fielding the same experiment in several countries, asks which countries to include and why, asks how to translate or adapt an instrument without breaking comparability, asks whether to pool or estimate per country, or asks how large each country sample must be. Question wording goes to survey-design, per-country CONSORT and DA-RT reporting to methods-reporting, and the study-level plan to pre-registration-writing.
- ▌ Research Grill · scdenneyInterviews a researcher, in rounds, until a research idea, design, or draft has no silently assumed decision left — every question numbered, each with a plain-language "why this matters" and a recommended answer, facts fetched by the assistant rather than asked, and every settled decision written to a file. Three stages, idea (a topic or hunch → a falsifiable question and a contribution claim), design (a question → estimand, identification, sample and power, measurement, pre-registration, analysis plan, venue), and defend (a finished design or draft → the objections a reviewer would raise, before the reviewer does). Use when the user says "grill me", "grill this", "stress-test my idea/design/plan", "interview me about this project", "poke holes in this", "what am I assuming", or brings a research idea that is not yet a design. Suitable for BA, MA, and PhD students as well as grant designs; it never answers the research question for the researcher. Adapted for research from Matt Pocock's grill-me. For a softwa
- ▌ LLM Calibration Logprobs · scdenney bundleReads a model's own uncertainty off its token log-probabilities — collecting logprobs and aggregating multi-token labels, confidence tiers and margins for triage, calibration assessment with ECE, Brier scores, and reliability diagrams, using confidence downstream without laundering it into evidence, and what to archive for reproducibility. Use when the user asks how confident a classifier was on each item, asks about logprobs, top-k tokens, calibration, ECE, Brier, or reliability diagrams, or wants low-confidence cases routed to human review. This is within-model confidence — agreement across several independent model coders goes to model-council-voting, and codebook and validation design to text-classification.
- ▌ Pre Registration Writing · scdenney bundleWrites a pre-analysis plan before data collection — registry selection (OSF, AsPredicted, AEA, EGAP), PAP document structure, an analytical strategy specified down to the model and the decision rule, analysis code pre-registered against simulated data, contingency planning for attrition, failed manipulations, and exclusions, deviation documentation, and timeline. Operationalizes the pre-data-collection side of DA-RT. Use when the user asks to write or review a pre-registration or PAP, asks which registry to use, asks what to lock down versus leave exploratory, or asks how to handle a deviation later. Hypotheses and estimands come from hypothesis-building, post-hoc reporting from methods-reporting.
- ▌ Literature Review · scdenneyBuilds or audits a literature review — evidence map, closest prior work, source clusters, gap verdict, and a synthesis plan that feeds the introduction. Use when the user asks whether a contribution is novel, who else has studied a question, what the literature actually establishes, how to organize or restructure a review section, or wants a reading list, Zotero export, or pile of papers turned into a review. Produces a systematic-review protocol scaffold and hands off to registration (PROSPERO, OSF) and screening tools when the user asks for PRISMA or a meta-analysis.
- ▌ Survey Data Audit · scdenneyAudit fielded survey response data for registered elements, data quality, bot and AI-automation screening, and sample integrity. Emits an appendix-ready quality report.
- ▌ Fact Check · scdenney bundleFact-check manuscript claims against cited sources in a per-source Markdown knowledge base. Use to audit claim support, overclaiming, direction, scope, and misattribution after source intake is complete.
- ▌ Orchestrate · scdenney bundleOrchestrate complex work from an active gpt-6-astra or gpt-5.6-sol Codex session. The detected lead owns decomposition, integration, and verification; Astra keeps the hardest reasoning in the lead, while Sol escalates unusually difficult units to Astra. Both route bounded work to cheaper GPT-5.6 tiers and can use a cross-vendor Claude peer. Not for routine single-context work. Use when the user explicitly asks to orchestrate, delegate, fan out, parallelize, assign subagents, obtain independent checks, or have Codex act as tech lead.
- ▌ Qualtrics Ops · scdenney bundleOperate or audit a live Qualtrics survey via the v3 APIs without breaking fielding — publish gating, quotas, flow routing, embedded data, panel-vendor redirects, read-back verification, and a read-only pre-fielding audit. Use when publishing or patching a fielding instrument, when a quota counts but never blocks, when wiring panel-vendor redirects or flow gates, or when auditing a survey before launch (consent-before-anything gates, force-response completeness, quota and redirect checks, anti-bot instrumentation, and language-arm symmetry).
- ▌ Research Repo · scdenney bundleScaffold or audit an entire research project repository organized around its source library. Use when starting, structuring, organizing, or reviewing a research repo. Build sources/{og,md,unprocessed}, references.bib, a PDF-to-Markdown converter, a repo-local process-source skill, AGENTS.md, .gitignore, .venv, and the relevant analysis/manuscript/review folders; or audit an existing layout. Not for one-PDF intake or publication replication packages.
- ▌ Survey Design · scdenney bundleDesigns survey instruments — item-specific question wording that avoids acquiescence, double-barreling, and leading language, scale construction (points, labeling, polarity, feeling thermometers, index reliability with Cronbach's alpha and McDonald's omega), flow and ordering with treatment placement and buffer items, pretesting through cognitive interviews and soft launches, respondent burden with attention checks and speeding rules, sensitive questions and social-desirability bias, and treatment delivery with comprehension gates. Use when the user is writing or revising survey questions, choosing a scale, ordering blocks, planning a pilot, or worried respondents will not answer honestly. Indirect measurement goes to list-experiment, multi-country instruments to cross-national-design, and operating a live survey to qualtrics-ops.
- ▌ Citation Check · scdenney bundleAudits the citation layer of a manuscript — in-text and reference-list parity, fabricated or nonexistent sources, DOIs that resolve to a different work, style and completeness against APA 7 or a named journal style, and whether each cited work actually supports the claim attached to it. Verifies against Crossref, OpenAlex, DataCite, and Semantic Scholar, drives LaTeX audits off the keys actually cited, and marks anything it cannot check as NOT CHECKED rather than guessing. Use when the user asks to check citations or references, suspects an AI-invented source, wants a .bib checked against the text, or asks whether the DOIs are right. Figures and tables go to figure-table-audit.
- ▌ Journal Review · scdenney bundleDrafts a referee report on someone else's manuscript for a journal editor — a recommendation, a summary of the claim and design, three to six major concerns that drive the decision, additional concerns, and a coherent revision plan, produced by five parallel adversarial finders (Breaker, Butcher, Shredder, Void, Situator), a Blue Team error filter, a Chief Reviewer synthesis, and a Tone Guard legal pass, with an optional confidential note to the editor. Use when the user has been asked to referee a manuscript for a journal. Self-audit of the user's own draft goes to paper-review-lite or presubmit, and writing an author-side response to reviewers goes to referee-response.
- ▌ Topic Modeling · scdenney bundleSpecifies and diagnoses structural topic models for survey and experimental text — choosing among STM, LDA, and BERTopic, preprocessing decisions and their consequences, prevalence and content formulas with spectral initialization and a recorded seed, selecting the topic count across semantic coherence, exclusivity and FREX, held-out likelihood, and residuals rather than one metric, interpretation and validation against representative documents, robustness checks, and DA-RT-compliant reporting. Use when the user has open-ended responses or another corpus and asks what topics are in it, asks how many topics to use, asks about STM, searchK, coherence, or FREX, or wants topic prevalence compared across treatment arms or countries. Coding against a fixed codebook goes to text-classification.
- ▌ Conjoint Design · scdenney bundleDesigns conjoint and factorial-vignette experiments end to end — attribute architecture and randomization restrictions, effective-N power from the closed-form AMCE standard error, treatment realism, estimand choice among AMCE, marginal means, and AMIE, design variants such as forced choice versus rating, PAP tiers for conjoint flexibility, and the regression models and R packages that implement each. Use when the user is planning a conjoint, drafting or critiquing an attribute table, asking how many respondents or tasks are needed, asking whether to report AMCEs or marginal means, or asking how to test interactions. Reviewing an existing design goes to conjoint-diagnostics, cleaning the export to conjoint-cleaning.
- ▌ Doc To Markdown · scdenney bundleRead or convert any document a research workflow hands you — PDF, Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, or CSV. Use whenever a document has to be read, opened, quoted, summarized, searched, extracted, or added to a source library. Decides whether to read the file directly or convert it, picks the converter from the document's actual structure, and decides whether the resulting Markdown is a tracked artifact or a scratch file to delete.
- ▌ List Experiment · scdenney bundleDesigns and diagnoses list experiments (the item count technique) — whether indirect measurement is warranted at all, control-list construction against ceiling and floor effects, design variants such as double list and direct-question pairing, difference-in-means and maximum-likelihood estimators via ictreg, the no-design-effect and no-liar assumptions with ict.test and ict.hausman.test, corrections for common failures, and power planning that accounts for the large precision penalty. Use when the user is measuring a sensitive attitude or behavior and asks about list experiments, item count, unmatched count, or veiled questioning, or has one returning a negative or implausible prevalence estimate. Deciding whether sensitivity bias exists at all goes to survey-design.
- ▌ Model Committee · scdenney bundleRun one consequential, contestable decision through a deliberative committee of GPT-6 Astra and Claude Opus 5, chaired by fable (default), astra, opus, or sol. Not for factual lookups, brainstorming, routine implementation, independent-coder reliability, or final professional judgment. Use when several options are defensible and the two model families should propose, critique, revise, and cross-rank under a predeclared rubric.
- ▌ Vlm Ocr · scdenney bundleOCR a scanned or image-only corpus with vision-language models, in three phases — evaluate compares candidate OCR systems against stratified human ground truth and picks one on measured CER/WER; run builds the production pipeline (model selection, image handling, prompts, architecture, batching, accuracy evaluation, reproducibility); clean corrects the raw OCR text with LLM and rule-based passes, quality diagnostics, multilingual handling, and span-level provenance. Use when the question is which OCR or VLM to run on a corpus or how accurate one is on your own pages, when scanned documents have to be transcribed at scale, or when raw OCR output needs correction, QA, or a provenance log. Not for born-digital documents with a text layer — those go to doc-to-markdown.
- ▌ Conjoint Cleaning · scdenney bundleTurns a raw Qualtrics conjoint export into an analysis-ready long-format dataset — export settings and metadata rows, identifying which Qualtrics implementation produced the file, reshaping wide to long by hand or with cjoint read.qualtrics and cjdata reshape_conjoint, mapping the choice variable onto profiles, ratings and secondary DVs, attribute-level translation and factor ordering, pilot data-quality diagnostics, and subgroup merges. Use when the user has a conjoint CSV and asks how to clean, reshape, restructure, or debug it, cannot line the choice column up with the profiles, or asks what the data must look like before AMCE estimation. Design questions go to conjoint-design, validity review to conjoint-diagnostics.
- ▌ Methods Reporting · scdenney bundleAudits a methods section against a 45-item checklist synthesized from the APSA Experimental Section rubric, JARS-Quant, CONSORT, and DA-RT — pre-registration and design documentation, subjects and recruitment, randomization and treatment detail, CONSORT-style sample flow with attrition, statistical analysis with sample-size justification and three-tier results labeling, conjoint-specific reporting, the four validity types, and open-science infrastructure. Use when the user asks whether a methods section reports enough, is preparing for submission or a replication archive, asks what CONSORT, JARS, or DA-RT require, or asks how to document a deviation from the pre-analysis plan. Writing that plan beforehand goes to pre-registration-writing.
- ▌ Paper Review Lite · scdenney bundleRun a pre-submission manuscript audit covering argument, numerics, references, writing, figures, methods, preregistration, and replication readiness. Add --codex for a cross-model pass in which Codex and Claude review independently and verify each other's findings.
- ▌ Spawn · scdenney bundleSpawn full peer Codex or Claude sessions in their own panes and git worktrees — real sessions, not subagents — briefed by the lead and merged back. Not for bounded consults a subagent covers. Use when work must outlive or run beside this session, needs its own worktree, or must stay user-steerable.
- ▌ Tables · scdenney bundleDesign and format publication-quality tables. Use for column order, row grouping, notes, statistical precision, accessibility, and reproducibility.
- ▌ Advisor · scdenney bundleConsult an independent, read-only GPT-6 Astra advisor before committing to a substantive interpretation, approach, or final result. Not for routine work, implementation, or file edits. Defaults to gpt-6-astra at xhigh for demanding reviews.
- ▌ Diverge · scdenney bundleGenerate 3-5 conceptually distinct approaches labeled by creativity dimension (Novel, Surprising, Diverse, Conventional) and hold for selection instead of implementing the first idea; the --codex mode delegates the brainstorm to a fresh Codex subagent for a clean, unanchored context and hands it the selected approach to implement. Use when a task has more than one non-obvious solution — creative, architectural, or analytical work — and before committing to an approach; use --codex when the user explicitly asks for a subagent, an independent context, or a delegated brainstorm.
- ▌ Figures · scdenney bundleDesign and format publication-quality figures. Use for chart choice, color, scales, legends, captions, accessibility, and reproducible figure workflows.