github-repo-skill
overview
To create and maintain high-quality repos that conform to Mungall group / BBOP best practice, use this skill. Use this skill regardless of whether the repo is for code or non-code (ontology, linkml schemas, curated content, analyses, websites). Use this skill for both new repos, for migrating legacy repos, or for ongoing maintenance.
Principles
Follow existing copier templates
The Mungall group favors the use of copier and blesses the following templates:
These should always be used for new repos. Pre-existing repos should try and follow these or migrate towards them.
Additionally the group uses additional drop-in templates for AI integrations:
Favored tools
These are included in the templates above but some general over-arching preferences:
- modern python dev stack:
uv, ruff (currently mypy for typing but we may switch to https://docs.astral.sh/ty/)
- for builds, both
just and make are favored, with just favored for non-pipeline cases
Engineering best practice
- pydantic or pydantic generated from LinkML for data models and data access objects (dataclasses are fine for engine objects)
- always make sure function/method parameters and return objects are typed. Use mypy or equivalent to test.
- testing:
- follow TDD, use pytest-style tests,
@pytest.mark.parametrize is good for combinatorial testing
- always use doctests: make them informative for humans but also serving as additional tests
- ensure unit tests and tests that depend on external APIs, infrastructure etc are separated (e.g.
pytest.mark.integration)
- for testing external APIs or services, use vcrpy
- do not create mock tests unless explicitly requested
- for data-oriented projects, yaml, tsvs, etc can go in
tests/input or smilar
- for schemas, follow the linkml copier template, and ensure schemas and example test data is validated
- for ontologies, follow ODK best practice and ensure ontologies are comprehensively axiomatized to allow for reasoner-based checking
- jupyter notebooks are good for documentation, dashboards, and analysis, but ensure that core logic is separated out and has unit tests
- CLI:
- Every library should have a fully featured CLI
- typer is favored, but click is also good.
- CLIs, APIs, and MCPs should be shims on top of core logic
- have separate test for both core logic and CLIs.
- Use idiomatic options and clig conventions. Group standards:
-i/--input, -o/--output (default stdout), -f/--format (input format), -O/--output-format, -v/-vv
- When testing Typer/Rich CLIs, set
NO_COLOR=1 and TERM=dumb env vars to avoid ANSI escape codes breaking string assertions in CI.
- Exceptions
- In general you should not need to worry about catching exceptions, although for a well-polished CLI some catching in the CLI layer is fine
- IMPORTANT: better to fail fast and know there is a problem than to defensively catch and carry on as if everything is OK (general principle: transparency)
Dependency management
uv add to add new dependencies (or uv add --dev or similar for dev dependencies)
- libraries should allow somewhat relaxed dependencies to avoid diamond dependency problems. applications and infra can pin more tightly.
Git and GitHub Practices
- always work on branches, commit early and often, make PRs early
- in general, one PR = one issue (avoid mixing orthogonal concerns). Always reference issues in commits/PR messages
- use
gh on command line for operations like finding issues, creating PRs
- all the group copier templates include extensive github actions for ensuring PRs are high quality
- github repo should have adequate metadata, links to docs, tags
Source of truth
- always have a clear source of truth (SoT) for all content, with github favored
- where projects dictate SoT is google docs/sheets, use https://rclone.org/ to sync
Documentation
- markdown is always favored, but older sites may use sphinx
- Follow Diátaxis framework: tutorial, how-to, reference, explanation
- Use examples extensively - examples can double as tests
- frameworks: mkdocs is generally favored due to simplicity but sphinx is ok for older projects
- Every project must have a comprehensive up to date README.md (or the README.md can point to site generated from mkdocs)
- jupyter notebooks can serve as combined integration tests/docs, use mkdocs-jupyter, for CLI examples, use
%%bash
- Formatting tips: lists should be preceded by a blank line to avoid formatting issues with mkdocs
1---2name: github-repo-skill3description: Guide for creating new GitHub repos and best practice for existing GitHub repos, applicable to both code and non-code projects4license: CC-05---67# github-repo-skill89## overview1011To create and maintain high-quality repos that conform to Mungall group / BBOP best practice, use this skill. Use this skill regardless of whether the repo is for code or non-code (ontology, linkml schemas, curated content, analyses, websites). Use this skill for both new repos, for migrating legacy repos, or for ongoing maintenance.1213---1415# Principles1617## Follow existing copier templates1819The Mungall group favors the use of [copier](https://copier.readthedocs.io/) and blesses the following templates:2021* For LinkML schemas: https://github.com/linkml/linkml-project-copier22* For code: https://github.com/monarch-initiative/monarch-project-copier23* For ontologies: https://github.com/INCATools/ontology-development-kit (uses bespoke framework, not copier)2425These should always be used for new repos. Pre-existing repos should try and follow these or migrate towards them.2627Additionally the group uses additional drop-in templates for AI integrations:2829* https://github.com/ai4curation/github-ai-integrations3031## Favored tools3233These are included in the templates above but some general over-arching preferences:3435* modern python dev stack: `uv`, `ruff` (currently `mypy` for typing but we may switch to https://docs.astral.sh/ty/)36* for builds, both `just` and `make` are favored, with `just` favored for non-pipeline cases3738## Engineering best practice3940* pydantic or pydantic generated from LinkML for data models and data access objects (dataclasses are fine for engine objects)41* always make sure function/method parameters and return objects are typed. Use mypy or equivalent to test.42* testing:43 * follow TDD, use pytest-style tests, `@pytest.mark.parametrize` is good for combinatorial testing44 * always use doctests: make them informative for humans but also serving as additional tests45 * ensure unit tests and tests that depend on external APIs, infrastructure etc are separated (e.g. `pytest.mark.integration`)46 * for testing external APIs or services, use vcrpy47 * do not create mock tests unless explicitly requested48 * for data-oriented projects, yaml, tsvs, etc can go in `tests/input` or smilar49 * for schemas, follow the linkml copier template, and ensure schemas and example test data is validated50 * for ontologies, follow ODK best practice and ensure ontologies are comprehensively axiomatized to allow for reasoner-based checking51* jupyter notebooks are good for documentation, dashboards, and analysis, but ensure that core logic is separated out and has unit tests52* CLI:53 * Every library should have a fully featured CLI54 * typer is favored, but click is also good.55 * CLIs, APIs, and MCPs should be shims on top of core logic56 * have separate test for both core logic and CLIs.57 * Use idiomatic options and clig conventions. Group standards: `-i/--input`, `-o/--output` (default stdout), `-f/--format` (input format), `-O/--output-format`, `-v/-vv`58 * When testing Typer/Rich CLIs, set `NO_COLOR=1` and `TERM=dumb` env vars to avoid ANSI escape codes breaking string assertions in CI.59* Exceptions60 * In general you should not need to worry about catching exceptions, although for a well-polished CLI some catching in the CLI layer is fine61 * IMPORTANT: better to fail fast and know there is a problem than to defensively catch and carry on as if everything is OK (general principle: transparency)6263## Dependency management6465* `uv add` to add new dependencies (or `uv add --dev` or similar for dev dependencies)66* libraries should allow somewhat relaxed dependencies to avoid diamond dependency problems. applications and infra can pin more tightly.6768## Git and GitHub Practices6970* always work on branches, commit early and often, make PRs early71* in general, one PR = one issue (avoid mixing orthogonal concerns). Always reference issues in commits/PR messages72* use `gh` on command line for operations like finding issues, creating PRs73* all the group copier templates include extensive github actions for ensuring PRs are high quality74* github repo should have adequate metadata, links to docs, tags7576## Source of truth7778* always have a clear source of truth (SoT) for all content, with github favored79* where projects dictate SoT is google docs/sheets, use https://rclone.org/ to sync8081## Documentation8283* markdown is always favored, but older sites may use sphinx84* Follow [Diátaxis framework](https://diataxis.fr/): tutorial, how-to, reference, explanation85* Use examples extensively - examples can double as tests86* frameworks: mkdocs is generally favored due to simplicity but sphinx is ok for older projects87* Every project must have a comprehensive up to date README.md (or the README.md can point to site generated from mkdocs)88* jupyter notebooks can serve as combined integration tests/docs, use mkdocs-jupyter, for CLI examples, use `%%bash`89* Formatting tips: lists should be preceded by a blank line to avoid formatting issues with mkdocs