# Python Bio Expressions

> Use Python arithmetic and comparison operators to compute GC content, codon/frame math, protein MW, and primer Tm. Use when calculating GC%, codon counts, reading frames, or fixing operator-precedence bugs in bio scripts.

- Skill: `pavel-kravchenko/python-bio-expressions` (Agent Skill)
- Install (CLI): `npx skillmds@latest add pavel-kravchenko/python-bio-expressions`
- Raw SKILL.md: https://api.skillmd.com/api/skills/pavel-kravchenko/python-bio-expressions/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: pavel-kravchenko (https://skillmd.com/u/pavel-kravchenko)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/pavel-kravchenko/python-bio-expressions

---


# Python Operators for Bioinformatics Calculations

## When to Use

- Computing GC content, codon counts, or reading-frame positions from sequence length/position math (`/`, `//`, `%`).
- Writing a sequence QC filter that chains length, GC%, and ambiguity checks (`and`, chained comparisons).
- Checking stop-codon or restriction-site membership efficiently (`in` on `set`/`frozenset` vs `list`).
- Estimating protein molecular weight or primer melting temperature from simple arithmetic formulas.
- Debugging a script that silently gives wrong numbers because of operator precedence (e.g. GC% off by a lot) or an `== None` / `is` mixup.

## Version Compatibility

Python ≥3.8 (uses f-strings; no third-party dependencies — pure stdlib). Applies to any CPython version currently in use; operator semantics have not changed across 3.x.

## Prerequisites

- Basic Python: variables, `def` functions, `for`/`if`.
- Related skills: `bio-sequence-manipulation-sequence-properties`, `bio-sequence-manipulation-codon-usage`.

## Codon and Reading-Frame Arithmetic

**Goal:** derive codon counts, leftover nucleotides, and the reading frame of a given position from a sequence length using integer arithmetic.
**Approach:** floor-divide (`//`) for whole codons, modulo (`%`) for the remainder/frame — never plain `/`, which returns a float.

```python
def codon_frame_report(seq_length: int, position: int) -> dict:
    """Report codon-count and reading-frame stats for a CDS length and a query position (0-based).

    seq_length // 3  -> number of complete codons
    seq_length % 3   -> leftover (partial) nucleotides
    position % 3     -> which reading frame (0, 1, 2) the position falls in
    position // 3    -> index of the codon containing that position
    """
    return {
        "complete_codons": seq_length // 3,
        "leftover_nt": seq_length % 3,
        "frame": position % 3,
        "codon_index": position // 3,
    }

result = codon_frame_report(1000, 47)
print(result)  # -> complete_codons=333, leftover_nt=1, frame=2, codon_index=15
```

## GC Content and Classification

**Goal:** compute GC% and bucket it into a biologically meaningful category.
**Approach:** count bases, divide by total, guard the empty-sequence edge case, then branch on thresholds.

```python
def gc_content(seq: str) -> float:
    """Return GC percentage (0-100) of a nucleotide sequence."""
    seq = seq.upper()
    total = len(seq)
    if total == 0:
        return 0.0
    return (seq.count("G") + seq.count("C")) / total * 100

def classify_gc(gc_pct: float) -> str:
    """Bucket a GC percentage into a rough taxonomic-flavor category."""
    if gc_pct < 30:
        return "Very AT-rich (e.g. Plasmodium)"
    if gc_pct < 40:
        return "AT-rich"
    if gc_pct < 60:
        return "Moderate"
    if gc_pct < 70:
        return "GC-rich"
    return "Very GC-rich (e.g. Streptomyces)"

seq = "ATGCGATCGATCGTAGC"
pct = gc_content(seq)
print(f"{pct:.2f}% -> {classify_gc(pct)}")
```

## Sequence QC Filter, Stop-Codon Check, and Identity

**Goal:** gate sequences on multiple numeric/logical criteria before downstream analysis, and score two aligned sequences for percent identity.
**Approach:** chained comparisons (`a <= x <= b`) read cleanly and short-circuit like `and`; use a `frozenset` for O(1) membership on stop codons; use `zip` to walk equal-length sequences pairwise.

```python
STOP_CODONS = frozenset({"TAA", "TAG", "TGA"})

def passes_qc(seq_length: int, gc_pct: float, has_ambiguous_bases: bool = False) -> bool:
    """QC gate: length 500-3000 bp, GC% 30-70, and no ambiguous (non-ACGT) bases."""
    return (
        500 <= seq_length <= 3000
        and 30.0 <= gc_pct <= 70.0
        and not has_ambiguous_bases
    )

def sequence_identity(seq_a: str, seq_b: str) -> float:
    """Percent identity between two equal-length sequences."""
    if len(seq_a) != len(seq_b):
        raise ValueError("sequences must be the same length")
    matches = sum(1 for a, b in zip(seq_a, seq_b) if a == b)
    return matches / len(seq_a) * 100

codon = "TAG"
print(codon in STOP_CODONS)                 # O(1) set lookup, not `in ["TAA","TAG","TGA"]`
print(passes_qc(1500, 48.0))
print(f"{sequence_identity('ATGCGATC', 'ATGCAATC'):.1f}%")
```

## Protein MW and Primer Tm

**Goal:** estimate protein molecular weight from residue count, and primer melting temperature from base composition.
**Approach:** MW ≈ 110 Da/AA minus 18 Da per peptide bond; Tm uses the Wallace rule for short oligos (<14 nt) and a salt-adjusted formula for longer ones.

```python
def protein_mw_da(n_aa: int) -> float:
    """Rough protein MW in Daltons: ~110 Da/residue minus water lost per peptide bond."""
    return n_aa * 110 - (n_aa - 1) * 18

def melting_temperature(primer: str) -> float:
    """Estimate primer Tm in Celsius (Wallace rule <14 nt, salt-adjusted formula otherwise)."""
    primer = primer.upper()
    a, t = primer.count("A"), primer.count("T")
    g, c = primer.count("G"), primer.count("C")
    if len(primer) < 14:
        return 2 * (a + t) + 4 * (g + c)
    return 64.9 + 41 * (g + c - 16.4) / (a + t + g + c)

print(f"p53 (393 aa): {protein_mw_da(393):,.0f} Da")
print(f"Tm: {melting_temperature('ATGCGATCGATCGATCGATCG'):.1f} C")
```

## Pitfalls

- **`/` vs `//`**: `/` always returns a float; `//` truncates. `seq_length // 3` gives complete codons — `seq_length / 3` does not.
- **Precedence**: `g + c / total * 100` computes `g + ((c / total) * 100)` — wrong. Use `(g + c) / total * 100`. `*`/`/`/`//`/`%` bind tighter than `+`/`-`; comparisons bind tighter than `not` > `and` > `or`. Parenthesize when mixing.
- **`==` vs `is`**: use `==` for value equality; reserve `is` for identity (`is None`, `is True`). Never `if x == None`.
- **`and`/`or` short-circuit**: `A and B` skips `B` when `A` is falsy — can hide bugs if `B` has side effects.
- **`in` on a dict checks keys, not values**: `"ATG" in codon_table` tests keys only; check `codon_table.values()` (O(n)) or invert the dict for value lookups.
- **`in` on list vs set**: list membership is O(n); set/`frozenset` membership is O(1). Use a `frozenset` for stop-codon or restriction-site lookups over large collections.
- **Equal-length assumption**: `sequence_identity`/`zip`-based comparisons silently truncate to the shorter sequence if you don't check lengths first — this skill's version raises instead.

## See Also

- `bio-sequence-manipulation-sequence-properties`
- `bio-sequence-manipulation-codon-usage`
- `bio-primer-design-primer-basics`
- `bio-restriction-analysis-restriction-sites`

