# Python Excellence

> Enforces principal-engineer-level Python code quality: formatting, type safety, idiomatic patterns, class and function design, and anti-pattern prevention.

- Skill: `paramchordiya/python-excellence` (Agent Skill)
- Install (CLI): `npx skillmds@latest add paramchordiya/python-excellence`
- Raw SKILL.md: https://api.skillmd.com/api/skills/paramchordiya/python-excellence/raw
- Safety review: PASS (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools, Code Review, Refactoring
- Tags: Best Practices, Black, Code Quality, Dataclasses, Isort, Python, Ruff, Type Annotations
- Author: ParamChordiya (https://skillmd.com/u/paramchordiya)
- Updated: 2026-08-22
- Page: https://skillmd.com/skills/paramchordiya/python-excellence

---


## Intent

This skill ensures that every Python file produced or modified in this session meets the bar of a principal engineer's code review: fully typed, black-formatted, idiomatic, and free of the anti-patterns that cause subtle production bugs. It enforces the modern Python style (3.10+) for all new code, treats type annotations as a correctness tool rather than documentation, and replaces every "it works" pattern with the pattern that is readable, maintainable, and provably correct. Loading this skill means no Python leaves the session untyped, unformatted, or using a Python 2-era idiom when a modern equivalent exists.

See also: [`software-architecture/SKILL.md`](../software-architecture/SKILL.md) for project structure and module design, [`code-review-standards/SKILL.md`](../code-review-standards/SKILL.md) for the self-review checklist applied to all output.

---

## 1. Code Style and Formatting

**All code must be formatted with `black` (line length 88) and imports sorted with `isort` before being considered correct.** Any code that `black` would reformat is not finished.

**All linting must pass `ruff` with zero warnings.** Configure in `pyproject.toml`:

```toml
[tool.black]
line-length = 88
target-version = ["py311"]

[tool.isort]
profile = "black"
line_length = 88

[tool.ruff]
line-length = 88
target-version = "py311"
select = ["E", "F", "W", "I", "N", "UP", "B", "SIM", "RUF"]
ignore = []

[tool.ruff.per-file-ignores]
"tests/**" = ["S101"]  # allow assert in tests
```

**All tool configuration lives in `pyproject.toml`. Never use scattered config files** (`.flake8`, `setup.cfg`, `.isort.cfg`).

---

## 2. Type Annotations — Mandatory

**Every function signature must have full type annotations: all parameters and the return type.**

```python
# WRONG — untyped
def process_batch(records, batch_size, callback):
    ...

# CORRECT — fully typed
def process_batch(
    records: list[dict[str, Any]],
    batch_size: int,
    callback: Callable[[list[dict[str, Any]]], None],
) -> int:
    ...
```

**Add `from __future__ import annotations` at the top of every file** that uses forward references or wants the lazy evaluation behavior.

**Use modern Python 3.10+ union syntax for new code:**
```python
# Python 3.10+ (preferred)
def find_user(user_id: int) -> User | None: ...

# Python <3.10 (only if supporting older versions)
from typing import Optional
def find_user(user_id: int) -> Optional[User]: ...
```

**Use built-in generic types, not `typing` aliases, for Python 3.9+:**
```python
# WRONG (old style)
from typing import List, Dict, Tuple, Set
def process(items: List[str]) -> Dict[str, int]: ...

# CORRECT (modern style)
def process(items: list[str]) -> dict[str, int]: ...
```

**Use `TypeVar` and `Generic` for generic containers. Use `Protocol` for structural typing:**

```python
from typing import TypeVar, Generic, Protocol

T = TypeVar("T")

class Repository(Protocol[T]):
    def find_by_id(self, id: int) -> T | None: ...
    def save(self, entity: T) -> None: ...

class Stack(Generic[T]):
    def __init__(self) -> None:
        self._items: list[T] = []

    def push(self, item: T) -> None:
        self._items.append(item)

    def pop(self) -> T:
        return self._items.pop()
```

**Never use `Any` without a `# type: ignore` comment or an inline comment explaining why the type cannot be specified.** `Any` is a last resort, not a default.

**For NumPy arrays, use proper annotations:**
```python
import numpy as np
import numpy.typing as npt

def normalize(
    X: npt.NDArray[np.float64],
    axis: int = 0,
) -> npt.NDArray[np.float64]: ...
```

---

## 3. Pythonic Patterns — Always Enforce

**Use comprehensions over explicit loops for simple transformations:**
```python
# WRONG
squared = []
for x in numbers:
    if x > 0:
        squared.append(x ** 2)

# CORRECT
squared = [x ** 2 for x in numbers if x > 0]
```

**Use `dataclasses` or Pydantic models over plain dicts for structured data:**
```python
# WRONG — opaque dict, no type safety
user = {"id": 1, "name": "Alice", "email": "alice@example.com"}

# CORRECT — self-documenting, typed
@dataclass
class User:
    id: int
    name: str
    email: str
```

**Use context managers for all resource management:**
```python
# WRONG — file may not be closed on exception
f = open("data.csv")
data = f.read()
f.close()

# CORRECT — guaranteed cleanup
with open("data.csv") as f:
    data = f.read()
```

**Use `pathlib.Path` for all file system operations. Never use `os.path`:**
```python
# WRONG
import os
path = os.path.join(base_dir, "data", "raw", "file.csv")
if os.path.exists(path):
    with open(path) as f: ...

# CORRECT
from pathlib import Path
path = Path(base_dir) / "data" / "raw" / "file.csv"
if path.exists():
    with path.open() as f: ...
```

**Use f-strings exclusively for new code. Never use `%` formatting or `.format()` in new code:**
```python
# WRONG
msg = "User %s has %d items" % (name, count)
msg = "User {} has {} items".format(name, count)

# CORRECT
msg = f"User {name} has {count} items"
```

**Use `enumerate()` over manual index tracking:**
```python
# WRONG
for i in range(len(items)):
    print(f"{i}: {items[i]}")

# CORRECT
for i, item in enumerate(items):
    print(f"{i}: {item}")
```

**Use `zip()` for parallel iteration. Use `zip(..., strict=True)` in Python 3.10+ to catch length mismatches:**
```python
# WRONG — silent bug if lengths differ
for name, score in zip(names, scores):
    ...

# CORRECT — raises ValueError if lengths differ (Python 3.10+)
for name, score in zip(names, scores, strict=True):
    ...
```

**Use the walrus operator (`:=`) only when it eliminates a repeated expression and improves readability:**
```python
# Good use — eliminates double computation
if (match := pattern.search(text)) is not None:
    print(match.group(0))

# Bad use — makes code harder to read, not easier
result = [y := f(x), y**2, y**3]  # obscure
```

---

## 4. Class Design

**Use `@dataclass` for classes that primarily hold data:**
```python
from dataclasses import dataclass, field

@dataclass
class TrainingConfig:
    learning_rate: float = 1e-3
    batch_size: int = 32
    max_epochs: int = 100
    tags: list[str] = field(default_factory=list)  # correct mutable default
```

**Use `@dataclass(frozen=True)` for immutable value objects:**
```python
@dataclass(frozen=True)
class ModelVersion:
    name: str
    version: str
    git_hash: str

    def __str__(self) -> str:
        return f"{self.name}@{self.version}+{self.git_hash[:8]}"
```

**Use `__slots__` for classes with many instances where memory efficiency matters:**
```python
@dataclass
class Point:
    __slots__ = ("x", "y")
    x: float
    y: float
```

**Use `@property` for computed attributes; use `@functools.cached_property` for expensive computed attributes:**
```python
class BoundingBox:
    def __init__(self, x1: float, y1: float, x2: float, y2: float) -> None:
        self.x1, self.y1, self.x2, self.y2 = x1, y1, x2, y2

    @property
    def area(self) -> float:
        return (self.x2 - self.x1) * (self.y2 - self.y1)

    @functools.cached_property
    def diagonal(self) -> float:
        # Expensive: cache after first access
        return math.sqrt((self.x2 - self.x1)**2 + (self.y2 - self.y1)**2)
```

**Implement `__repr__` for every class that will appear in logs or debugging:**
```python
def __repr__(self) -> str:
    return f"TrainingConfig(lr={self.learning_rate}, batch={self.batch_size}, epochs={self.max_epochs})"
```

**Implement `__eq__` and `__hash__` consistently.** If you implement `__eq__`, you must implement `__hash__` or explicitly set `__hash__ = None`. Dataclasses handle this automatically when `frozen=True`.

---

## 5. Function Design

**Maximum function length: 40 lines.** If a function exceeds this, split it into well-named helper functions. The split should be at a semantic boundary, not an arbitrary line count.

**Maximum number of parameters: 5.** If a function needs more than 5 parameters, group related parameters into a config dataclass:

```python
# WRONG — too many parameters
def train_model(X, y, lr, batch_size, epochs, dropout, weight_decay, patience, seed, device):
    ...

# CORRECT — grouped into a config
def train_model(X: np.ndarray, y: np.ndarray, config: TrainingConfig) -> TrainedModel:
    ...
```

**Never use boolean parameters that switch behavior.** Split into two functions:

```python
# WRONG — boolean flag changes behavior
def process(data, validate=True):
    if validate:
        ...
    else:
        ...

# CORRECT — two explicit functions
def process_with_validation(data: pd.DataFrame) -> pd.DataFrame: ...
def process_without_validation(data: pd.DataFrame) -> pd.DataFrame: ...
```

**Use `*` to enforce keyword-only arguments for clarity in public APIs:**

```python
# Without *, callers can write train(data, 0.01, 32) — what are 0.01 and 32?
# With *, callers must write train(data, learning_rate=0.01, batch_size=32)
def train(
    data: pd.DataFrame,
    *,
    learning_rate: float,
    batch_size: int,
    epochs: int = 100,
) -> TrainedModel: ...
```

---

## 6. Async Python

**Use `asyncio` for I/O-bound concurrency. Document why synchronous vs. asynchronous was chosen for any service or function:**

```python
# Async chosen: this function calls three external APIs whose latency is independent.
# Running them concurrently reduces total latency from ~900ms to ~300ms.
async def enrich_user(user_id: int) -> EnrichedUser:
    profile, orders, recommendations = await asyncio.gather(
        fetch_profile(user_id),
        fetch_orders(user_id),
        fetch_recommendations(user_id),
    )
    return EnrichedUser(profile=profile, orders=orders, recommendations=recommendations)
```

**Never mix sync and async without an explicit bridge.** Use `asyncio.run()` at the entry point or `loop.run_in_executor()` for blocking calls inside async code.

**Always handle `asyncio.CancelledError` in long-running async tasks:**

```python
async def long_running_job() -> None:
    try:
        while True:
            await process_next_batch()
            await asyncio.sleep(1)
    except asyncio.CancelledError:
        await cleanup()  # flush buffers, close connections
        raise  # always re-raise CancelledError
```

---

## 7. Package and Module Design

**Every public module must define `__all__`** to explicitly declare its public interface:

```python
# src/mypackage/services/__init__.py
__all__ = ["UserService", "TransactionService", "NotificationService"]

from mypackage.services.user_service import UserService
from mypackage.services.transaction_service import TransactionService
from mypackage.services.notification_service import NotificationService
```

**`__init__.py` must only re-export. Never put logic in `__init__.py`.** The `__init__.py` is a public API declaration file, not a script.

**Use `importlib` for dynamic imports and document why they are necessary:**

```python
import importlib

# Dynamic import required: plugin architecture where plugin names are user-configured.
# Static imports would require modifying this file for each new plugin.
plugin_module = importlib.import_module(f"mypackage.plugins.{plugin_name}")
```

**Circular imports are a design smell, not a Python limitation to work around.** If you encounter a circular import, refactor the modules to break the cycle by extracting shared types into a third module. Never use `TYPE_CHECKING` hacks to paper over circular dependencies caused by bad design — `TYPE_CHECKING` is acceptable only for type-only imports that have no runtime impact.

---

## 8. Anti-Patterns — Always Refuse These

**Never use mutable default arguments.** Python evaluates default argument values once at function definition time.

```python
# WRONG — all callers share the same list object
def add_item(item: str, items: list[str] = []) -> list[str]:
    items.append(item)
    return items

# CORRECT
def add_item(item: str, items: list[str] | None = None) -> list[str]:
    if items is None:
        items = []
    items.append(item)
    return items

# OR use dataclasses with field(default_factory=list)
```

**Never use bare `except:` or `except Exception:` without re-raising or specific handling.** Always catch specific exception types.

**Never use `global` or `nonlocal` without a comment explaining why encapsulation cannot solve the problem.**

**Never use `eval()` or `exec()` in production code.** These introduce security vulnerabilities (code injection) and make code impossible to analyze statically.

**Never use `assert` for input validation in production code.** Assertions can be disabled with `python -O`. Use explicit `if`/`raise` for validation:

```python
# WRONG — silently disabled in optimized mode
assert user_id > 0, "user_id must be positive"

# CORRECT — always enforced
if user_id <= 0:
    raise ValueError(f"user_id must be positive, got {user_id}")
```

**Never shadow built-in names.** The following names are forbidden as variable, parameter, or function names: `list`, `dict`, `set`, `tuple`, `type`, `id`, `input`, `print`, `min`, `max`, `sum`, `len`, `range`, `filter`, `map`, `zip`, `open`, `hash`, `format`, `object`, `str`, `int`, `float`, `bool`, `bytes`.

**Never use `type(x) == SomeType` for type checking. Use `isinstance(x, SomeType)` instead** — it handles subclasses correctly.

