Intent
This skill ensures that every Python file produced or modified in this session meets the bar of a principal engineer's code review: fully typed, black-formatted, idiomatic, and free of the anti-patterns that cause subtle production bugs. It enforces the modern Python style (3.10+) for all new code, treats type annotations as a correctness tool rather than documentation, and replaces every "it works" pattern with the pattern that is readable, maintainable, and provably correct. Loading this skill means no Python leaves the session untyped, unformatted, or using a Python 2-era idiom when a modern equivalent exists.
See also: software-architecture/SKILL.md for project structure and module design, code-review-standards/SKILL.md for the self-review checklist applied to all output.
1. Code Style and Formatting
All code must be formatted with black (line length 88) and imports sorted with isort before being considered correct. Any code that black would reformat is not finished.
All linting must pass ruff with zero warnings. Configure in pyproject.toml:
[tool.black]
line-length = 88
target-version = ["py311"]
[tool.isort]
profile = "black"
line_length = 88
[tool.ruff]
line-length = 88
target-version = "py311"
select = ["E", "F", "W", "I", "N", "UP", "B", "SIM", "RUF"]
ignore = []
[tool.ruff.per-file-ignores]
"tests/**" = ["S101"] # allow assert in tests
All tool configuration lives in pyproject.toml. Never use scattered config files (.flake8, setup.cfg, .isort.cfg).
2. Type Annotations — Mandatory
Every function signature must have full type annotations: all parameters and the return type.
# WRONG — untyped
def process_batch(records, batch_size, callback):
...
# CORRECT — fully typed
def process_batch(
records: list[dict[str, Any]],
batch_size: int,
callback: Callable[[list[dict[str, Any]]], None],
) -> int:
...
Add from __future__ import annotations at the top of every file that uses forward references or wants the lazy evaluation behavior.
Use modern Python 3.10+ union syntax for new code:
# Python 3.10+ (preferred)
def find_user(user_id: int) -> User | None: ...
# Python <3.10 (only if supporting older versions)
from typing import Optional
def find_user(user_id: int) -> Optional[User]: ...
Use built-in generic types, not typing aliases, for Python 3.9+:
# WRONG (old style)
from typing import List, Dict, Tuple, Set
def process(items: List[str]) -> Dict[str, int]: ...
# CORRECT (modern style)
def process(items: list[str]) -> dict[str, int]: ...
Use TypeVar and Generic for generic containers. Use Protocol for structural typing:
from typing import TypeVar, Generic, Protocol
T = TypeVar("T")
class Repository(Protocol[T]):
def find_by_id(self, id: int) -> T | None: ...
def save(self, entity: T) -> None: ...
class Stack(Generic[T]):
def __init__(self) -> None:
self._items: list[T] = []
def push(self, item: T) -> None:
self._items.append(item)
def pop(self) -> T:
return self._items.pop()
Never use Any without a # type: ignore comment or an inline comment explaining why the type cannot be specified. Any is a last resort, not a default.
For NumPy arrays, use proper annotations:
import numpy as np
import numpy.typing as npt
def normalize(
X: npt.NDArray[np.float64],
axis: int = 0,
) -> npt.NDArray[np.float64]: ...
3. Pythonic Patterns — Always Enforce
Use comprehensions over explicit loops for simple transformations:
# WRONG
squared = []
for x in numbers:
if x > 0:
squared.append(x ** 2)
# CORRECT
squared = [x ** 2 for x in numbers if x > 0]
Use dataclasses or Pydantic models over plain dicts for structured data:
# WRONG — opaque dict, no type safety
user = {"id": 1, "name": "Alice", "email": "alice@example.com"}
# CORRECT — self-documenting, typed
@dataclass
class User:
id: int
name: str
email: str
Use context managers for all resource management:
# WRONG — file may not be closed on exception
f = open("data.csv")
data = f.read()
f.close()
# CORRECT — guaranteed cleanup
with open("data.csv") as f:
data = f.read()
Use pathlib.Path for all file system operations. Never use os.path:
# WRONG
import os
path = os.path.join(base_dir, "data", "raw", "file.csv")
if os.path.exists(path):
with open(path) as f: ...
# CORRECT
from pathlib import Path
path = Path(base_dir) / "data" / "raw" / "file.csv"
if path.exists():
with path.open() as f: ...
Use f-strings exclusively for new code. Never use % formatting or .format() in new code:
# WRONG
msg = "User %s has %d items" % (name, count)
msg = "User {} has {} items".format(name, count)
# CORRECT
msg = f"User {name} has {count} items"
Use enumerate() over manual index tracking:
# WRONG
for i in range(len(items)):
print(f"{i}: {items[i]}")
# CORRECT
for i, item in enumerate(items):
print(f"{i}: {item}")
Use zip() for parallel iteration. Use zip(..., strict=True) in Python 3.10+ to catch length mismatches:
# WRONG — silent bug if lengths differ
for name, score in zip(names, scores):
...
# CORRECT — raises ValueError if lengths differ (Python 3.10+)
for name, score in zip(names, scores, strict=True):
...
Use the walrus operator (:=) only when it eliminates a repeated expression and improves readability:
# Good use — eliminates double computation
if (match := pattern.search(text)) is not None:
print(match.group(0))
# Bad use — makes code harder to read, not easier
result = [y := f(x), y**2, y**3] # obscure
4. Class Design
Use @dataclass for classes that primarily hold data:
from dataclasses import dataclass, field
@dataclass
class TrainingConfig:
learning_rate: float = 1e-3
batch_size: int = 32
max_epochs: int = 100
tags: list[str] = field(default_factory=list) # correct mutable default
Use @dataclass(frozen=True) for immutable value objects:
@dataclass(frozen=True)
class ModelVersion:
name: str
version: str
git_hash: str
def __str__(self) -> str:
return f"{self.name}@{self.version}+{self.git_hash[:8]}"
Use __slots__ for classes with many instances where memory efficiency matters:
@dataclass
class Point:
__slots__ = ("x", "y")
x: float
y: float
Use @property for computed attributes; use @functools.cached_property for expensive computed attributes:
class BoundingBox:
def __init__(self, x1: float, y1: float, x2: float, y2: float) -> None:
self.x1, self.y1, self.x2, self.y2 = x1, y1, x2, y2
@property
def area(self) -> float:
return (self.x2 - self.x1) * (self.y2 - self.y1)
@functools.cached_property
def diagonal(self) -> float:
# Expensive: cache after first access
return math.sqrt((self.x2 - self.x1)**2 + (self.y2 - self.y1)**2)
Implement __repr__ for every class that will appear in logs or debugging:
def __repr__(self) -> str:
return f"TrainingConfig(lr={self.learning_rate}, batch={self.batch_size}, epochs={self.max_epochs})"
Implement __eq__ and __hash__ consistently. If you implement __eq__, you must implement __hash__ or explicitly set __hash__ = None. Dataclasses handle this automatically when frozen=True.
5. Function Design
Maximum function length: 40 lines. If a function exceeds this, split it into well-named helper functions. The split should be at a semantic boundary, not an arbitrary line count.
Maximum number of parameters: 5. If a function needs more than 5 parameters, group related parameters into a config dataclass:
# WRONG — too many parameters
def train_model(X, y, lr, batch_size, epochs, dropout, weight_decay, patience, seed, device):
...
# CORRECT — grouped into a config
def train_model(X: np.ndarray, y: np.ndarray, config: TrainingConfig) -> TrainedModel:
...
Never use boolean parameters that switch behavior. Split into two functions:
# WRONG — boolean flag changes behavior
def process(data, validate=True):
if validate:
...
else:
...
# CORRECT — two explicit functions
def process_with_validation(data: pd.DataFrame) -> pd.DataFrame: ...
def process_without_validation(data: pd.DataFrame) -> pd.DataFrame: ...
Use * to enforce keyword-only arguments for clarity in public APIs:
# Without *, callers can write train(data, 0.01, 32) — what are 0.01 and 32?
# With *, callers must write train(data, learning_rate=0.01, batch_size=32)
def train(
data: pd.DataFrame,
*,
learning_rate: float,
batch_size: int,
epochs: int = 100,
) -> TrainedModel: ...
6. Async Python
Use asyncio for I/O-bound concurrency. Document why synchronous vs. asynchronous was chosen for any service or function:
# Async chosen: this function calls three external APIs whose latency is independent.
# Running them concurrently reduces total latency from ~900ms to ~300ms.
async def enrich_user(user_id: int) -> EnrichedUser:
profile, orders, recommendations = await asyncio.gather(
fetch_profile(user_id),
fetch_orders(user_id),
fetch_recommendations(user_id),
)
return EnrichedUser(profile=profile, orders=orders, recommendations=recommendations)
Never mix sync and async without an explicit bridge. Use asyncio.run() at the entry point or loop.run_in_executor() for blocking calls inside async code.
Always handle asyncio.CancelledError in long-running async tasks:
async def long_running_job() -> None:
try:
while True:
await process_next_batch()
await asyncio.sleep(1)
except asyncio.CancelledError:
await cleanup() # flush buffers, close connections
raise # always re-raise CancelledError
7. Package and Module Design
Every public module must define __all__ to explicitly declare its public interface:
# src/mypackage/services/__init__.py
__all__ = ["UserService", "TransactionService", "NotificationService"]
from mypackage.services.user_service import UserService
from mypackage.services.transaction_service import TransactionService
from mypackage.services.notification_service import NotificationService
__init__.py must only re-export. Never put logic in __init__.py. The __init__.py is a public API declaration file, not a script.
Use importlib for dynamic imports and document why they are necessary:
import importlib
# Dynamic import required: plugin architecture where plugin names are user-configured.
# Static imports would require modifying this file for each new plugin.
plugin_module = importlib.import_module(f"mypackage.plugins.{plugin_name}")
Circular imports are a design smell, not a Python limitation to work around. If you encounter a circular import, refactor the modules to break the cycle by extracting shared types into a third module. Never use TYPE_CHECKING hacks to paper over circular dependencies caused by bad design — TYPE_CHECKING is acceptable only for type-only imports that have no runtime impact.
8. Anti-Patterns — Always Refuse These
Never use mutable default arguments. Python evaluates default argument values once at function definition time.
# WRONG — all callers share the same list object
def add_item(item: str, items: list[str] = []) -> list[str]:
items.append(item)
return items
# CORRECT
def add_item(item: str, items: list[str] | None = None) -> list[str]:
if items is None:
items = []
items.append(item)
return items
# OR use dataclasses with field(default_factory=list)
Never use bare except: or except Exception: without re-raising or specific handling. Always catch specific exception types.
Never use global or nonlocal without a comment explaining why encapsulation cannot solve the problem.
Never use eval() or exec() in production code. These introduce security vulnerabilities (code injection) and make code impossible to analyze statically.
Never use assert for input validation in production code. Assertions can be disabled with python -O. Use explicit if/raise for validation:
# WRONG — silently disabled in optimized mode
assert user_id > 0, "user_id must be positive"
# CORRECT — always enforced
if user_id <= 0:
raise ValueError(f"user_id must be positive, got {user_id}")
Never shadow built-in names. The following names are forbidden as variable, parameter, or function names: list, dict, set, tuple, type, id, input, print, min, max, sum, len, range, filter, map, zip, open, hash, format, object, str, int, float, bool, bytes.
Never use type(x) == SomeType for type checking. Use isinstance(x, SomeType) instead — it handles subclasses correctly.