Polars Python
Produce version-grounded Polars code whose object types, row grain, schema,
missing-value policy, ordering, and execution boundary are deliberate and
tested.
Boundary
Use this skill when the project already uses Polars or the user explicitly
requests Polars or a pandas-to-Polars migration. Polars SQL is in scope only
inside a Polars pipeline. Do not introduce Polars into a library-neutral,
pandas-only, PySpark, Rust, standalone-SQL, Cloud, On-Prem, distributed, or GPU
task. Preserve the caller's public return type unless the task changes it.
Know the objects before choosing an API
| Object |
Runtime meaning |
Use it for |
DataFrame |
An eager, materialized two-dimensional table of uniquely named, equal-length typed columns. |
In-memory work, an eager interface, or a collected result whose final table fits memory. |
Series |
One eager, materialized, named one-dimensional column with one dtype. |
A one-column input/output boundary. |
Expr |
An unevaluated recipe for producing values from columns or literals; it is not data. |
Reusable computation inside an expression context. |
LazyFrame |
A logical query plan that produces a table only when executed. |
File-backed, multi-stage, large, or optimization-sensitive work. |
Schema |
The ordered mapping from column names to dtypes. |
The table interface contract. |
| Selector |
A schema-aware expression expander such as cs.numeric(). |
Apply one rule to columns selected by name or dtype. |
DataFrame columns are Series. Use frame methods to change table shape and
Expr objects to define column computations or predicates inside those methods.
Eager and lazy frames share most expression APIs: eager contexts execute now,
while lazy contexts extend a plan. GroupBy and LazyGroupBy are builders, not
result frames; finish them with a terminal group operation, normally .agg(...).
Polars has no pandas-style semantic row index—keep record labels as columns.
Read the object and shape model whenever object
type, dimensionality, positional selection, dtype namespaces, or return shape
is uncertain.
Ordered workflow
- Recover the contract from the request, callers, schema, and tests: public
return object, row grain, keys, dtypes, missing values, and required order.
- Classify the operation's shape effect, then choose its context or API family.
- Choose eager, lazy, streaming, or sink execution from the real input and
consumer boundaries.
- Apply the relevant schema, join, time-series, reshape, migration, or UDF rule.
- Keep native expressions in the pipeline; in lazy work, materialize only at a
required consumer or unsupported-operation boundary.
- Add the smallest falsifier: typed empty input, null/
NaN, duplicate keys,
order permutation, ambiguous rows, or dirty input.
- Run targeted tests. Inspect a plan or measure scale only for a performance
claim.
Rigorous analytical contract
Separate four stages and report which ones actually ran:
| Stage |
Required output |
| Frame |
Population, row grain, keys, metric formulas, units, missingness, and required order. |
| Transform |
Input/output objects, shape change, schema, execution boundary, and failure policy. |
| Validate |
Cardinality, reconciliation, invariants, adversarial fixtures, and execution evidence. |
| Interpret |
Evidence-supported result plus unresolved semantic, data, or numerical risk. |
Do not infer semantics from column names alone. Distinguish row count, entity
count, non-null count, numerator, denominator, and units before aggregation.
Compute a rate from aggregated numerator and denominator unless the contract
specifically defines a justified weighted alternative; never average row or
subgroup rates merely because they are available.
For consequential transformations, use at least one independent falsifier:
input-order permutation, batch partition/recombination for row-local work,
component-to-total reconciliation, eager/lazy equivalence, or a second query
formulation. Passing execution is not validation. Read rigorous Polars
practice for the full checklist and
evaluated anchors.
Choose by intent and output shape
| Required result |
Use |
Shape contract |
| Only selected/computed columns |
select |
Only requested outputs; expression lengths must be compatible. |
| Original columns plus additions/replacements |
with_columns |
Input height and unspecified columns remain. |
| Rows satisfying a predicate |
filter |
Same columns; only predicate True survives. |
| One materialized column |
get_column |
Returns Series; select("x") instead returns a one-column frame. |
| One row per group |
group_by(...).agg(...) |
Output grain is distinct group keys. |
| Group result aligned to every row |
expression .over(...) |
Default group_to_rows mapping preserves input grain; explode is shape-changing. |
| Rows chosen by position |
slice, head, tail, or gather |
Positions, not labels, define the result. |
| Records stacked by schema |
pl.concat |
Schema compatibility or union policy is explicit. |
| Tables combined by keys |
join |
Cardinality follows key multiplicity and join type. |
| Nested elements/fields expanded |
explode / unnest |
Row or column grain changes and is tested. |
| Columns converted to/from rows |
unpivot / pivot |
Identifier grain and generated-column policy are explicit. |
Read the DataFrame operation map for
construction, inspection, selectors, ordinary transformations, sorting,
deduplication, reshaping, and export.
Canonical object flow
import polars as pl
def normalized_email(column: str) -> pl.Expr:
return pl.col(column).str.strip_chars().str.to_lowercase()
def enrich(frame: pl.DataFrame) -> pl.DataFrame:
return (
frame.with_columns(
normalized_email("email").alias("email"),
net=pl.col("quantity") * pl.col("unit_price"),
)
.with_columns(tax=pl.col("net") * pl.lit(0.19))
.filter(pl.col("net").is_not_null())
.select("order_id", "email", "net", "tax")
)
The helper returns a symbolic Expr, so the caller chooses eager or lazy
execution. The second with_columns is required because sibling expressions
see the context's input schema, not aliases created by siblings. pl.lit
unambiguously represents a value; strings in expression-input positions often
mean column names.
Execution boundary
- Return
LazyFrame when the public contract requires a plan; never collect
inside that function or collect and re-lazify to imitate laziness.
- For file-backed or multi-stage work whose final table fits memory, start with
scan_*, keep one plan, and collect() once at the eager consumer.
- For a file or batch consumer, use an installed
sink_* or batch API when it
avoids materializing an oversized final DataFrame.
- For a small existing
DataFrame with an eager consumer, eager expressions
are valid; do not add .lazy().collect() ceremonially.
- If a required operation has no usable lazy form, materialize immediately
before it only when the public contract permits an eager result. Record the
optimization barrier.
- A streaming collect still materializes its final
DataFrame; it only can
bound intermediate memory. A LazyFrame is a plan, not a cached result.
Read lazy execution and performance before making
an engine, streaming, caching, plan, or memory claim.
High-risk semantic rules
Expressions and schema
- Use parenthesized predicates with
&, |, and ~, never Python boolean
operators on expressions. Every conditional branch must be independently
valid.
- Alias derived scalar outputs. Selector and dtype expressions can expand to
zero, one, or many columns; test expected names.
- Choose
.str, .dt, .list, .arr, .struct, .cat, or .bin from the
dtype rather than converting values to Python.
pl.len() counts rows; expression .count() counts non-null values; a bare
grouped column produces a List.
- Pin constructor or scan schemas when identifiers, temporal/nested values,
empty inputs, or late dirty rows make inference unsafe. For ambiguous
row-oriented constructor data, pass
orient="row".
- Keep casts strict unless invalid values becoming null is the declared policy,
then test or count introduced nulls. Null and floating
NaN are distinct.
Read expression and shape rules and schema and
missing-data rules for these branches.
Relational, ordered, and shape-changing work
Before a join, state preserved rows, key domains/dtypes, uniqueness on both
sides, null-key behavior, output-key behavior, and order. Validate known
cardinality; otherwise prove uniqueness on the constrained side. Never cast
keys unless both represent the same domain and conversion is lossless. Read
join rules.
After an important join, measure result rows, null fact keys, unmatched
non-null fact rows, unused dimension keys, and any conserved measure required
by the analysis. A cardinality declaration proves multiplicity, not semantic
coverage or reconciliation.
Do not rely on undocumented output order from group, join, unique, pivot, or
unpivot. Preserve input order only where the chosen API guarantees it; otherwise
encode the sort keys and tie-breakers that define a survivor, list, cumulative
result, or final table. Read join rules for as-of joins and
time-series rules for rolling, dynamic grouping,
calendar durations, and time zones.
Python escape hatch
Prefer native expressions, selectors, typed namespaces, folds, windows, joins,
and reshapes. If they cannot express the operation, choose the narrowest
installed batch UDF before scalar map_elements; use a whole-frame UDF only
under an explicit schema contract. Declare dtype/schema, purity, null/exception
behavior, and empty/all-null behavior. Claim streamability only if arbitrary
batch boundaries cannot change results. Read Python UDF
contracts before using any callback.
Version grounding and completion
Inspect the installed version when a signature, keyword, dtype, engine,
warning, or capability can drift. Run python scripts/inspect_polars.py from
the installed skill directory for JSON evidence and read API
grounding. Do not copy stale syntax from a prompt.
Translate pandas semantics rather than spellings; read pandas
migration before a migration. Use Polars
testing helpers and typed expected frames; read testing Polars
behavior.
Do not declare completion until the return object, grain, schema, cardinality,
missing-value policy, and required order match the contract; the relevant
dirty/empty/null/duplicate/permutation fixtures pass; no accidental
materialization, unjustified Python row path, arbitrary key cast, or stale API
remains; aggregate components reconcile where required; validation status says
what executed versus what was only generated; and project checks pass or
skipped evidence and its consequence are reported.
References
- Core solution recipes
- Lazy and temporal recipes
- Rigorous analytical practice and recipes
- Object and shape model
- DataFrame operation map
- Expression and shape rules
- Schema and missing-data rules
- Join rules
- Time-series rules
- Lazy execution and performance
- Python UDF contracts
- Pandas migration
- Testing Polars behavior
- API grounding
1---2name: polars-python3description: Write, review, debug, test, and optimize Python Polars code with version-grounded object types, schemas, and execution boundaries.4---56# Polars Python78Produce version-grounded Polars code whose object types, row grain, schema,9missing-value policy, ordering, and execution boundary are deliberate and10tested.1112## Boundary1314Use this skill when the project already uses Polars or the user explicitly15requests Polars or a pandas-to-Polars migration. Polars SQL is in scope only16inside a Polars pipeline. Do not introduce Polars into a library-neutral,17pandas-only, PySpark, Rust, standalone-SQL, Cloud, On-Prem, distributed, or GPU18task. Preserve the caller's public return type unless the task changes it.1920## Know the objects before choosing an API2122| Object | Runtime meaning | Use it for |23|---|---|---|24| `DataFrame` | An eager, materialized two-dimensional table of uniquely named, equal-length typed columns. | In-memory work, an eager interface, or a collected result whose final table fits memory. |25| `Series` | One eager, materialized, named one-dimensional column with one dtype. | A one-column input/output boundary. |26| `Expr` | An unevaluated recipe for producing values from columns or literals; it is not data. | Reusable computation inside an expression context. |27| `LazyFrame` | A logical query plan that produces a table only when executed. | File-backed, multi-stage, large, or optimization-sensitive work. |28| `Schema` | The ordered mapping from column names to dtypes. | The table interface contract. |29| Selector | A schema-aware expression expander such as `cs.numeric()`. | Apply one rule to columns selected by name or dtype. |3031`DataFrame` columns are `Series`. Use frame methods to change table shape and32`Expr` objects to define column computations or predicates inside those methods.33Eager and lazy frames share most expression APIs: eager contexts execute now,34while lazy contexts extend a plan. `GroupBy` and `LazyGroupBy` are builders, not35result frames; finish them with a terminal group operation, normally `.agg(...)`.36Polars has no pandas-style semantic row index—keep record labels as columns.3738Read [the object and shape model](references/object-model.md) whenever object39type, dimensionality, positional selection, dtype namespaces, or return shape40is uncertain.4142## Ordered workflow43441. Recover the contract from the request, callers, schema, and tests: public45 return object, row grain, keys, dtypes, missing values, and required order.462. Classify the operation's shape effect, then choose its context or API family.473. Choose eager, lazy, streaming, or sink execution from the real input and48 consumer boundaries.494. Apply the relevant schema, join, time-series, reshape, migration, or UDF rule.505. Keep native expressions in the pipeline; in lazy work, materialize only at a51 required consumer or unsupported-operation boundary.526. Add the smallest falsifier: typed empty input, null/`NaN`, duplicate keys,53 order permutation, ambiguous rows, or dirty input.547. Run targeted tests. Inspect a plan or measure scale only for a performance55 claim.5657## Rigorous analytical contract5859Separate four stages and report which ones actually ran:6061| Stage | Required output |62|---|---|63| Frame | Population, row grain, keys, metric formulas, units, missingness, and required order. |64| Transform | Input/output objects, shape change, schema, execution boundary, and failure policy. |65| Validate | Cardinality, reconciliation, invariants, adversarial fixtures, and execution evidence. |66| Interpret | Evidence-supported result plus unresolved semantic, data, or numerical risk. |6768Do not infer semantics from column names alone. Distinguish row count, entity69count, non-null count, numerator, denominator, and units before aggregation.70Compute a rate from aggregated numerator and denominator unless the contract71specifically defines a justified weighted alternative; never average row or72subgroup rates merely because they are available.7374For consequential transformations, use at least one independent falsifier:75input-order permutation, batch partition/recombination for row-local work,76component-to-total reconciliation, eager/lazy equivalence, or a second query77formulation. Passing execution is not validation. Read [rigorous Polars78practice](references/recipes-rigorous-analysis.md) for the full checklist and79evaluated anchors.8081## Choose by intent and output shape8283| Required result | Use | Shape contract |84|---|---|---|85| Only selected/computed columns | `select` | Only requested outputs; expression lengths must be compatible. |86| Original columns plus additions/replacements | `with_columns` | Input height and unspecified columns remain. |87| Rows satisfying a predicate | `filter` | Same columns; only predicate `True` survives. |88| One materialized column | `get_column` | Returns `Series`; `select("x")` instead returns a one-column frame. |89| One row per group | `group_by(...).agg(...)` | Output grain is distinct group keys. |90| Group result aligned to every row | expression `.over(...)` | Default `group_to_rows` mapping preserves input grain; `explode` is shape-changing. |91| Rows chosen by position | `slice`, `head`, `tail`, or `gather` | Positions, not labels, define the result. |92| Records stacked by schema | `pl.concat` | Schema compatibility or union policy is explicit. |93| Tables combined by keys | `join` | Cardinality follows key multiplicity and join type. |94| Nested elements/fields expanded | `explode` / `unnest` | Row or column grain changes and is tested. |95| Columns converted to/from rows | `unpivot` / `pivot` | Identifier grain and generated-column policy are explicit. |9697Read [the DataFrame operation map](references/dataframe-operations.md) for98construction, inspection, selectors, ordinary transformations, sorting,99deduplication, reshaping, and export.100101## Canonical object flow102103```python104import polars as pl105106107def normalized_email(column: str) -> pl.Expr:108 return pl.col(column).str.strip_chars().str.to_lowercase()109110111def enrich(frame: pl.DataFrame) -> pl.DataFrame:112 return (113 frame.with_columns(114 normalized_email("email").alias("email"),115 net=pl.col("quantity") * pl.col("unit_price"),116 )117 .with_columns(tax=pl.col("net") * pl.lit(0.19))118 .filter(pl.col("net").is_not_null())119 .select("order_id", "email", "net", "tax")120 )121```122123The helper returns a symbolic `Expr`, so the caller chooses eager or lazy124execution. The second `with_columns` is required because sibling expressions125see the context's input schema, not aliases created by siblings. `pl.lit`126unambiguously represents a value; strings in expression-input positions often127mean column names.128129## Execution boundary130131- Return `LazyFrame` when the public contract requires a plan; never collect132 inside that function or collect and re-lazify to imitate laziness.133- For file-backed or multi-stage work whose final table fits memory, start with134 `scan_*`, keep one plan, and `collect()` once at the eager consumer.135- For a file or batch consumer, use an installed `sink_*` or batch API when it136 avoids materializing an oversized final `DataFrame`.137- For a small existing `DataFrame` with an eager consumer, eager expressions138 are valid; do not add `.lazy().collect()` ceremonially.139- If a required operation has no usable lazy form, materialize immediately140 before it only when the public contract permits an eager result. Record the141 optimization barrier.142- A streaming collect still materializes its final `DataFrame`; it only can143 bound intermediate memory. A `LazyFrame` is a plan, not a cached result.144145Read [lazy execution and performance](references/performance.md) before making146an engine, streaming, caching, plan, or memory claim.147148## High-risk semantic rules149150### Expressions and schema151152- Use parenthesized predicates with `&`, `|`, and `~`, never Python boolean153 operators on expressions. Every conditional branch must be independently154 valid.155- Alias derived scalar outputs. Selector and dtype expressions can expand to156 zero, one, or many columns; test expected names.157- Choose `.str`, `.dt`, `.list`, `.arr`, `.struct`, `.cat`, or `.bin` from the158 dtype rather than converting values to Python.159- `pl.len()` counts rows; expression `.count()` counts non-null values; a bare160 grouped column produces a `List`.161- Pin constructor or scan schemas when identifiers, temporal/nested values,162 empty inputs, or late dirty rows make inference unsafe. For ambiguous163 row-oriented constructor data, pass `orient="row"`.164- Keep casts strict unless invalid values becoming null is the declared policy,165 then test or count introduced nulls. Null and floating `NaN` are distinct.166167Read [expression and shape rules](references/expressions.md) and [schema and168missing-data rules](references/schema-missing.md) for these branches.169170### Relational, ordered, and shape-changing work171172Before a join, state preserved rows, key domains/dtypes, uniqueness on both173sides, null-key behavior, output-key behavior, and order. Validate known174cardinality; otherwise prove uniqueness on the constrained side. Never cast175keys unless both represent the same domain and conversion is lossless. Read176[join rules](references/joins.md).177178After an important join, measure result rows, null fact keys, unmatched179non-null fact rows, unused dimension keys, and any conserved measure required180by the analysis. A cardinality declaration proves multiplicity, not semantic181coverage or reconciliation.182183Do not rely on undocumented output order from group, join, unique, pivot, or184unpivot. Preserve input order only where the chosen API guarantees it; otherwise185encode the sort keys and tie-breakers that define a survivor, list, cumulative186result, or final table. Read [join rules](references/joins.md) for as-of joins and187[time-series rules](references/time-series.md) for rolling, dynamic grouping,188calendar durations, and time zones.189190### Python escape hatch191192Prefer native expressions, selectors, typed namespaces, folds, windows, joins,193and reshapes. If they cannot express the operation, choose the narrowest194installed batch UDF before scalar `map_elements`; use a whole-frame UDF only195under an explicit schema contract. Declare dtype/schema, purity, null/exception196behavior, and empty/all-null behavior. Claim streamability only if arbitrary197batch boundaries cannot change results. Read [Python UDF198contracts](references/udf-contracts.md) before using any callback.199200## Version grounding and completion201202Inspect the installed version when a signature, keyword, dtype, engine,203warning, or capability can drift. Run `python scripts/inspect_polars.py` from204the installed skill directory for JSON evidence and read [API205grounding](references/api-grounding.md). Do not copy stale syntax from a prompt.206207Translate pandas semantics rather than spellings; read [pandas208migration](references/pandas-migration.md) before a migration. Use Polars209testing helpers and typed expected frames; read [testing Polars210behavior](references/testing.md).211212Do not declare completion until the return object, grain, schema, cardinality,213missing-value policy, and required order match the contract; the relevant214dirty/empty/null/duplicate/permutation fixtures pass; no accidental215materialization, unjustified Python row path, arbitrary key cast, or stale API216remains; aggregate components reconcile where required; validation status says217what executed versus what was only generated; and project checks pass or218skipped evidence and its consequence are reported.219220## References221222- [Core solution recipes](references/recipes-core.md)223- [Lazy and temporal recipes](references/recipes-scale.md)224- [Rigorous analytical practice and recipes](references/recipes-rigorous-analysis.md)225- [Object and shape model](references/object-model.md)226- [DataFrame operation map](references/dataframe-operations.md)227- [Expression and shape rules](references/expressions.md)228- [Schema and missing-data rules](references/schema-missing.md)229- [Join rules](references/joins.md)230- [Time-series rules](references/time-series.md)231- [Lazy execution and performance](references/performance.md)232- [Python UDF contracts](references/udf-contracts.md)233- [Pandas migration](references/pandas-migration.md)234- [Testing Polars behavior](references/testing.md)235- [API grounding](references/api-grounding.md)