PyArrow Python
Produce schema-explicit columnar code whose logical object, batch boundary,
null behavior, memory lifetime, and materialization point are deliberate.
Boundary
Use this skill when the implementation directly uses pyarrow or must expose an
Arrow-compatible boundary. Do not introduce PyArrow merely to perform a small
Python-list or pandas transformation. Keep storage-format questions in scope
only when PyArrow reads or writes them. Arrow Flight servers, Acero internals,
and non-Python implementations require their own guidance.
Classify the object first
| Object |
Meaning |
Use it for |
DataType / Field / Schema |
Immutable type and named-field contracts, including nullability and metadata. |
Pinning a boundary before constructing or scanning data. |
Scalar |
One typed value, possibly null. |
Kernel arguments and scalar results. |
Array |
One contiguous typed column backed by buffers. |
A single physical column segment. |
ChunkedArray |
One logical column made of zero or more same-typed arrays. |
Table columns and multi-source results; chunk count is not row count. |
RecordBatch |
Equal-length arrays under one schema. |
Bounded transport or processing batches. |
Table |
A logical table whose columns may be chunked. |
Materialized tabular results that fit the required memory boundary. |
RecordBatchReader |
A schema plus a consumable stream of batches. |
Streaming interchange when the consumer can process batches once. |
Dataset |
A logical collection of fragments with a unified schema. |
Discovering and querying multi-file or partitioned data. |
Scanner |
A dataset scan with projection, filter, and batching options bound. |
Deferred dataset execution and bounded batch iteration. |
Arrays and schemas are immutable. A Table is materialized but may be
physically chunked; combine_chunks() can allocate and is not routine cleanup.
A Dataset describes sources, while a Scanner describes the read. Calling
to_table() materializes all selected rows; to_batches() preserves a batch
boundary. Read the object and memory model when
chunking, buffers, ownership, dictionaries, nested types, or zero-copy claims
matter.
Ordered workflow
- Recover the boundary contract: input object, output object, schema, field
nullability, row order if required, batch size, and ownership lifetime.
- Choose
Array, RecordBatch, Table, RecordBatchReader, Dataset, or
Scanner from that contract; do not default every job to Table.
- Pin types where empty input, identifiers, timestamps, decimals, nested data,
dictionary encoding, or cross-language interchange make inference unsafe.
- Push dataset projection and predicates into the scan. Materialize only at a
consumer that requires all rows.
- Use
pyarrow.compute kernels for Arrow data. Do not convert to Python rows
or pandas solely to perform a supported kernel operation.
- Test schema equality separately from value equality, and include empty,
null, multi-chunk, and multi-batch inputs.
- Inspect the installed API before using a drifting option or claiming a
zero-copy, pushdown, or memory property.
Choose by required result
| Required result |
Use |
Critical condition |
| One typed column |
pa.array(..., type=...) |
Conversion and null policy are explicit. |
| One bounded set of equal-length columns |
pa.record_batch(..., schema=...) |
All arrays conform to the same schema and length. |
| One logical materialized table |
pa.table(..., schema=...) |
Total selected result fits the boundary. |
| Consumable batches with a common schema |
RecordBatchReader |
The consumer accepts a one-pass stream. |
| Multi-file discovery and pruning |
pyarrow.dataset.dataset |
Format, filesystem, partitioning, and schema are known or inspected. |
| Deferred filtered/projected scan |
dataset.scanner(...) |
Use dataset expressions such as ds.field, not eager masks. |
| Elementwise, vector, aggregate, or selection work |
pyarrow.compute |
Choose kernel null semantics and checked/safe behavior deliberately. |
| Portable Arrow stream/file serialization |
pyarrow.ipc |
Choose stream for sequential batches, file for random-access footer metadata. |
| Analytical columnar storage |
Parquet APIs or Dataset writer |
Schema, partition layout, row groups, overwrite policy, and reader compatibility are explicit. |
Read the operation map before choosing construction,
compute, dataset, serialization, or conversion APIs.
Canonical typed scan
from collections.abc import Iterator
from pathlib import Path
import pyarrow as pa
import pyarrow.dataset as ds
ORDER_SCHEMA = pa.schema(
[
pa.field("order_id", pa.uint64(), nullable=False),
pa.field("country", pa.string(), nullable=False),
pa.field("amount", pa.decimal128(18, 2), nullable=True),
]
)
def scan_orders(root: Path, country: str) -> Iterator[pa.RecordBatch]:
dataset = ds.dataset(
root,
format="parquet",
schema=ORDER_SCHEMA,
partitioning="hive",
)
scanner = dataset.scanner(
columns=["order_id", "country", "amount"],
filter=ds.field("country") == country,
batch_size=65_536,
)
yield from scanner.to_batches()
This returns bounded batches rather than concatenating the whole dataset. The
explicit schema prevents empty or heterogeneous fragments from silently
changing the boundary. Whether the filter is pruned by partitions/statistics or
evaluated after reading depends on the fragments and expression; verify a
performance claim instead of inferring it from the code.
High-risk rules
Types, nulls, and chunks
- Treat field nullability as a contract. A nullable type does not mean nulls
have been checked for domain validity.
- Do not assume
nullable=False validates constructed values. PyArrow table
construction can attach a non-nullable field to an array that still contains
nulls. Build typed arrays, reject a nonzero null_count for required fields,
then construct the batch or table under the exact schema.
- Use safe casts by default. Permit truncation, overflow, temporal-unit loss, or
invalid UTF-8 only when the contract names that loss and a test proves it.
- Arrow null and floating
NaN are distinct. Choose kernel options explicitly
when either affects filtering, aggregation, or equality.
- Compute functions commonly propagate nulls; Boolean kernels also have Kleene
variants. Select the truth table required by the domain.
- Do not assume one chunk. Test with a multi-chunk
Table, and call
combine_chunks() only for an API that requires contiguous buffers or after
measurement justifies the allocation.
- Dictionary, list, large-list, struct, map, decimal, and timestamp types carry
semantics that Python containers do not preserve automatically. Pin the exact
Arrow type at external boundaries.
Datasets and materialization
- Use dataset
Expression objects for scan filters and projections. A Python
function or eager Boolean array cannot provide file/row-group pruning.
- Include partition columns in the declared schema and specify partitioning
when directory names encode values. Do not infer a partition convention from
one path example.
to_table() loads the selected result; use to_batches() or a
RecordBatchReader when total size is not proven safe.
- Preserve scanner/reader lifetime until consumption completes. Do not return a
view over memory whose producer or foreign owner is already gone.
- “Zero-copy” is conditional on type, layout, alignment, mutability, and the
source/consumer protocol. Verify actual buffers or document the allowed copy;
never promise zero-copy from an API name alone.
Conversion and storage
Conversions to pandas, NumPy, or Python objects can allocate, change nullable
representations, or lose nested/dictionary metadata. State the consumer and
test round trips only for properties the contract requires. For Parquet or IPC,
test the persisted schema and representative values after reopening the
artifact; an in-memory pre-write assertion is insufficient. Run the installed-API inspector, read version and
API grounding, and use testing Arrow
boundaries.
Completion gate
Do not declare completion until the public object type and consumption model
match the caller; schema names, order, types, nullability, metadata requirements,
and time-zone/decimal semantics are asserted; empty, null, multi-chunk, and
multi-batch cases pass; dataset projection/filtering occurs before
materialization where required; conversions and writes are reopened or
round-tripped; and every copy, allocation, ordering, and pushdown claim is
measured or qualified. Report skipped checks and their consequence.
References
- Object and memory model
- Operation map
- Version and API grounding
- Testing Arrow boundaries
1---2name: pyarrow-python3description: Write, review, debug, test, or optimize Python code using PyArrow arrays, schemas, tables, compute kernels, datasets, Parquet, and Arrow IPC.4---56# PyArrow Python78Produce schema-explicit columnar code whose logical object, batch boundary,9null behavior, memory lifetime, and materialization point are deliberate.1011## Boundary1213Use this skill when the implementation directly uses `pyarrow` or must expose an14Arrow-compatible boundary. Do not introduce PyArrow merely to perform a small15Python-list or pandas transformation. Keep storage-format questions in scope16only when PyArrow reads or writes them. Arrow Flight servers, Acero internals,17and non-Python implementations require their own guidance.1819## Classify the object first2021| Object | Meaning | Use it for |22|---|---|---|23| `DataType` / `Field` / `Schema` | Immutable type and named-field contracts, including nullability and metadata. | Pinning a boundary before constructing or scanning data. |24| `Scalar` | One typed value, possibly null. | Kernel arguments and scalar results. |25| `Array` | One contiguous typed column backed by buffers. | A single physical column segment. |26| `ChunkedArray` | One logical column made of zero or more same-typed arrays. | Table columns and multi-source results; chunk count is not row count. |27| `RecordBatch` | Equal-length arrays under one schema. | Bounded transport or processing batches. |28| `Table` | A logical table whose columns may be chunked. | Materialized tabular results that fit the required memory boundary. |29| `RecordBatchReader` | A schema plus a consumable stream of batches. | Streaming interchange when the consumer can process batches once. |30| `Dataset` | A logical collection of fragments with a unified schema. | Discovering and querying multi-file or partitioned data. |31| `Scanner` | A dataset scan with projection, filter, and batching options bound. | Deferred dataset execution and bounded batch iteration. |3233Arrays and schemas are immutable. A `Table` is materialized but may be34physically chunked; `combine_chunks()` can allocate and is not routine cleanup.35A `Dataset` describes sources, while a `Scanner` describes the read. Calling36`to_table()` materializes all selected rows; `to_batches()` preserves a batch37boundary. Read [the object and memory model](references/object-model.md) when38chunking, buffers, ownership, dictionaries, nested types, or zero-copy claims39matter.4041## Ordered workflow42431. Recover the boundary contract: input object, output object, schema, field44 nullability, row order if required, batch size, and ownership lifetime.452. Choose `Array`, `RecordBatch`, `Table`, `RecordBatchReader`, `Dataset`, or46 `Scanner` from that contract; do not default every job to `Table`.473. Pin types where empty input, identifiers, timestamps, decimals, nested data,48 dictionary encoding, or cross-language interchange make inference unsafe.494. Push dataset projection and predicates into the scan. Materialize only at a50 consumer that requires all rows.515. Use `pyarrow.compute` kernels for Arrow data. Do not convert to Python rows52 or pandas solely to perform a supported kernel operation.536. Test schema equality separately from value equality, and include empty,54 null, multi-chunk, and multi-batch inputs.557. Inspect the installed API before using a drifting option or claiming a56 zero-copy, pushdown, or memory property.5758## Choose by required result5960| Required result | Use | Critical condition |61|---|---|---|62| One typed column | `pa.array(..., type=...)` | Conversion and null policy are explicit. |63| One bounded set of equal-length columns | `pa.record_batch(..., schema=...)` | All arrays conform to the same schema and length. |64| One logical materialized table | `pa.table(..., schema=...)` | Total selected result fits the boundary. |65| Consumable batches with a common schema | `RecordBatchReader` | The consumer accepts a one-pass stream. |66| Multi-file discovery and pruning | `pyarrow.dataset.dataset` | Format, filesystem, partitioning, and schema are known or inspected. |67| Deferred filtered/projected scan | `dataset.scanner(...)` | Use dataset expressions such as `ds.field`, not eager masks. |68| Elementwise, vector, aggregate, or selection work | `pyarrow.compute` | Choose kernel null semantics and checked/safe behavior deliberately. |69| Portable Arrow stream/file serialization | `pyarrow.ipc` | Choose stream for sequential batches, file for random-access footer metadata. |70| Analytical columnar storage | Parquet APIs or Dataset writer | Schema, partition layout, row groups, overwrite policy, and reader compatibility are explicit. |7172Read [the operation map](references/operations.md) before choosing construction,73compute, dataset, serialization, or conversion APIs.7475## Canonical typed scan7677```python78from collections.abc import Iterator79from pathlib import Path8081import pyarrow as pa82import pyarrow.dataset as ds838485ORDER_SCHEMA = pa.schema(86 [87 pa.field("order_id", pa.uint64(), nullable=False),88 pa.field("country", pa.string(), nullable=False),89 pa.field("amount", pa.decimal128(18, 2), nullable=True),90 ]91)929394def scan_orders(root: Path, country: str) -> Iterator[pa.RecordBatch]:95 dataset = ds.dataset(96 root,97 format="parquet",98 schema=ORDER_SCHEMA,99 partitioning="hive",100 )101 scanner = dataset.scanner(102 columns=["order_id", "country", "amount"],103 filter=ds.field("country") == country,104 batch_size=65_536,105 )106 yield from scanner.to_batches()107```108109This returns bounded batches rather than concatenating the whole dataset. The110explicit schema prevents empty or heterogeneous fragments from silently111changing the boundary. Whether the filter is pruned by partitions/statistics or112evaluated after reading depends on the fragments and expression; verify a113performance claim instead of inferring it from the code.114115## High-risk rules116117### Types, nulls, and chunks118119- Treat field nullability as a contract. A nullable type does not mean nulls120 have been checked for domain validity.121- Do not assume `nullable=False` validates constructed values. PyArrow table122 construction can attach a non-nullable field to an array that still contains123 nulls. Build typed arrays, reject a nonzero `null_count` for required fields,124 then construct the batch or table under the exact schema.125- Use safe casts by default. Permit truncation, overflow, temporal-unit loss, or126 invalid UTF-8 only when the contract names that loss and a test proves it.127- Arrow null and floating `NaN` are distinct. Choose kernel options explicitly128 when either affects filtering, aggregation, or equality.129- Compute functions commonly propagate nulls; Boolean kernels also have Kleene130 variants. Select the truth table required by the domain.131- Do not assume one chunk. Test with a multi-chunk `Table`, and call132 `combine_chunks()` only for an API that requires contiguous buffers or after133 measurement justifies the allocation.134- Dictionary, list, large-list, struct, map, decimal, and timestamp types carry135 semantics that Python containers do not preserve automatically. Pin the exact136 Arrow type at external boundaries.137138### Datasets and materialization139140- Use dataset `Expression` objects for scan filters and projections. A Python141 function or eager Boolean array cannot provide file/row-group pruning.142- Include partition columns in the declared schema and specify partitioning143 when directory names encode values. Do not infer a partition convention from144 one path example.145- `to_table()` loads the selected result; use `to_batches()` or a146 `RecordBatchReader` when total size is not proven safe.147- Preserve scanner/reader lifetime until consumption completes. Do not return a148 view over memory whose producer or foreign owner is already gone.149- “Zero-copy” is conditional on type, layout, alignment, mutability, and the150 source/consumer protocol. Verify actual buffers or document the allowed copy;151 never promise zero-copy from an API name alone.152153### Conversion and storage154155Conversions to pandas, NumPy, or Python objects can allocate, change nullable156representations, or lose nested/dictionary metadata. State the consumer and157test round trips only for properties the contract requires. For Parquet or IPC,158test the persisted schema and representative values after reopening the159artifact; an in-memory pre-write assertion is insufficient. Run [the installed-API inspector](scripts/inspect_pyarrow.py), read [version and160API grounding](references/version-grounding.md), and use [testing Arrow161boundaries](references/testing.md).162163## Completion gate164165Do not declare completion until the public object type and consumption model166match the caller; schema names, order, types, nullability, metadata requirements,167and time-zone/decimal semantics are asserted; empty, null, multi-chunk, and168multi-batch cases pass; dataset projection/filtering occurs before169materialization where required; conversions and writes are reopened or170round-tripped; and every copy, allocation, ordering, and pushdown claim is171measured or qualified. Report skipped checks and their consequence.172173## References174175- [Object and memory model](references/object-model.md)176- [Operation map](references/operations.md)177- [Version and API grounding](references/version-grounding.md)178- [Testing Arrow boundaries](references/testing.md)