本文へ移動
cccskills
無料GitHub で公開

pandera-polars

Use for writing, reviewing, debugging, or testing Pandera schemas and runtime validation for Polars DataFrame or LazyFrame pipelines installed with pandera[polars]. Trigger on pandera.polars DataFrameSchema, DataFrameModel, Column, Field, Check, PolarsData, decorators, coercion, strictness, lazy error collection, and validation-depth decisions. Do not use for pandas-backed Pandera, Pydantic object models, Polars transformations without Pandera, static dataframe typing alone, or generic data-quality platforms.

インストール方法を見る

含まれるファイル(6)

  • SKILL.md8.7 KB
  • references/object-model.md1.6 KB
  • references/operations.md1.4 KB
  • references/testing.md978 B
  • references/version-grounding.md809 B
  • scripts/inspect_pandera_polars.py1.6 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Pandera for Polars

Create executable Polars dataframe contracts whose backend, validation depth, coercion, failure aggregation, and pipeline boundary are explicit.

Boundary

Use this skill only for Pandera's Polars backend. Import it as pandera.polars; pandas, Ibis, PySpark, and Narwhals-backed behavior differs. Use the Polars skill for transformation semantics and this skill for runtime dataframe contracts. Do not replace ordinary Python object validation with a one-row dataframe schema.

Know the objects and overloaded words

ObjectRuntime meaningUse it for
DataFrameSchemaAn executable schema object containing Polars column and dataframe checks.Dynamic/programmatic schemas and schema composition.
ColumnA named column contract: dtype, nullability, requirement, uniqueness, coercion, and checks.Per-column structural and value rules.
CheckA predicate contract evaluated by the backend.Domain constraints not captured by dtype/nullability.
DataFrameModelA class-declared schema compiled from annotations, Fields, checks, and config.Reusable named contracts with type-checker-friendly declarations.
FieldDeclarative column constraints inside a DataFrameModel.Built-in comparisons, membership, aliases, nullable/unique behavior.
PolarsDataCustom-check input holding a LazyFrame and optional column key.Native vectorized Polars checks.
SchemaError / SchemaErrorsOne validation failure or an aggregate of failures.Machine-readable failure handling and diagnostics.

Two kinds of “lazy” must remain separate:

  • pl.LazyFrame is a deferred Polars query. Pandera's native Polars validation checks schema-level properties by default and does not automatically execute all data-level checks on an uncollected plan.
  • schema.validate(..., lazy=True) requests accumulation of multiple validation failures before raising; it does not make eager validation computationally lazy.

Read the schema and validation model before choosing a schema style or claiming that values were checked.

Ordered workflow

  1. Recover the data contract: accepted frame type, ordered/required columns, exact dtypes, nullable fields, uniqueness, extra-column policy, allowed coercions, cross-column invariants, and failure interface.
  2. Import pandera.polars as pa. Confirm the installed Pandera and Polars versions before copying a backend feature or signature.
  3. Choose DataFrameSchema for dynamic composition or DataFrameModel for a stable named contract. Do not maintain both as independent sources of truth.
  4. Encode structural rules first, then built-in vectorized checks, then the smallest native custom check that remains.
  5. Choose validation depth from the input object. If data-level checks are required for a LazyFrame, collect or otherwise establish a supported execution boundary; do not report schema-only validation as full validation.
  6. Choose coerce, strict, and lazy independently. Each changes a different contract.
  7. Validate at ingress, after an untrusted/shape-changing stage, or before egress—not after every expression by habit.
  8. Test one failure for every rule and inspect failure_cases/error categories, not exception prose alone.

Decision table

NeedUseDo not confuse it with
Programmatic/reusable schema valuepa.DataFrameSchemaA Python class instance model.
Declarative named dataframe contractpa.DataFrameModelPydantic BaseModel.
Reject unspecified columnsstrict=Truerequired=True, which concerns declared columns.
Drop unspecified columns intentionallyInstalled strict="filter" supportSilent schema drift; test passthrough loss.
Convert compatible inputscoerce=True at the chosen scopeValidation-only behavior; coercion mutates the returned representation.
Gather all failuresvalidate(..., lazy=True)A Polars LazyFrame.
Vectorized custom column ruleCheck receiving PolarsData and returning Boolean LazyFrame outputPython element callbacks.
Cross-column invariantDataframe-level native checkA column check that cannot see the needed peer fields.

Read the operation map for schemas, models, checks, decorators, and pipeline placement.

Canonical strict contract

import pandera.polars as pa
import polars as pl


ORDERS = pa.DataFrameSchema(
    {
        "order_id": pa.Column(
            pl.Int64,
            checks=pa.Check.ge(1),
            nullable=False,
            unique=True,
        ),
        "country": pa.Column(
            pl.String,
            checks=pa.Check.isin(["DE", "FR", "NL"]),
            nullable=False,
        ),
        "amount": pa.Column(
            pl.Float64,
            checks=pa.Check.ge(0),
            nullable=True,
        ),
    },
    strict=True,
    coerce=False,
)


def validate_orders(frame: pl.DataFrame) -> pl.DataFrame:
    return ORDERS.validate(frame, lazy=True)

This rejects extra columns, does not silently coerce identifiers or amounts, checks all data-level rules on an eager frame, and aggregates failures. If the boundary intentionally normalizes compatible types, turn on coercion and test the returned dtypes and failed dirty values; do not use coercion to make an unknown schema “pass.”

Custom checks without Python row paths

import pandera.polars as pa
import polars as pl
from pandera.polars import PolarsData


def start_before_end(data: PolarsData) -> pl.LazyFrame:
    return data.lazyframe.select(pl.col("start") <= pl.col("end"))


INTERVALS = pa.DataFrameSchema(
    {
        "start": pa.Column(pl.Datetime),
        "end": pa.Column(pl.Datetime),
    },
    checks=pa.Check(start_before_end),
)

A native custom check returns a Boolean LazyFrame shape accepted by the backend. Avoid element_wise=True for a vectorizable condition; the Polars backend implements elementwise Python callbacks through a slower row path.

High-risk rules

  • Backend imports are part of correctness. Do not use top-level or pandera.pandas examples in Polars code.
  • nullable, required, unique, strict, and coerce are independent. State each required policy rather than relying on defaults.
  • A schema validates its declared contract, not business completeness. Add cross-column checks for relational invariants and tests for duplicate/null combinations where needed.
  • DataFrame validation can return a coerced/filtered frame. Use the returned value; do not validate and then continue with the unvalidated original.
  • For a LazyFrame, establish whether schema-only validation is acceptable. If row values must be proven, validate an eager boundary or a currently documented supported full-data path.
  • Treat validation errors as structured evidence. Preserve useful failure cases without leaking sensitive source values in logs or API responses.
  • Feature support differs by backend. Do not copy pandas-only index, groupby check, parser, hypothesis, synthesis, or custom-registration patterns into a Polars schema without current evidence.

Run the backend inspector, read version and backend grounding, and use the validation test matrix.

Completion gate

Do not declare completion until the import selects the Polars backend; the schema style has one source of truth; exact dtypes, required/null/unique/extra and coercion policies are asserted; eager versus lazy-frame validation depth is tested; vectorizable rules stay native; the returned validated frame is used; every rule has a failing example; aggregated failures are asserted structurally; and installed-version checks cover backend-specific syntax. Report any schema-only LazyFrame validation or skipped dependency execution explicitly.

References

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Use this skill whenever the user wants to create or improve a presentation for an academic context — conference papers, seminar talks, thesis defenses, grant briefings, lab meetings, invited lectures, or any presentation where the audience will evaluate reasoning and evidence. Triggers include: 'conference talk', 'seminar slides', 'thesis defense', 'research presentation', 'academic deck', 'academic presentation'. Also triggers when the user asks to 'make slides' in combination with academic content (e.g., 'make slides for my paper on X', 'create a presentation for my dissertation defense', 'build a deck for my grant proposal'). This skill governs CONTENT and STRUCTURE decisions. For the technical work of creating or editing the .pptx file itself, also read the pptx SKILL.md.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd1202026年10月9日 更新

adp

無料

Redpanda's Agentic Data Plane: governance infrastructure for building, running, and governing AI agents and MCP servers, plus a proxying AI Gateway for LLM providers, operated via `rpk ai` and the ADP API. Use when creating or managing AI agents (managed or self-managed) via `rpk ai agent` or `AgentRegistryService`; configuring MCP servers (remote or managed catalog, code mode, auth); setting up LLM providers or querying models via `rpk ai llm`/`rpk ai model` or the AI Gateway proxy; or configuring budgets, guardrails, or Cedar access-control policies through the governance APIs. Also covers reading agent transcripts and spending insights, and wiring OAuth clients or providers to the aigw Authorization Server. For the separate rpk cloud mcp control-plane MCP server, see `/redpanda:rpk-cloud`.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd1202026年10月9日 更新

ads

無料

Operate professional paid advertising across Google, Meta, YouTube, LinkedIn, TikTok, Microsoft, Apple, Amazon, Reddit, Pinterest, Snapchat, and X. Use for account intake, source-grounded audits, strategy, budget and measurement planning, creative production, experiments, reporting, monitoring, and explicitly approved campaign changes. Also trigger on PPC, paid social, retail media, attribution, tracking, landing pages, cross-platform conversion totals, negative keywords or search terms, beta-feature scoring, stale platform claims, API-token or credential setup, campaign deletion, and safe Claude Ads installation or uninstall.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd1202026年10月9日 更新

ads-apple

無料

Audit Apple Ads measurement, AdServices and AdAttributionKit, campaign and keyword structure, Search Match, App Store placements, custom product pages, bidding, budgets, MMP reconciliation, and policy. Use for Apple Ads, Apple Search Ads, App Store ads, Search Match, custom product pages, AdServices, or Apple app-install campaigns.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd1202026年10月9日 更新

Research competitor paid-ad presence, messaging, creative, formats, landing pages, keyword and auction signals, transparent ad libraries, and strategic gaps across supported platforms. Use for competitor ads, ad libraries, ad spy, competitive PPC analysis, competitor creative, Google Ads Transparency, Meta Ad Library, or paid-media competitor research.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd1202026年10月9日 更新

ads-dna

無料

Extract a public-safe brand and offer profile for paid advertising from an authorized website and operator input. Triggers on: brand DNA, brand profile, brand identity, brand style, brand colors, brand voice, visual identity, style guide, website brand analysis.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd1202026年10月9日 更新

skillmds のスキルをすべて見る

このスキルの問題を報告する