本文へ移動
cccskills
無料GitHub で公開

dask

Use this repo skill for Dask, the Python parallel computing library, when working with lazy task graphs, schedulers, Dask Array, Dask DataFrame, Dask Bag/bytes IO, configuration, diagnostics, CLI usage, or contributor validation.

インストール方法を見る

含まれるファイル(37)

  • SKILL.md5.0 KB
  • references/installation-and-environment.md2.8 KB
  • references/repo-provenance.md2.1 KB
  • references/repo-routing-metadata.json310 B
  • references/troubleshooting.md3.2 KB
  • scripts/dask_package_smoke.py1.1 KB
  • sub-skills/array-workflows/references/api-reference.md6.6 KB
  • sub-skills/array-workflows/references/chunking-and-performance.md5.2 KB
  • sub-skills/array-workflows/references/troubleshooting.md5.2 KB
  • sub-skills/array-workflows/references/workflows.md7.5 KB
  • sub-skills/array-workflows/scripts/array_smoke.py1.1 KB
  • sub-skills/array-workflows/SKILL.md3.3 KB
  • sub-skills/bag-bytes-workflows/references/api-reference.md6.8 KB
  • sub-skills/bag-bytes-workflows/references/io-and-data-formats.md5.9 KB
  • sub-skills/bag-bytes-workflows/references/troubleshooting.md5.6 KB
  • sub-skills/bag-bytes-workflows/references/workflows.md5.4 KB
  • sub-skills/bag-bytes-workflows/scripts/bag_text_smoke.py4.2 KB
  • sub-skills/bag-bytes-workflows/SKILL.md3.1 KB
  • sub-skills/configuration-diagnostics-cli/references/cli-reference.md3.3 KB
  • sub-skills/configuration-diagnostics-cli/references/configuration.md5.3 KB
  • sub-skills/configuration-diagnostics-cli/references/development-and-testing.md3.2 KB
  • sub-skills/configuration-diagnostics-cli/references/diagnostics-and-performance.md4.6 KB
  • sub-skills/configuration-diagnostics-cli/references/troubleshooting.md5.2 KB
  • sub-skills/configuration-diagnostics-cli/scripts/dask_cli_smoke.py2.5 KB
  • sub-skills/configuration-diagnostics-cli/SKILL.md2.7 KB
  • sub-skills/core-graphs-schedulers/references/api-reference.md9.9 KB
  • sub-skills/core-graphs-schedulers/references/troubleshooting.md6.4 KB
  • sub-skills/core-graphs-schedulers/references/workflows.md7.2 KB
  • sub-skills/core-graphs-schedulers/scripts/core_smoke.py3.5 KB
  • sub-skills/core-graphs-schedulers/SKILL.md2.9 KB
  • sub-skills/dataframe-workflows/references/api-reference.md7.5 KB
  • sub-skills/dataframe-workflows/references/io-and-data-formats.md5.7 KB
  • sub-skills/dataframe-workflows/references/troubleshooting.md6.3 KB
  • sub-skills/dataframe-workflows/references/workflows.md6.1 KB
  • sub-skills/dataframe-workflows/scripts/dataframe_demo_smoke.py4.2 KB
  • sub-skills/dataframe-workflows/scripts/dataframe_smoke.py5.0 KB
  • sub-skills/dataframe-workflows/SKILL.md4.0 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Dask Repo Skill

Use this skill when a task involves Dask's public APIs, internal task-graph model, collection workflows, configuration, diagnostics, or contributor test/build practices. Dask provides lazy parallel collections and schedulers for Python analytics: delayed objects, arrays, dataframes, bags, low-level task graphs, and local/distributed execution integrations.

First Checks

  • Import check: python -c "import dask; print(dask.__version__)".
  • CLI check: dask --help should show config, docs, and info command groups.
  • Minimal compute check: python scripts/dask_package_smoke.py --scheduler synchronous.
  • Install extras: use dask[array] for array workflows, dask[dataframe] for dataframe workflows, dask[diagnostics] for local dashboard/profiling helpers, and dask[distributed] only when remote cluster APIs are needed.

Route By Task

  • Core graphs and schedulers: Use sub-skills/core-graphs-schedulers/SKILL.md for dask.delayed, compute, persist, optimize, annotations, tokenization, HighLevelGraph, low-level task specs, scheduler choice, and graph debugging.
  • Array workflows: Use sub-skills/array-workflows/SKILL.md for dask.array, chunk planning, from_array, slicing, map_blocks, blockwise, reductions, overlap, rechunking, gufuncs, random arrays, linalg, FFT, stats, and array backend caveats.
  • DataFrame workflows: Use sub-skills/dataframe-workflows/SKILL.md for dask.dataframe, pandas-like APIs, CSV/Parquet/JSON/SQL IO, partitions/divisions, joins, groupby, shuffle, repartitioning, pyarrow strings, categoricals, and dask_expr query planning.
  • Bag and bytes workflows: Use sub-skills/bag-bytes-workflows/SKILL.md for dask.bag, dask.bytes, text and byte IO, JSON-like records, Avro, fsspec URLs, compression, foldby, and small-file object pipelines.
  • Configuration, diagnostics, and CLI: Use sub-skills/configuration-diagnostics-cli/SKILL.md for dask.config, YAML config paths, environment variables, dask config, dask info, local profilers, progress bars, callbacks/cache, install extras, and contributor validation commands.

Shared References

  • Read references/installation-and-environment.md when choosing extras, verifying imports, or diagnosing optional dependency availability.
  • Read references/troubleshooting.md for cross-cutting failures that affect multiple Dask collections or schedulers.
  • Read references/repo-provenance.md before relying on this skill for a changed checkout; refresh the skill when commit, package version, or major evidence paths drift.

Dask Working Rules

  • Keep graph construction lazy. Do not call .compute() or .persist() inside collection methods or while defining a reusable graph unless the user explicitly wants materialization.
  • Use metadata, divisions, chunks, dtypes, and meta objects to infer output shape/schema instead of computing sample data.
  • Choose the smallest collection that matches the data model: delayed for arbitrary Python functions, array for NumPy-like blocked tensors, dataframe for pandas-like tabular data, bag for unordered Python records, and bytes for low-level file blocks.
  • Prefer explicit scheduler and config scopes when reproducing behavior: with dask.config.set({...}): ... and compute(..., scheduler="synchronous") for deterministic local debugging.
  • Treat array.query-planning and dataframe.query-planning as import-time configuration; set them before importing the relevant collection modules in a fresh process.
  • For dataframe and array APIs, validate meta, chunks, divisions, and optional dependency requirements before changing algorithms.

Validation Pattern

  1. Run the root smoke script for package-level sanity.
  2. Run the owning sub-skill smoke script for the collection or support workflow.
  3. Add a tiny fixture or in-memory example that exercises the changed route.
  4. For repo contribution work, run the narrowest relevant pytest selection before broader suites.
  5. Record skipped native cases separately when they require network, GPU, distributed services, large data, or credentials.

Common Pitfalls

  • Importing dask.dataframe without the dataframe extra can fail due to missing pandas or pyarrow.
  • Unknown chunks or divisions often make otherwise valid operations expensive, ambiguous, or unsupported.
  • Multiprocessing schedulers require pickleable functions and guarded entry points; use threads or synchronous mode for quick debugging.
  • Config files may be read from multiple paths; dask config find <key> is the safest way to diagnose where a value comes from.
  • Optional GPU/cloud/storage backends are not part of the base package; keep those checks conditional and skip unsafe native cases unless dependencies and hardware are available.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Routes 3D ResNets PyTorch video action-recognition workflows across training, inference, and data preparation.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

3ddfa

無料

Guide 3DDFA Python inference, geometry rendering, training/evaluation, and optional C++ ONNX workflows for 3D dense face alignment.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

3ddfa-v2

無料

Routes 3DDFA_V2 face-alignment setup, still-image demos, video tracking, and ONNX benchmarking workflows.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

ab3dmot

無料

Operate AB3DMOT 3D multi-object tracking workflows for KITTI and nuScenes data, tracking, evaluation, and visualization.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

Use Hugging Face Accelerate for PyTorch training-loop migration, distributed launch/configuration, DeepSpeed/FSDP/TPU backend setup, big-model inference/offload, checkpointing, tracking, and troubleshooting.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

acme

無料

Route Acme reinforcement-learning framework tasks across core loops, replay/data, JAX agents, and TensorFlow/Sonnet agents.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

VectorSpaceLab のスキルをすべて見る

このスキルの問題を報告する