本文へ移動
cccskills
無料GitHub で公開

datachain-knowledge

Use whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure Blob, or local storage. Also use when running any script that may create datasets as a side effect. Maintains a knowledge base at dc-knowledge/ (JSON + markdown). ALWAYS use this skill when the user creates a dataset, saves pipeline output, runs a data script, or references any storage bucket.

インストール方法を見る

含まれるファイル(21)

  • SKILL.md12.0 KB
  • __init__.py0 B
  • collect.py2.6 KB
  • prompts/enrich_bucket.md5.4 KB
  • prompts/enrich.md7.4 KB
  • scripts/__init__.py0 B
  • scripts/bucket_overview.py3.9 KB
  • scripts/bucket_scan.py16.8 KB
  • scripts/changes.py2.0 KB
  • scripts/cleanup_json.py1.8 KB
  • scripts/dataset_all.py11.5 KB
  • scripts/dataset.py7.2 KB
  • scripts/db_mtime.py815 B
  • scripts/list_datasets.py1.1 KB
  • scripts/plan.py8.2 KB
  • scripts/render_index.py8.0 KB
  • scripts/schema.py2.6 KB
  • scripts/summary.py18.3 KB
  • scripts/utils.py16.5 KB
  • snapshot.py7.5 KB
  • types.py3.1 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Maintain a knowledge base at dc-knowledge/. .md files are the persistent output. .json files are intermediate (generated in Step 3, consumed in Step 4, then deleted).

{core_skill_dir}/SDK.md owns how the pipeline code is written — dataset shape, row grain, provenance, naming. This file owns the knowledge base and how the work is run: what already exists, what a run will cost, and where the results land.

Critical Rules

  1. Path is dc-knowledge/ — NOT .datachain/. The .datachain/ directory is the internal database; the knowledge base lives at dc-knowledge/.
  2. Never pass update=True to dc.read_storage() in query or exploration code unless the user explicitly asks to refresh the listing. Build scripts that read storage are the exception — they pass update=True, delta=True.
  3. Prefer DataChain operations over plain Python for all metadata analysis.
  4. Bounded output — JSON and markdown files stay small regardless of data size.
  5. Stop on auth/connection errors — bucket_scan.py runs a fast access check. If it exits with an error JSON on stderr, stop immediately and show the error to the user. Do not retry with different regions, profiles, or endpoints — ask for the missing credentials.
  6. Follow the enrichment prompt template literally in Step 4. Downstream tooling (render_index.py) parses the exact frontmatter the prompt prescribes.

Common gotchas in UDF scripts

  • parallel=N vs workers=N. parallel=N is local multiprocessing (works anywhere). workers=N is Studio-only and MUST be guarded: chain = chain.settings(parallel=N); if dc.is_studio(): chain = chain.settings(workers=N).
  • No from __future__ import annotations in UDF modules. It stringifies type hints and DataChain's signal-schema resolution rejects the string-vs-class mismatch.
  • Type the UDF return precisely. Iterator[object] / Iterator[Any] / bare dict fail schema resolution. Return a specific Iterator[T], a Pydantic BaseModel, or a primitive.
  • Generators aren't subscriptable. Iterators returned by file APIs do not support [:N]. Use enumerate + break, or list(...) only when the result is genuinely small.
  • Use datachain.__version__ to get the package version (e.g. dc.__version__).

Running the work

Reuse before building

Read dc-knowledge/index.md first. When an existing dataset covers the task — even partially — read it with dc.read_dataset(...) and filter / merge / extend from there instead of going back to raw storage. Re-running a pass that already ran is the most expensive mistake available here. Say which dataset was reused and what it saved.

Save what was expensive

A UDF that ran a model, decoded file bodies, or called a paid API produces rows worth keeping: save that operation's full output, unfiltered, under a descriptive name with a description=. Chains that only list, filter, or select are cheap to recompute and need no dataset. Cost is the only criterion — there is no hierarchy of datasets that has to be built.

Estimate before a long run

Quote a number before starting anything that may run for minutes:

wall ≈ files × per-row × 1.5 / parallel
Op classPer-row
header / metadata parse (bounded-prefix reads)~1 ms
file-body decodesize / 10-50 MB/s
small CPU model (text, light CV)5-50 ms
mid CPU model (detection, segmentation)50-500 ms
streaming CPU model (ASR, audio)0.1-0.5× realtime
local GPU10-100× faster than the CPU row
paid API (LLM / VLM)$0.001-0.01 per row + 0.5-2 s, rate-limited

Measure instead of estimating when the implementation is untested, the model or library has no row in the table, or files are large enough that decode dominates: run 3-5 items with .persist() (never .save()), budget 60 s, and extrapolate wall_full = (wall_sample / N) × total_files × 1.5. Kill at 60 s and fall back to the estimate.

Label which is which — estimated ~X or measured on N=5: ~X. Never present an estimate as a measurement, and write not measured literally when nothing was.

Watch the first minutes

For any run estimated over 5 minutes, read the throughput line DataChain prints (Processed: N rows [elapsed, rate]) over the first 60-90 s:

  • at or above ~0.66× the expected rate → carry on;
  • below ~0.5× → kill it, report the gap and the revised estimate;
  • no throughput line within 2 minutes → kill it and investigate (model download, auth retry, startup cost).

Results land in datasets

Aggregations and final answers are DataChain chains — .filter(), .group_by(), .mutate(), .distinct() — ending in .save(). Two bypasses are forbidden: writing results to .json / .csv / .parquet through open(), json.dump or pandas.to_csv, and pulling rows out with .to_iter() / .to_list() to walk them in Python loops and print the answer. Both leave no dataset, no lineage and no KB record, so the next session recomputes everything. .show() on a saved dataset is fine. If answering needs a Python loop over nested lists, the row grain is wrong — see "Row shape" in SDK.md.

One script per stage

A pipeline that produces several datasets is several scripts, each named after the dataset it produces, each with exactly one .save(). Never batch or shard by hand: DataChain checkpoints UDF progress, so re-running a killed script resumes where it stopped.


Workflow Mode Detection

Mode A — Discovery/Exploration (e.g., "what datasets exist", "show schema", "explore bucket"): → If the user references a specific bucket URI, run Step 1 (Bucket Enlistment) for its root first. → Then run Steps 2–7.

Mode B — Dataset Creation/Pipeline (e.g., "create dataset X from ...", "process files and save"):

Precondition (do this FIRST — before ANY tool call):

$ cat dc-knowledge/index.md

If index.md exists and the task can be solved by reading an existing dataset, do not write a pipeline — read it directly with dc.read_dataset("name") and filter/merge/extend from there. This avoids recomputing expensive operations.

Never parse files under dc-knowledge/datasets/*.json or dc-knowledge/buckets/**/*.json directly — those are pre-render intermediates that get deleted. The information you need is in index.md.

If dc-knowledge/index.md does not exist, proceed with Steps 1–7 to build it.

→ If the pipeline reads from a bucket, run Step 1 (Bucket Enlistment) for the bucket root first. → Run the access check (if not already done in Step 1): datachain bucket status <uri>. If not found / denied, stop and ask for credentials. → Read {core_skill_dir}/SDK.md for DataChain SDK rules. → Work through "Running the work" above — reuse, estimate, then write the script. → While the pipeline is running, enrich any Step 1 bucket JSON that does not yet have a .md (parallel work). → After the pipeline completes, run Steps 2–7 to update the knowledge base. → Report both: pipeline result AND knowledge base update status.

Mode C — Script Execution (e.g., user runs an existing .py file that touches data): → If the script references bucket URIs, run Step 1 for each bucket root first. → Scripts can create datasets as side effects. → While the script is running, enrich Step 1 bucket JSON in parallel. → After ANY data-related script finishes, run Steps 2–7 to detect and record new/changed datasets.

Mode D — Knowledge Base Maintenance (e.g., "update the knowledge base", "refresh dataset docs"): → Run Steps 2–7. Existing session context in .md files is preserved automatically during re-enrichment.


Step 1 — Bucket Enlistment

When any storage URI is encountered, enlist the whole bucket first.

  1. Extract bucket root. From any URI, derive {scheme}://{bucket}/.
  2. Check if already enlisted. Look for dc-knowledge/buckets/{scheme}/{bucket_slug}.md or .json. If either exists, skip.
  3. Access check. Run datachain bucket status {root_uri}. If denied / not found, stop and ask.
  4. Scan with timeout. Default 60s; user can override:
    python3 {skill_dir}/scripts/bucket_scan.py {root_uri} \
      --output dc-knowledge/buckets/{scheme}/{bucket_slug}.json --timeout 60
    
  5. Handle timeout (exit code 124). Run the hierarchical fallback:
    python3 {skill_dir}/scripts/bucket_overview.py {root_uri} \
      --bucket-json dc-knowledge/buckets/{scheme}/{bucket_slug}.json
    
  6. Report. "Enlisted bucket {bucket} — {N} files, total size {size}, primarily {top 2-3 extensions}." Do not enrich here; Step 4 batches it.

Step 1 runs once per bucket root per session.


Step 2 — Sync

python3 {skill_dir}/scripts/plan.py [--studio] --output dc-knowledge/.plan.json

Buckets are auto-discovered from catalog listings. Do not add --studio unless requested. If "up_to_date": true, print "Knowledge base is up to date." and stop. Entries with status of "new" or "stale" need processing in Step 3.


Step 3 — Save Data

For each dataset where status != "ok":

python3 {skill_dir}/scripts/dataset_all.py <name> \
  --plan dc-knowledge/.plan.json --output dc-knowledge/<file_path>.json

For each bucket where status != "ok" (and not enlisted in Step 1):

python3 {skill_dir}/scripts/bucket_scan.py <uri> --output dc-knowledge/<file_path>.json

Run independent calls concurrently.


Step 4 — Enrich

Generate .md from .json for each entry processed in Step 3 (and any Step 1 bucket JSON that lacks a .md).

  • Datasets: read {skill_dir}/prompts/enrich.md, then write dc-knowledge/<file_path>.md per the template.
  • Buckets: read {skill_dir}/prompts/enrich_bucket.md, then write dc-knowledge/<file_path>.md.

The prompt template is authoritative — downstream tooling parses the exact frontmatter it prescribes. Skip this step only if the user requests raw output only.


Step 5 — Build Index

python3 {skill_dir}/scripts/render_index.py --plan dc-knowledge/.plan.json --output dc-knowledge/index.md

Step 6 — Cleanup

python3 {skill_dir}/scripts/cleanup_json.py --plan dc-knowledge/.plan.json

Keeps .plan.json for Step 7. Skip if the user asks to retain JSON for debugging.


Step 7 — Report

Knowledge base updated: <N> datasets (<M> updated, <K> unchanged), <B> buckets (<X> scanned, <Y> unchanged).

If any buckets have listing_expired: true, add:

Warning: Listing for <bucket> is expired (last scanned: <date>). Run dc.read_storage("<uri>", update=True) to refresh.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Use ONLY for abstract DataChain SDK questions — API usage, method signatures, or code patterns — when no specific dataset or bucket is referenced. If the request mentions creating, saving, listing, exploring datasets or buckets, use datachain-knowledge instead.

日本語の概要は準備中です。原文の説明を表示しています。

datachain-ai/datachain2,8252026年10月11日 更新

datachain-ai のスキルをすべて見る

このスキルの問題を報告する