本文へ移動
cccskills
無料GitHub で公開

fg-data-profiling

Guides fg-data-profiling data profiling, report generation, configuration, CLI, comparison, privacy, and optional Spark/notebook integration workflows.

インストール方法を見る

含まれるファイル(37)

  • SKILL.md5.3 KB
  • references/development-and-maintenance.md3.3 KB
  • references/package-overview.md4.0 KB
  • references/repo-provenance.md2.8 KB
  • references/repo-routing-metadata.json347 B
  • references/troubleshooting.md4.1 KB
  • scripts/check_environment.py4.1 KB
  • sub-skills/cli-and-automation/references/automation-recipes.md2.3 KB
  • sub-skills/cli-and-automation/references/cli-reference.md2.3 KB
  • sub-skills/cli-and-automation/references/troubleshooting.md2.1 KB
  • sub-skills/cli-and-automation/scripts/cli_smoke.py2.5 KB
  • sub-skills/cli-and-automation/SKILL.md2.9 KB
  • sub-skills/comparison-and-quality/references/comparison-workflows.md2.3 KB
  • sub-skills/comparison-and-quality/references/privacy-and-metadata.md2.0 KB
  • sub-skills/comparison-and-quality/references/quality-outputs.md1.6 KB
  • sub-skills/comparison-and-quality/references/troubleshooting.md1.8 KB
  • sub-skills/comparison-and-quality/scripts/compare_reports_smoke.py1.9 KB
  • sub-skills/comparison-and-quality/scripts/sensitive_report_smoke.py2.3 KB
  • sub-skills/comparison-and-quality/SKILL.md3.5 KB
  • sub-skills/configuration-and-output/references/cache-and-serialization.md1.9 KB
  • sub-skills/configuration-and-output/references/configuration-reference.md3.5 KB
  • sub-skills/configuration-and-output/references/output-and-rendering.md2.6 KB
  • sub-skills/configuration-and-output/references/troubleshooting.md1.9 KB
  • sub-skills/configuration-and-output/scripts/write_minimal_config.py2.8 KB
  • sub-skills/configuration-and-output/SKILL.md3.7 KB
  • sub-skills/integrations-and-backends/references/integrations.md1.9 KB
  • sub-skills/integrations-and-backends/references/optional-dependencies.md1.7 KB
  • sub-skills/integrations-and-backends/references/spark-backend.md1.6 KB
  • sub-skills/integrations-and-backends/references/troubleshooting.md1.4 KB
  • sub-skills/integrations-and-backends/scripts/check_spark_readiness.py3.5 KB
  • sub-skills/integrations-and-backends/SKILL.md3.0 KB
  • sub-skills/profiling-workflows/references/api-reference.md4.1 KB
  • sub-skills/profiling-workflows/references/data-formats.md2.7 KB
  • sub-skills/profiling-workflows/references/troubleshooting.md2.5 KB
  • sub-skills/profiling-workflows/references/workflows.md3.4 KB
  • sub-skills/profiling-workflows/scripts/profile_dataframe_smoke.py2.6 KB
  • sub-skills/profiling-workflows/SKILL.md4.2 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

fg-data-profiling

Use this repo skill when a task involves the fg-data-profiling Python package or its public import data_profiling: exploratory data analysis reports, data-quality profiling, HTML/JSON export, command-line report generation, comparison reports, sensitive-data-safe reports, configuration, or optional Spark/notebook integrations.

Do not use this skill as proof that the original repository checkout is present. All runnable helpers and operational references needed by future agents are bundled in this skill tree.

Install and import check

For normal package usage, install the public distribution and import the current module name:

python -m pip install -U fg-data-profiling
python - <<'PY'
import data_profiling
from data_profiling import ProfileReport, compare
print(data_profiling.__version__)
print(ProfileReport)
print(compare)
PY

The repository also exposes a deprecated compatibility import named ydata_profiling and a legacy CLI name pandas_profiling; prefer the data_profiling import and data_profiling CLI for new work.

Run scripts/check_environment.py when you need a safe package/import/CLI diagnostic before deciding which sub-skill to read. Read references/repo-provenance.md before refreshing this skill or comparing it with a new checkout.

Route by task

  • Read sub-skills/profiling-workflows/SKILL.md for core pandas ProfileReport usage, df.profile_report(), HTML/JSON/notebook output calls, time-series mode, type schemas, supported file shapes, and a tiny report-generation smoke helper.
  • Read sub-skills/cli-and-automation/SKILL.md for data_profiling / pandas_profiling command-line usage, supported input extensions, parser flags, default output naming, and automation in DAGs or IDE tasks.
  • Read sub-skills/configuration-and-output/SKILL.md for Settings, YAML config files, PROFILE_ environment variables, minimal/explorative section controls, HTML assets/themes, cache invalidation, and serialization.
  • Read sub-skills/comparison-and-quality/SKILL.md for comparing profile reports, sensitive-data redaction, custom samples, dataset metadata, data dictionaries, type schema, quality output access, and Great Expectations caveats.
  • Read sub-skills/integrations-and-backends/SKILL.md for optional Spark, notebook widgets, Bytewax/streaming snapshots, interactive app embedding, other DataFrame libraries, PyCharm integration, optional dependencies, and backend readiness checks.

Shared references

High-value entry points

  • from data_profiling import ProfileReport, compare
  • ProfileReport(df, minimal=True, progress_bar=False).to_file("report.html")
  • df.profile_report(...) after importing data_profiling
  • data_profiling --minimal --silent data.csv report.html
  • profile.to_json(), profile.to_html(), profile.get_description()
  • profile_a.compare(profile_b) or compare([profile_a, profile_b])

Backend and optional dependency honesty

The verified core scope is CPU/Pandas package usage. Spark support is a first-class optional workflow, but it requires Java plus PySpark and was not verified in the production environment used to create this skill. Notebook widgets require the notebook extra and active widget support. Great Expectations support remains in source APIs, but the public docs state that current versions no longer support that integration; treat it as a legacy/compatibility surface unless the user pins compatible versions.

Default response pattern

  1. Identify whether the user wants API, CLI, configuration, comparison/privacy, or optional integration guidance.
  2. Read the matching sub-skill and its nearest references/scripts.
  3. Prefer tiny local fixtures and bundled helpers for validation before running expensive profiles on real data.
  4. Avoid source-checkout dependencies: do not tell the user to open or run repository examples, docs, or tests unless they explicitly ask to maintain a repo checkout.
  5. Preserve privacy: for sensitive datasets, route to comparison-and-quality before suggesting samples, duplicates, or numeric treatment of identifiers.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Routes 3D ResNets PyTorch video action-recognition workflows across training, inference, and data preparation.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

3ddfa

無料

Guide 3DDFA Python inference, geometry rendering, training/evaluation, and optional C++ ONNX workflows for 3D dense face alignment.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

3ddfa-v2

無料

Routes 3DDFA_V2 face-alignment setup, still-image demos, video tracking, and ONNX benchmarking workflows.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

ab3dmot

無料

Operate AB3DMOT 3D multi-object tracking workflows for KITTI and nuScenes data, tracking, evaluation, and visualization.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

Use Hugging Face Accelerate for PyTorch training-loop migration, distributed launch/configuration, DeepSpeed/FSDP/TPU backend setup, big-model inference/offload, checkpointing, tracking, and troubleshooting.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

acme

無料

Route Acme reinforcement-learning framework tasks across core loops, replay/data, JAX agents, and TensorFlow/Sonnet agents.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

VectorSpaceLab のスキルをすべて見る

このスキルの問題を報告する