本文へ移動
cccskills
無料GitHub で公開

dataset-quality-audit

Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and actionable suggestions. Triggered when users ask to check data quality, find missing or duplicate values, detect outliers, validate formats, profile data, or clean data.

インストール方法を見る

含まれるファイル(3)

  • SKILL.md3.9 KB
  • LICENSE1.1 KB
  • scripts/data_quality_checker.py24.5 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

dataset-quality-audit

A data quality auditing tool that runs 12-dimension quality checks on tabular data, producing per-dimension scores (0–100), an overall grade, and actionable fix suggestions.

Capabilities

DimensionDescription
Missing ValuesCount and percentage of null/NaN values per column
Duplicate RowsNumber and percentage of fully duplicated rows
Type ConsistencyMixed types within a single column (e.g., numbers mixed with text)
Value Range / OutliersOutlier detection using the IQR method
Format ComplianceConsistency of date, email, phone number, and other formatted fields
Uniqueness ConstraintsWhether ID-type columns contain duplicates
Whitespace IssuesLeading/trailing spaces, empty strings, whitespace-only values
Constant ColumnsColumns with only a single unique value (zero information)
Distribution SkewnessWhether numeric columns have excessive skewness
Column NamingSpaces, special characters, or inconsistent casing in column names
Cardinality AnomaliesUnusually high or low number of unique values
Cross-Column ConsistencyLogical checks across columns (e.g., start date before end date)

Quick Start

# Basic quality check
python3 scripts/data_quality_checker.py data.csv

# Save report as JSON
python3 scripts/data_quality_checker.py data.csv --output report.json

# Specify ID columns (for uniqueness checks)
python3 scripts/data_quality_checker.py users.csv --id-columns "user_id,email"

# Specify date columns (for format checks)
python3 scripts/data_quality_checker.py orders.csv --date-columns "created_at,updated_at"

Detailed Usage

Basic Invocation

python3 scripts/data_quality_checker.py <data-file> [options]

Parameters

ParameterShortRequiredDefaultDescription
input—Yes—Path to input file (CSV/TSV/Excel/JSON)
--output-oNostdoutPath for the JSON report output
--id-columns-idNoAuto-detectComma-separated column names that should be unique
--date-columns-dcNoAuto-detectComma-separated column names containing dates
--sample-sNoAll rowsNumber of rows to sample (useful for large files)
--encoding-eNoutf-8File encoding

Output Format (JSON)

{
  "file": "data.csv",
  "rows": 10000,
  "columns": 15,
  "overall_score": 78.5,
  "grade": "B",
  "dimensions": {
    "missing_values": {
      "score": 85.0,
      "issues": [
        {"column": "age", "missing_count": 150, "missing_pct": 1.5, "suggestion": "Fill with median or mode"}
      ]
    },
    "duplicates": {
      "score": 95.0,
      "issues": [...]
    }
  },
  "top_suggestions": [
    "Column 'age' has 1.5% missing values — consider filling with the median",
    "Found 200 fully duplicated rows — consider deduplication"
  ]
}

Grading Scale

GradeScore RangeMeaning
A+95–100Excellent quality — ready for use as-is
A90–95Good quality — minor issues only
B80–90Moderate quality — recommended to fix before use
C60–80Poor quality — significant cleaning required
D40–60Very poor quality — many issues need attention
F0–40Essentially unusable — requires re-collection or major cleanup

Dependencies

  • Python 3.8+
  • pandas
  • numpy
pip install pandas numpy

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Simulates academic peer review, evaluating papers across Originality, Methodology, Results, and Writing to provide Major/Minor Revision recommendations with actionable feedback. Triggers when a user asks to "review my paper," "simulate peer review," or "give my paper a peer review.

日本語の概要は準備中です。原文の説明を表示しています。

zebbern/claude-code-guide4,6532026年10月10日 更新

This skill should be used when the user asks to "attack Active Directory", "exploit AD", "Kerberoasting", "DCSync", "pass-the-hash", "BloodHound enumeration", "Golden Ticket", "Silver Ticket", "AS-REP roasting", "NTLM relay", or needs guidance on Windows domain penetration testing.

日本語の概要は準備中です。原文の説明を表示しています。

zebbern/claude-code-guide4,6532026年10月10日 更新

This skill should be used when the user asks to "test API security", "fuzz APIs", "find IDOR vulnerabilities", "test REST API", "test GraphQL", "API penetration testing", "bug bounty API testing", or needs guidance on API security assessment techniques.

日本語の概要は準備中です。原文の説明を表示しています。

zebbern/claude-code-guide4,6532026年10月10日 更新

Generate multiple radically different interface designs for a module using parallel sub-agents. Use when user wants to design an API, explore interface options, compare module shapes, or mentions "design it twice".

日本語の概要は準備中です。原文の説明を表示しています。

zebbern/claude-code-guide4,6532026年10月10日 更新

Interactive system flow tracing across CODE, API, AUTH, DATA, NETWORK layers with SQLite persistence and Mermaid export. Use for security audits, compliance documentation, flow tracing, feature ideation, brainstorming, debugging, architecture reviews, or incident post-mortems. Triggers on audit, trace flow, document flow, security review, debug flow, brainstorm, architecture review, post-mortem, incident review.

日本語の概要は準備中です。原文の説明を表示しています。

zebbern/claude-code-guide4,6532026年10月10日 更新

Authentication patterns: session vs JWT vs OAuth comparison, provider selection (NextAuth, Clerk, Supabase Auth), security checklist, and common mistakes. Use when implementing auth, reviewing auth flows, or choosing auth providers.

日本語の概要は準備中です。原文の説明を表示しています。

zebbern/claude-code-guide4,6532026年10月10日 更新

zebbern のスキルをすべて見る

このスキルの問題を報告する