本文へ移動
cccskills
無料GitHub で公開

data-catalog

Create and enrich durable data catalogs using the native DS_CATALOG_V1 Markdown contract, declared entity relationships, privacy citation fields, and stable relationship IDs. Use when inventorying engagement data, recording semantic relationships, or preparing a catalog for ERD rendering.

インストール方法を見る

含まれるファイル(18)

  • SKILL.md6.1 KB
  • assets/ds-catalog-v1.schema.json5.9 KB
  • examples/northwind-catalog.md8.4 KB
  • pyproject.toml698 B
  • references/catalog-contract.md11.2 KB
  • references/dcat-crosswalk.md5.4 KB
  • references/provenance.md3.1 KB
  • scripts/validate_catalog.py15.0 KB
  • templates/ds-catalog-v1.md5.3 KB
  • tests/corpus/0_valid_frontmatter90 B
  • tests/corpus/1_empty_frontmatter8 B
  • tests/corpus/2_unclosed_sequence20 B
  • tests/corpus/3_duplicate_key24 B
  • tests/corpus/4_no_frontmatter15 B
  • tests/corpus/README.md1.2 KB
  • tests/fuzz_harness.py1.2 KB
  • tests/test_validate_catalog.py24.0 KB
  • uv.lock107.7 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Data Catalog Workflow

Goal

Produce a customer-readable Markdown catalog whose YAML frontmatter is a valid DS_CATALOG_V1 machine contract. Preserve uncertainty explicitly so inferred or assumed relationships never appear confirmed.

Flow

  1. Confirm the engagement name and the caller-approved durable output path.
  2. Inventory entities at business grain. Record source access, tier, volume, profile pointer, classification, lineage, and open questions without copying column-level profile data.
  3. Assign every relationship a stable rel-* identifier. Record endpoints, maximum cardinality, both endpoint minimums, one or more paired join-key fields, confidence, and evidence basis.
  4. Reconcile coverage counts with the entity and relationship records.
  5. Render the human-readable sections from the YAML facts, ending with the canonical Data Science and Engineering Coaching disclaimer footer. Narrative can explain facts but cannot redefine them.
  6. Validate the artifact with scripts/validate_catalog.py before treating it as ready for review.

Inputs

  • Engagement context and a caller-approved output path
  • Data source inventory and access status
  • Business entity names, grain, and declared relationships
  • Existing per-dataset profile paths, when available
  • Privacy classifications or standards citations produced by the owning privacy workflow

Success criteria

  • The frontmatter declares exactly catalog_version: DS_CATALOG_V1 and validates against assets/ds-catalog-v1.schema.json.
  • Entity IDs and relationship IDs are unique and stable. Every endpoint and lineage reference resolves.
  • Every relationship declares cardinality as its maximum multiplicity plus from_minimum and to_minimum as zero or one.
  • Join keys use one string on both sides or paired arrays of equal length, and record field names only without primary-key, foreign-key, or uniqueness roles.
  • Relationship confidence is one of confirmed, inferred, or assumed, and every relationship records its basis.
  • Classification uses the privacy-standards citation-field names. The catalog does not invent standards identifiers.
  • Column statistics and feature metadata remain behind profile_ref rather than being copied into the catalog.
  • Every customer-facing catalog ends with the canonical Data Science and Engineering Coaching disclaimer footer.

Constraints

  • Treat catalog enrichment as user-driven. Offer enrichment when a source is missing, but do not change the catalog silently.
  • Record source locations as paths or connection references, never embedded credentials.
  • Keep DCAT alignment non-binding. The catalog is native YAML, not RDF, and makes no DCAT conformance claim.
  • Keep tier recording separate from tier behavior. data-catalog records the tier; dataops owns what that tier means.
  • Keep rendering separate from semantic authority. Diagram tools consume declared relationships and do not infer new ones.
  • End every customer-facing catalog with the canonical Data Science and Engineering Coaching disclaimer from disclaimer-language.instructions.md.

Stop rules

  • Stop and request clarification when an entity grain, relationship endpoint, join-key pairing, or endpoint minimum is ambiguous.
  • Stop and retain inferred or assumed confidence when evidence does not support confirmed.
  • Stop before a durable write when the caller has not confirmed the destination.
  • Stop and route privacy interpretation to privacy-standards or the Privacy Planner when citation values or DPIA status are unknown.

Package resources

ResourceUse
catalog-contract.mdRead for the authoritative field, identity, multiplicity, and body-generation rules
dcat-crosswalk.mdRead when planning a DCAT, DCAT-AP, or catalog-platform export
provenance.mdRead for standards selection, licensing, and non-conformance boundaries
ds-catalog-v1.mdCopy when starting a new catalog
northwind-catalog.mdRead as a complete valid example with scalar and composite join keys
assets/ds-catalog-v1.schema.jsonUse as the structural JSON Schema for parsed frontmatter
scripts/validate_catalog.pyExecute with uv run python scripts/validate_catalog.py <catalog.md> to validate a catalog

Attribution

The native contract, workflow, schema, template, examples, and validator are repository-original content licensed CC BY 4.0.

The crosswalk paraphrases selected concepts from W3C DCAT 3, W3C PROV-O, and Frictionless Table Schema v2. Standard names and term identifiers are factual citations, the crosswalk prose and the paired-array join-key contract are independently authored, and no upstream schema, example, table, or substantial excerpt is reproduced. Because no upstream expression is redistributed, the package remains solely CC BY 4.0. See provenance.md.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Consolidated accessibility skill entrypoint for WCAG 2.2, ARIA Authoring Practices, cognitive accessibility, Section 508, EN 301 549, design intent verification, and the Accessibility Planner workflow.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月10日 更新

Build, refresh, report, or probe an accessibility coverage matrix across criteria, surfaces, and evidence methods. Use when assessing coverage with the accessibility runtime harness and generated evidence bundle.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月10日 更新

Authoring skill for Architecture Decision Records (ADRs) supporting capture, from-planner-handoff, and adopt-template entry modes with selectable Y-Statement or MADR v4.0.0 output templates, supersession lineage, and ASR trigger evaluation.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月10日 更新

Authoring conventions for exploratory data analysis notebooks and analytical dashboards, covering section sequence, visualization selection, scale thresholds, caching and state, and dashboard validation budgets. Use when composing or reviewing an EDA notebook, an analytical dashboard, or a dashboard test pass.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月10日 更新

Architecture diagram authoring for cloud infrastructure and declared data catalogs. Use when rendering Azure IaC or DS_CATALOG_V1 relationships as caller-selected ASCII or Mermaid diagrams.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月10日 更新

Create a durable Architecture Review Record from a confirmed System Architecture Reviewer scope, evidence, pillar analysis, trade-offs, and dispositions

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月10日 更新

microsoft のスキルをすべて見る

このスキルの問題を報告する