本文へ移動
cccskills
無料GitHub で公開

dataops

DataOps and DS/MLOps testing reference for data tiering, Bronze-to-Silver validation placement, pipeline invariants, pytest categories, and validation-versus-drift. Use when designing, reviewing, or generating data pipelines, transformation code, data validation, or data-science test suites.

インストール方法を見る

含まれるファイル(16)

  • SKILL.md7.6 KB
  • assets/synthetic-data-operation-v1.schema.json9.2 KB
  • pyproject.toml681 B
  • references/data-tiers-and-pipeline-invariants.md10.2 KB
  • references/provenance.md10.8 KB
  • references/synthetic-data-operation-contract.md6.4 KB
  • references/testing-data-science-and-mlops.md8.8 KB
  • references/validation-drift-and-observability.md7.2 KB
  • scripts/synthetic_data_operation.py21.2 KB
  • tests/corpus/0_empty0 B
  • tests/corpus/1_empty_object2 B
  • tests/corpus/2_non_object_json2 B
  • tests/corpus/3_minimal_preflight27 B
  • tests/fuzz_harness.py1.2 KB
  • tests/test_synthetic_data_operation.py23.9 KB
  • uv.lock101.9 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

DataOps Reference Pack

Goal

Ground pipeline and test generation in the Microsoft CSE engineering playbook so that data tier semantics, validation placement, recovery invariants, and DS/MLOps test technique are applied consistently and attributed accurately.

Inputs

  • The pipeline, transformation, validation, or test work under discussion
  • The tier of each dataset involved, when the consuming workflow records one
  • Existing test layout, package structure, and data-access boundaries
  • A data classification produced elsewhere, when sensitivity matters

Reference index

Read only the reference that matches the active concern.

ReferenceRead this when
data-tiers-and-pipeline-invariants.mdAssigning tier meaning, placing validation, routing malformed records, or asserting replay, idempotency, testability, source-control, and configuration invariants
testing-data-science-and-mlops.mdWriting or reviewing tests for data loading, transformation, model load or predict, data validation, or model robustness
validation-drift-and-observability.mdDistinguishing data validation from drift detection, choosing remediation, or deciding which data and model signals matter
synthetic-data-operation-contract.mdValidating synthetic-data preflight authority, field lineage, conditional subgroup evidence, result linkage, or guarded local replacement
provenance.mdConfirming what is upstream guidance, what is HVE Core derivation, and where upstream is silent

Success criteria

  • Tier language distinguishes the three upstream quality tiers from the additional storage areas.
  • Validation is placed at the Bronze-to-Silver boundary, and the faithful-copy rationale with both replay purposes travels with that placement whether or not the request challenges it.
  • Test guidance names the operation category, its technique, and where mocking stops.
  • Validation and drift keep their distinct definitions and their distinct remediation paths.
  • Every claim traces to an attributed upstream source or is labelled as HVE Core guidance.
  • Synthetic-data generation stops before source access or writes unless its applicable SYNTHETIC_DATA_OPERATION_V1 preflight passes deterministic validation.

Constraints

  • This skill generates code, assertions, and review guidance. It does not execute pipelines, transformation engines, or telemetry backends.
  • Reproduce only the minimum text necessary for a specific technical point, and paraphrase everything else. Attribute every reference and describe accurately what each reference reproduces.
  • Label HVE Core derivations as such. Do not present a derived consequence or a repository convention as upstream guidance.
  • Where upstream is silent, say so rather than inventing an upstream-sounding rule.

Ownership boundaries

This skill decides which data and model signals matter. It does not own the vocabulary, the classification, or the surrounding workflow.

ConcernOwner
Metric names, instrument types, units, cardinality discipline, and the PII emission denylisttelemetry-foundations
Data sensitivity classification and DPIA thresholdsprivacy-standards
Entity semantics, relationships, and which tier a dataset is recorded asThe calling workflow
Feasibility assessment and go/no-go recommendationThe calling workflow

No repository artifact currently owns the last two rows. When the caller supplies neither, state the gap rather than deciding tier assignment or feasibility here.

This skill never decides what is sensitive. It reads a classification produced elsewhere.

For synthetic-data operations, read synthetic-data-operation-contract.md and validate records against synthetic-data-operation-v1.schema.json. The contract records qualified decisions by immutable reference; it does not store source values or replace the authority of data, privacy, Responsible AI, fairness, or domain owners.

Stop rules

  • Stop and route metric naming, units, and cardinality to telemetry-foundations, the OpenTelemetry-aligned vocabulary and instrumentation skill. Route data sensitivity to privacy-standards, the privacy classification and DPIA-threshold reference.
  • Stop and state the gap when the request depends on guidance the playbook does not provide, such as a drift threshold or an alerting policy.
  • Stop and offer the correct placement when asked to validate before Bronze landing, rather than complying or refusing without an alternative.

Attribution

This pack declares CC-BY-4.0 because both bodies of content it holds carry that license.

Source pages are Microsoft CSE Code With Engineering Playbook documentation licensed CC BY 4.0. The upstream project applies MIT through a separate LICENSE-CODE file to code samples only, which this pack does not reproduce. The references derive from those documentation pages and have been changed: upstream guidance is paraphrased, and only identifiers and structural names are carried across as facts. THIRD-PARTY-NOTICES carries the attribution CC BY 4.0 requires and states that the content has been changed. Each reference cites its own upstream URL and states what it reproduces.

Content labelled as HVE Core derivation is repository-original material under CC BY 4.0. See provenance.md for the consolidated source map and derivation labels.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Consolidated accessibility skill entrypoint for WCAG 2.2, ARIA Authoring Practices, cognitive accessibility, Section 508, EN 301 549, design intent verification, and the Accessibility Planner workflow.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月10日 更新

Build, refresh, report, or probe an accessibility coverage matrix across criteria, surfaces, and evidence methods. Use when assessing coverage with the accessibility runtime harness and generated evidence bundle.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月10日 更新

Authoring skill for Architecture Decision Records (ADRs) supporting capture, from-planner-handoff, and adopt-template entry modes with selectable Y-Statement or MADR v4.0.0 output templates, supersession lineage, and ASR trigger evaluation.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月10日 更新

Authoring conventions for exploratory data analysis notebooks and analytical dashboards, covering section sequence, visualization selection, scale thresholds, caching and state, and dashboard validation budgets. Use when composing or reviewing an EDA notebook, an analytical dashboard, or a dashboard test pass.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月10日 更新

Architecture diagram authoring for cloud infrastructure and declared data catalogs. Use when rendering Azure IaC or DS_CATALOG_V1 relationships as caller-selected ASCII or Mermaid diagrams.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月10日 更新

Create a durable Architecture Review Record from a confirmed System Architecture Reviewer scope, evidence, pillar analysis, trade-offs, and dispositions

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月10日 更新

microsoft のスキルをすべて見る

このスキルの問題を報告する