Consolidated accessibility skill entrypoint for WCAG 2.2, ARIA Authoring Practices, cognitive accessibility, Section 508, EN 301 549, design intent verification, and the Accessibility Planner workflow.
日本語の概要は準備中です。原文の説明を表示しています。
Design evaluation datasets and supporting documentation for AI systems and agents, covering the scoping interview, difficulty distribution, dataset contract, sample review, and metric and tooling selection. Use when building or reviewing an evaluation set for a conversational agent, assistant, or retrieval-grounded AI system.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Produce an evaluation dataset and its supporting documentation that measure whether an AI system does its job, refuses what it should refuse, and behaves acceptably under pressure. The dataset is a durable customer artifact, so its scope, balance, and rationale are recorded rather than implied.
uv run python scripts/validate_evaluation_dataset.py --json <dataset.json> --csv <dataset.csv> before treating machine facts as evidence.Activate rpi-research only when the confirmed evaluation job needs current evaluator names, availability, preview state, platform compatibility, or prerequisites that the authoritative live source must establish. Provide the platform and product-version scope, system context, tool-use pattern, metric-plan decision, source and date boundaries, evidence criteria, non-goals, and supplied evaluation evidence. Use analysis or comparison mode and the default Research evidence root.
Read the completed primary artifact before naming current evaluators or committing a platform-backed metric plan. Record the authoritative source and retrieval date. Research establishes current facts; this skill still selects metrics from the confirmed system grounding, tool use, risk profile, operating constraints, and evaluation cadence.
Treat Blocked and Needs clarification as unresolved current-fact evidence. Stop only the catalog-backed recommendation, record the smallest gap, and do not invent names, availability, compatibility, or preview state. If rpi-research or a required lookup capability is unavailable, report the limitation rather than substituting training-data claims.
| Concern | Owner |
|---|---|
| Trained-model evaluation, tracking, reproducibility, and production readiness | ml-experimentation |
| Responsible AI assessment, risk classification, and approval | rai-planner |
| Session state, job lifecycle, and durable-write gating | data-science-engineering-foundation |
| Notebook and dashboard authoring conventions | analysis-authoring |
| Dataset entity semantics and profile contracts | data-catalog |
This skill covers evaluation of AI systems whose output is a response: assistants, conversational agents, and retrieval-grounded applications. Evaluating a trained model's predictive performance is a different concern and belongs to ml-experimentation.
rai-planner when the risk needs assessment, classification, severity, likelihood, approval, or another decision beyond detection coverage.| Resource | Use |
|---|---|
| evaluation-interview-and-review.md | Read before scoping; contains the interview areas, distribution rules, and sample-review protocol |
| metric-selection-and-tooling.md | Read when selecting metrics and recommending evaluation tooling |
| provenance.md | Read for source, licensing, and currency posture on external evaluator vocabulary |
| evaluation-dataset-contract.md | Copy as the dataset's machine-readable shape |
| supporting-documents.md | Copy as the single sectioned evaluation-guide skeleton |
| evaluation-dataset-v1.schema.json | Execute through the validator to check the version 1.0.0 JSON contract |
| validate_evaluation_dataset.py | Execute against sibling JSON and CSV files before using their machine facts as evidence |
The interview structure, distribution defaults, dataset contract, review protocol, and document skeletons are repository-original content licensed CC BY 4.0. External evaluator names are cited as factual identifiers; their authoritative definitions remain with the vendor documentation identified in provenance.md. No upstream text is reproduced.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Consolidated accessibility skill entrypoint for WCAG 2.2, ARIA Authoring Practices, cognitive accessibility, Section 508, EN 301 549, design intent verification, and the Accessibility Planner workflow.
日本語の概要は準備中です。原文の説明を表示しています。
Build, refresh, report, or probe an accessibility coverage matrix across criteria, surfaces, and evidence methods. Use when assessing coverage with the accessibility runtime harness and generated evidence bundle.
日本語の概要は準備中です。原文の説明を表示しています。
Authoring skill for Architecture Decision Records (ADRs) supporting capture, from-planner-handoff, and adopt-template entry modes with selectable Y-Statement or MADR v4.0.0 output templates, supersession lineage, and ASR trigger evaluation.
日本語の概要は準備中です。原文の説明を表示しています。
Authoring conventions for exploratory data analysis notebooks and analytical dashboards, covering section sequence, visualization selection, scale thresholds, caching and state, and dashboard validation budgets. Use when composing or reviewing an EDA notebook, an analytical dashboard, or a dashboard test pass.
日本語の概要は準備中です。原文の説明を表示しています。
Architecture diagram authoring for cloud infrastructure and declared data catalogs. Use when rendering Azure IaC or DS_CATALOG_V1 relationships as caller-selected ASCII or Mermaid diagrams.
日本語の概要は準備中です。原文の説明を表示しています。
Create a durable Architecture Review Record from a confirmed System Architecture Reviewer scope, evidence, pillar analysis, trade-offs, and dispositions
日本語の概要は準備中です。原文の説明を表示しています。