本文へ移動
cccskills
無料GitHub で公開

docs-grounding-verifier

Use this skill to verify CLAIM-LEVEL grounding of a documentation page (or set of pages) against the source code. Activate when you have specific pages to check for factual accuracy -- not when sweeping a whole corpus (use docs-corpus-audit for that) and not when triaging a PR diff (use docs-sync for that). Trigger nouns: "is this doc accurate", "verify the page against the code", "fact-check this section", "any claims that drifted from source", "fact-checking", "grounding audit", "drift hunt", "claim verification". Returns per-claim verdicts (GROUNDED | PARTIAL | CONTRADICTED | UNSUPPORTED) with file:line evidence citations. Catches paragraph-level inaccuracies that page-level audit averages over -- e.g. a paragraph with 5 claims where 4 are grounded and 1 is fabricated. Does NOT modify files (returns advisory only); does NOT re-architect the docs; does NOT triage PRs.

インストール方法を見る

含まれるファイル(101)

  • SKILL.md7.3 KB
  • assets/judge-prompt.md2.6 KB
  • evals/content-evals.json2.2 KB
  • evals/run-evals.sh5.1 KB
  • evals/runs/20260527-194228/seeded-corpus/drift-install-flag/page.md14.1 KB
  • evals/runs/20260527-194228/seeded-corpus/drift-install-flag/scenario.json371 B
  • evals/runs/20260527-194228/seeded-corpus/drift-policy-reject/page.md7.2 KB
  • evals/runs/20260527-194228/seeded-corpus/drift-policy-reject/scenario.json446 B
  • evals/runs/20260527-194228/seeded-corpus/drift-registry-resolver/page.md21.4 KB
  • evals/runs/20260527-194228/seeded-corpus/drift-registry-resolver/scenario.json450 B
  • evals/runs/20260527-194228/trigger-prompts.txt2.1 KB
  • evals/runs/proof/claims/copilot.json4.4 KB
  • evals/runs/proof/claims/install.json4.6 KB
  • evals/runs/proof/claims/pkgtypes.json4.2 KB
  • evals/runs/proof/claims/policy.json3.8 KB
  • evals/runs/proof/claims/registries.json5.0 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c1.json2.6 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c10.json2.0 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c11.json1.8 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c12.json2.7 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c13.json1.6 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c14.json2.4 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c15.json2.0 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c2.json2.6 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c3.json3.5 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c4.json2.6 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c5.json2.6 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c6.json2.2 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c7.json1.9 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c8.json2.8 KB
  • evals/runs/proof/copilot/evidence/docs_src_content_docs_integrations_copilot-app_md_c9.json2.8 KB
  • evals/runs/proof/copilot/judge-batch-docs_src_content_docs_integrations_copilot-app_md.txt45.0 KB
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c1.json3.1 KB
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c10.json1.8 KB
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c11.json971 B
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c12.json1.9 KB
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c13.json1.7 KB
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c14.json3.4 KB
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c15.json3.2 KB
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c2.json973 B
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c3.json1.9 KB
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c4.json1.1 KB
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c5.json1.5 KB
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c6.json1.5 KB
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c7.json5.2 KB
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c8.json1.1 KB
  • evals/runs/proof/install/evidence/docs_src_content_docs_reference_cli_install_md_c9.json896 B
  • evals/runs/proof/install/judge-batch-docs_src_content_docs_reference_cli_install_md.txt38.1 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c1.json3.4 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c10.json2.3 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c11.json1.9 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c12.json2.2 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c13.json2.6 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c14.json1.5 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c15.json3.5 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c2.json3.6 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c3.json2.7 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c4.json4.2 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c5.json3.3 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c6.json2.6 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c7.json2.7 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c8.json1.7 KB
  • evals/runs/proof/pkgtypes/evidence/docs_src_content_docs_reference_package-types_md_c9.json2.0 KB
  • evals/runs/proof/pkgtypes/judge-batch-docs_src_content_docs_reference_package-types_md.txt50.5 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c1.json3.6 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c10.json2.2 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c11.json2.5 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c12.json2.2 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c13.json2.5 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c14.json2.5 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c15.json4.2 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c2.json2.6 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c3.json3.3 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c4.json3.4 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c5.json1.9 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c6.json3.2 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c7.json2.4 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c8.json2.3 KB
  • evals/runs/proof/policy/evidence/docs_src_content_docs_enterprise_apm-policy_md_c9.json1.4 KB
  • evals/runs/proof/policy/judge-batch-docs_src_content_docs_enterprise_apm-policy_md.txt50.8 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c1.json2.9 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c10.json2.5 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c11.json2.6 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c12.json3.3 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c13.json2.2 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c14.json3.3 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c15.json2.7 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c2.json2.7 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c3.json2.1 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c4.json3.0 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c5.json3.6 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c6.json2.6 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c7.json1.2 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c8.json2.8 KB
  • evals/runs/proof/registries/evidence/docs_src_content_docs_guides_registries_md_c9.json1.7 KB
  • evals/runs/proof/registries/judge-batch-docs_src_content_docs_guides_registries_md.txt49.4 KB
  • evals/runs/proof/trigger-responses.json1020 B
  • evals/trigger-evals.json1.5 KB
  • scripts/extract-claims.py2.8 KB
  • scripts/retrieve-evidence.sh2.7 KB
  • scripts/verify-page.sh2.4 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

docs-grounding-verifier

CLAIM-LEVEL grounding verification. Adapts the RAGAS faithfulness-eval pattern (proven in RAG literature) to docs/code instead of generated- answers/retrieved-context. Source code is the ground truth; docs paragraphs are the candidate text under audit.

python-architect persona doc-writer persona

Sibling contract

This skill is a SIBLING of docs-corpus-audit and docs-sync. The boundary is load-bearing:

SkillTriggerScopeGranularity
docs-syncPR opened/synchronizedPR diff onlyPage-level
docs-corpus-auditMaintainer asks for whole-corpus passEntire corpusPage-level
docs-grounding-verifierVerify specific pages factually1..N pagesCLAIM-level

docs-corpus-audit invokes this skill in its VERIFY phase on the highest-risk pages of each wave. docs-sync can invoke it on the specific pages in a PR diff. The skill is also runnable standalone.

When to activate

  • Maintainer says "verify <page> against the code".
  • An audit wave wants per-claim grounding scores for its highest-risk pages.
  • A PR review wants to confirm that prose changes are not just plausible but actually consistent with the implementation.
  • A "fact-check" or "grounding" or "drift hunt" request.

When NOT to activate

  • Whole-corpus sweep with no specific page list -> use docs-corpus-audit.
  • PR review with mixed code+docs diff -> use docs-sync.
  • Editorial / tone review -> use editorial-owner persona directly.

Architecture (PIPELINE-of-PANELS)

PARENT
  -> [Stage 1: EXTRACT claims, fan-out PANEL]
       per page -> LLM extracts atomic factual claims as JSON
       script: scripts/extract-claims.py
  -> [Stage 2: RETRIEVE evidence, deterministic S7]
       per claim -> grep over src/ via keywords + hints
       script: scripts/retrieve-evidence.sh   (NO LLM)
  -> [Stage 3: JUDGE grounding, adversarial A7]
       per (claim, evidence) -> LLM rules GROUNDED|PARTIAL|CONTRADICTED|UNSUPPORTED
       asset: assets/judge-prompt.md
  -> [Stage 4: SYNTHESIZE]
       aggregate ungrounded -> doc-writer for fix
       re-verify after fix (A8 ALIGNMENT LOOP)

Stage 2 is the load-bearing design choice: evidence retrieval is DETERMINISTIC (grep + AST hints), not LLM. The judge in Stage 3 can only rule on evidence it actually receives -- it cannot hallucinate support that the retriever did not find. This is the structural guard against the failure mode "the LLM convinces itself the docs match the code."

Phase 1: SCOPE

Input: list of page paths to verify (1..N). If a risk_class is attached (e.g. "high-stakes"), prefer it; otherwise treat all as equal.

Out-of-scope:

  • Pages outside docs/src/content/docs/ or packages/apm-guide/.apm/skills/apm-usage/.
  • Pages with no factual claims (pure editorial / landing). Skip rather than force-extract.

Phase 2: EXTRACT (parallel)

For each page, dispatch ONE claim-extractor agent:

  • Prompt template: scripts/extract-claims.py <page> produces the prompt and embeds the page content.
  • Returns: JSON {"page", "claims":[{"id","text","section","keywords", "expected_source_areas"}]} capped at 15 claims per page.

Parallel safe; no shared state between extractors.

Phase 3: RETRIEVE (deterministic, batched)

For each claim, pipe to scripts/retrieve-evidence.sh:

  • Uses keywords + expected_source_areas to grep src/.
  • Returns one-line JSON: {"claim_id","claim_text","evidence":[...], "evidence_count"}.

Sequential is fine (grep is fast). No LLM. Diagnostics on stderr, data on stdout.

Phase 4: JUDGE (parallel)

For each (claim, evidence) tuple, dispatch ONE grounding-judge agent:

  • Load assets/judge-prompt.md.
  • Send the prompt + the tuple.
  • Returns: JSON verdict per the schema in judge-prompt.md.

Batching across claims-of-one-page into a single judge call is fine (prompt with all tuples at once). Across pages, fan out.

Phase 5: SYNTHESIZE

Aggregate verdicts. Materialize the report:

{
  "summary": {
    "pages_verified": N,
    "claims_total": N,
    "grounded": N, "partial": N, "contradicted": N, "unsupported": N,
    "grounding_rate": N/total
  },
  "actionable": [
    {"page", "claim", "verdict", "evidence_cited", "fix_suggestion"}
  ]
}

CONTRADICTED and PARTIAL are doc-writer work items. UNSUPPORTED is split: if retrieval_fix_suggestion is plausible, retry retrieval with the suggested keywords; if still empty, treat as CONTRADICTED.

Phase 6: ALIGNMENT LOOP (A8)

Hand actionable items to doc-writer (one subagent per page). After edits, RE-RUN the pipeline on the same pages. The grounding_rate must MONOTONICALLY INCREASE between iterations or the loop has diverged -- stop and escalate to the operator.

Ship gate

  • grounding_rate >= 0.9 on each verified page after the alignment loop.
  • Every CONTRADICTED claim cited a specific code file:line that disproves it -- not vague "the code doesn't say that".
  • The eval-runner (see evals/) passes on the trigger evals and the content evals before the skill is treated as production-ready.

Bundled assets

  • scripts/extract-claims.py -- Stage 1 prompt builder. --help, --schema.
  • scripts/retrieve-evidence.sh -- Stage 2 retriever. Deterministic. --help.
  • scripts/verify-page.sh -- end-to-end orchestrator. --help.
  • assets/judge-prompt.md -- Stage 3 adversarial judge prompt.
  • evals/trigger-evals.json -- 20 dispatch queries (10 should, 10 shouldn't).
  • evals/content-evals.json -- seeded-drift recall scenarios.
  • evals/run-evals.sh -- the eval-runner that turns JSON into metrics.

Failure modes guarded against

  • Hallucinated grounding: Stage 2 is deterministic; judge sees only real evidence.
  • Adversarial weakness: Stage 3 prompt defaults to SKEPTICAL.
  • Page-level averaging: claim-level granularity surfaces partials.
  • Bundle leakage: design notes / one-time scripts stay in session state, never in references/.
  • Phantom dependency: SKILL.md links its persona deps via relative paths; A9 PROBE before invoking docs-corpus-audit's substrate.
  • Dispatch collision with sibling skills: trigger-eval validation split is the ship gate (must distinguish from docs-sync / docs-corpus-audit triggers).

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Use this skill to run a four-panel adversarial advisory review on any pull request that touches the OpenAPM specification artifact (docs/src/content/docs/specs/openapm-*.md), its inline / sidecar JSON Schemas (docs/src/content/docs/specs/schemas/*.schema.json), or the conformance fixture seed (tests/fixtures/spec-conformance/**). The panel fans out to four spec-ecosystem reviewers (swagger-openapi-editor, oci-distribution-editor, pkgmgr-registry-contract-editor, w3c-tag-architect), each running in its own agent thread, and a spec-editor synthesizer that produces a fold-now / defer-v0.1.1 / defer-v0.2 / reject list plus a ship decision keyed off a 1..10 shocked_meter scale. The orchestrator is the sole writer to the PR: ONE consolidated comment, no verdict labels, no merge gating. The panel is advisory -- it surfaces findings, prioritizes folds, and renders a ship recommendation that the maintainer weighs.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/apm3,9952026年10月10日 更新

Activate for changes to project positioning, release communication, community-facing artifacts, or breaking-change decisions in microsoft/apm. Triggers on README, MANIFESTO, PRD, CHANGELOG, release workflows, and issue templates.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/apm3,9952026年10月10日 更新

apm-usage

無料

Activate when the user asks about APM (Agent Package Manager): installing, configuring, authoring, or troubleshooting AI-agent packages, dependencies, compilation, MCP servers, policy, or any `apm` CLI command.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/apm3,9952026年10月10日 更新

auth

無料

Activate when code touches token management, credential resolution, git auth flows, GITHUB_APM_PAT, ADO_APM_PAT, AuthResolver, HostInfo, AuthContext, or any remote host authentication -- even if 'auth' isn't mentioned explicitly.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/apm3,9952026年10月10日 更新

Use this skill to post or patch ONE GitHub comment for a microsoft/apm autopilot run. It owns public vs debug body assembly and the AI disclaimer footer. Never labels, assigns, requests reviewers, or merges. Autopilot skills that write comments must load this skill and must not call gh comment APIs themselves.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/apm3,9952026年10月10日 更新

Use this skill to queue maintainer-accepted microsoft/apm issues (`status/accepted`) and fan them out through an isolated pool (default 2) of autopilot-issue-delivery-worker sessions. Any accepted type is eligible. Advisory `triage/recommended` is not authorization. Does not triage. Does not review PRs. Works in a local session, Copilot App automation, Cloud Agent, Remote Agent, or Agentic Workflow.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/apm3,9952026年10月10日 更新

microsoft のスキルをすべて見る

このスキルの問題を報告する