本文へ移動
cccskills
無料GitHub で公開

domain-validator

Validate agent output against declared domain rules and ground truth before trusting it downstream. Trigger after any agent produces output that will be used in a decision, stored persistently, or passed to another agent.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md3.6 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Domain Validator

Agent output is a hypothesis. Domain validation is the test.

An agent that produces output without validation is a system that produces hallucinations at scale. The domain validator is the check that separates "the agent said so" from "it is true."

When to use

  • After an agent produces output that feeds a downstream system or human decision
  • When an agent has reasoned over domain-specific data (financial figures, medical records, legal clauses, system configurations)
  • Before persisting agent-generated content to a database or document store
  • When an agent output will be presented to an end user as factual

Procedure

  1. Declare the domain rules — before running any validation, the domain rules must be explicit:

    • What are the invariants? (e.g. "a date range must have start < end", "a price must be positive", "a configuration must reference an existing resource")
    • What are the allowed value ranges or enumerations?
    • What is the ground truth source? (database record, API response, regulatory document, schema definition)
  2. Extract the claims — identify the specific assertions in the agent output that are subject to validation. Not every word in the output is a claim; focus on structured data, named values, and factual assertions.

  3. Validate each claim against the domain rules:

    • Structural validation: does the output conform to the expected schema or format?
    • Range and constraint validation: are values within allowed bounds?
    • Referential integrity: do referenced entities exist in the ground truth source?
    • Logical consistency: are the claims internally consistent? (e.g. no contradictory figures)
    • Freshness: is the ground truth source current, or could it be stale?
  4. Classify findings:

    • PASS: claim is valid against all domain rules
    • WARN: claim is plausible but cannot be fully verified (e.g. ground truth unavailable)
    • FAIL: claim violates a domain rule or contradicts ground truth
  5. Produce a validation report — for each claim: status (PASS/WARN/FAIL), the rule checked, and the evidence.

  6. Gate downstream use — FAIL findings block downstream use of the output. WARN findings require explicit human acknowledgement before proceeding. PASS findings may proceed automatically.

Outputs

  • Validation report: claim | status | rule checked | evidence
  • Overall verdict: PASS / WARN / FAIL
  • List of FAIL and WARN findings for human review

Guardrails

  • Domain rules must be declared before validation runs. Validating against implicit rules produces false confidence.
  • WARN is not PASS. A WARN finding means uncertainty, not safety.
  • Ground truth must be identified. If there is no ground truth source, the output cannot be validated — flag this explicitly rather than assuming it is correct.
  • Validation is not proofreading. Grammar and style are not domain rules. Focus on factual and structural correctness.

Anti-rationalization table

ExcuseCounter
"The model is reliable enough"Reliability is a statistical claim. Domain validation is a deterministic check. Run it.
"We'll catch errors in review"Human review misses structured errors that automated validation catches. Both are needed.
"The domain rules aren't defined yet"Then the output cannot be trusted yet. Define the rules before relying on the output.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Google ADK (Agent Development Kit) orchestration patterns — boundaries, agent composition, and tool seams. Trigger when designing or reviewing multi-agent systems built on ADK. Authoritative source: adk.dev.

日本語の概要は準備中です。原文の説明を表示しています。

jpantsjoha/ai-native-developer-experience122026年10月6日 更新

JP's signature red-team pass — "how would I break this?" Argue against your own approach before proceeding. Trigger on any high-stakes decision, architecture choice, or before marking work complete.

日本語の概要は準備中です。原文の説明を表示しています。

jpantsjoha/ai-native-developer-experience122026年10月6日 更新

Read-only SRE checkup of any GCP project: deterministic probes of the edge, Cloud Run services, 7-day error logs, Cloud Scheduler, alert policies and uptime checks, Secret Manager and IAM, the data stores and the machine's own scheduled jobs, audited into one fixed status table (LIVE / WARNING / RED / INCONCLUSIVE) with evidence, findings by severity, what could not be checked, and a single OVERALL line delivered as one notification. Parametrised by a per-project manifest, so the same routine runs on every project. Use when the operator says "cloud checkup", "SRE check", "is everything live", "what's healthy / warning / red", "any errors this week", "audit the infra", "weekly checkup", "set up the weekly checkup", before a deploy or demo, or after an incident. Cloud Run first; App Engine and GKE differ only in the serving probes.

日本語の概要は準備中です。原文の説明を表示しています。

jpantsjoha/ai-native-developer-experience122026年10月6日 更新

Cloud guardrails for any vendor workload — Google Cloud (GCP, Vertex AI, GKE), AWS (IAM, EKS, Bedrock), Azure (Entra ID, Policy, AKS), Alibaba Cloud (RAM, mainland/international residency). Enforces identity least-privilege, mechanical policy, data boundaries, residency, cost caps, network egress and observability, with official-source validation before any claim. Trigger on any cloud infrastructure design, review, Terraform plan, or LLM/agent deployment; the-architect routes here.

日本語の概要は準備中です。原文の説明を表示しています。

jpantsjoha/ai-native-developer-experience122026年10月6日 更新

LLM and cloud cost awareness — model tiering, token budgets, right-sizing, and when a cheaper model suffices. Trigger before finalising any architecture that calls LLMs, before scaling a workload, or when a cost estimate is needed.

日本語の概要は準備中です。原文の説明を表示しています。

jpantsjoha/ai-native-developer-experience122026年10月6日 更新

Decompose an epic into atomic parallelizable tasks, route each to the right skill, and keep the four delivery records straight — issues, STATUS, ROADMAP, CHANGELOG. Use as a meta-router when several skills could apply, and as the baseline for how delivery state is recorded. Trigger at the start of any multi-track epic, when the skill count exceeds ~12, or when the records have drifted from reality.

日本語の概要は準備中です。原文の説明を表示しています。

jpantsjoha/ai-native-developer-experience122026年10月6日 更新

jpantsjoha のスキルをすべて見る

このスキルの問題を報告する