本文へ移動
cccskills
無料GitHub で公開

security

Review code for security problems; scan for vulnerabilities, secrets, dependency and prompt risks. Use when: asked whether code is safe to ship, even one small handler.

インストール方法を見る

含まれるファイル(10)

  • SKILL.md10.1 KB
  • references/agentops-redteam-pack.json5.7 KB
  • references/owasp-checklist.md4.0 KB
  • references/policy-example.json534 B
  • references/security-suite-runbook.md3.7 KB
  • references/security-suite.feature1.4 KB
  • references/security.feature1.4 KB
  • scripts/prompt_redteam.py11.3 KB
  • scripts/security_suite.py31.3 KB
  • scripts/validate.sh4.0 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Security Skill

Purpose: Find and report security weaknesses in code, scripts, authorized binaries, and repo-managed prompt surfaces, with honest coverage.

Use this skill for a caller-requested security review of code, a repository scan, authorized binary assurance, dependency risk, secrets, or offline prompt-surface redteam.

Critical Constraints

  • Scan only repositories, binaries, and prompt surfaces the operator owns or is explicitly authorized to assess. Why: a security review does not grant access to third-party systems or proprietary material.
  • Keep collection read-only by default; do not exfiltrate secrets, execute destructive payloads, or mutate policy/baselines to manufacture green. Why: the assessment must not become the incident or erase its evidence.
  • Treat missing/error scanners as a coverage gap, never a clean finding; use --require-tools when complete tool coverage is required. Why: absent evidence is not evidence of absence.
  • Use the current agent and local shell; do not start another runtime or orchestration substrate unless explicitly requested. Why: repository scanning is a bounded operation, not permission to fan out.
  • Report findings and coverage gaps, then stop. Remediation, risk acceptance, reruns, promotion, and any ship or merge call are caller decisions. Name each finding's remediation class in a few words; do not write the patch, a plan, an owner, or a priority.

What every review reports

Apply these to every review, scripted or manual. They are the rules most often skipped:

  1. Fail-open paths. For every guard, check, timeout, and exception handler on the surface, ask what happens when it errors or hangs. A control that grants access, skips a check, or continues as success on error is a finding even when its happy path is correct.
  2. Borrowed identity. Trace the effective identity at each hop (user, service, token, default, hook). A hop where identity is assumed, defaulted, or inherited instead of verified is the borrowed identity failure mode and a finding.
  3. Per-class coverage ledger. Walk every applicable class in the OWASP checklist (the attack pack for prompt surfaces), plus fail-open and identity, and give each a result: finding, clean, or not assessed. An unvisited class is a gap, never a clean. Chasing one lead to the exclusion of the taxonomy is the first-scent fixation failure mode.
  4. Proven versus suspected. A finding is proven only when you ran a concrete input, request, or command and observed the behavior; capture it. A finding reasoned from the code is suspected, even with a candidate input; give that input and rank it below proven findings.
target:   <paths, endpoints, or binary>; authorization: <boundary>
findings: <id> <severity> <file:line> <class>: <what>
          proven: <input run> | suspected: <candidate input, why not run>
          fix class: <a few words>
coverage: <class> -> finding <ids> | clean | not assessed (<why>)
tools:    <scanner or command> -> ran | missing | error
hunt:     converged after <n> passes | unconverged | not run

Manual hunt

Code-level review and redteam passes work in any repository, with or without AgentOps tooling. Walk the ledger against the full surface and probe fail-open behavior where that is safe. Repeat full passes until one complete pass adds no new finding and no new coverage gap; that quiet round is the stop condition. If the budget ends first, report the hunt as unconverged. The quiet-round rule applies only to the manual hunt.

Scripted scans

Each selected scan runs once per request; a rerun is a new caller decision.

SurfaceEntry pointLocation
Repository gate (quick or full)scripts/security-gate.shAgentOps repository root only
Composable suite for authorized binariesskills/security/scripts/security_suite.pythis skill's scripts/
Offline prompt-surface redteamskills/security/scripts/prompt_redteam.pythis skill's scripts/
  • No gate script (any other repository): run the scanners the project already uses, such as a dependency audit, secret scan, or static analyzer, record each one that is absent as a coverage gap, and do the manual hunt.
  • Redteam pack: the bundled attack pack targets AgentOps control surfaces. In another repository its cases fail with "no files matched target globs"; that is a pack mismatch, not a finding.
  • Read the suite runbook before binary, policy, baseline, or redteam work.

This is the canonical security runbook. Suite policy gating produces machine-consumable outputs, including policy/policy-verdict.json when a policy file is supplied.

Repository gate

scripts/security-gate.sh --mode quick   # changed scope
scripts/security-gate.sh --mode full    # repository-wide

Add --require-tools when skipped scanners would invalidate the assurance claim. Checkpoint: preserve the exit code and verify the reported security-gate-summary.json exists and parses before triage; report the result as incomplete unless the selected artifact validator and process both succeed.

Scheduled automation runs the full gate against the intended branch and retains its artifact directory. A failing scheduled run creates actionable tracked work; AgentOps itself does not supply the scheduler.

Triage

  1. Open the latest artifact and identify scanner, severity, file, and coverage gaps.
  2. Reproduce the finding with the narrowest safe command; an unreproduced hit stays suspected.
  3. Rank concrete findings and preserve coverage gaps.
  4. Stop. Remediation, risk acceptance, and any later scan are new caller decisions. Do not downgrade, suppress, or update a baseline merely to pass.

Output Specification

Artifact directory: repository gates write ${SECURITY_GATE_OUTPUT_DIR:-${TMPDIR:-/tmp}/agentops-security}/<run-id>/; composable-suite and redteam runs use their explicit --out-dir.

Filename convention: repository gates require security-gate-summary.json (and raw summary.json); suite runs require suite-summary.json; redteam runs require redteam/redteam-results.json.

Serialization/schema format: security-gate-summary.json is JSON with nonempty mode, run_id, output_dir, and gate_status, numeric missing_tool_count, boolean require_tools, and object toolchain.

Validator command: with OUT=<security-gate-run-dir>, run jq -e '(.mode|type)=="string" and (.mode|length)>0 and (.run_id|type)=="string" and (.run_id|length)>0 and (.output_dir|type)=="string" and (.output_dir|length)>0 and .gate_status=="PASS" and (.missing_tool_count|type)=="number" and (.require_tools|type)=="boolean" and (.toolchain|type)=="object"' "$OUT/security-gate-summary.json" >/dev/null.

Output: the review report above; for scripted scans also the artifact path, command/exit code, mode, and gate status. Do not add an owner, next action, approval, release, ship, or retry decision.

Quality Checklist

  • Target and authorization boundary are explicit; collection stayed within them.
  • Every applicable class has a result; unvisited classes are listed as not assessed.
  • Scanner availability and skipped/error coverage are visible in the report.
  • Findings include severity, location, proven-or-suspected evidence, and a remediation class, with no patch, plan, owner, or priority.
  • Artifacts contain no newly exposed secrets or unredacted sensitive payloads.
  • The report distinguishes a passing scan from permission to promote, ship, or release.
  • Suppressions, policy changes, baselines, and risk acceptance require explicit judgment.
  • The report stops after evidence and contains no continuation decision.

Validation

Run the skill and redteam validators:

bash skills/security/scripts/validate.sh
bash tests/scripts/test-security-suite-redteam.sh

For a bounded suite smoke test, use an owned binary and a temporary output directory as shown in the suite runbook.

Troubleshooting

ProblemResponse
Scanner missing/errorRecord the coverage gap; install it or rerun with --require-tools when required
Local/CI mismatchCompare scanner versions, config, mode, and both artifact directories
Suspected false positiveReproduce narrowly; document any authorized suppression and its owner
Suite/baseline failureInspect the named compare/policy artifact; never refresh baseline reflexively
Redteam failure after wording changeDecide whether the control regressed or the attack-pack matcher needs intentional revision

Reference Documents

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Use AtomLane to compile and execute safe atomic parallel plans on macOS and native Windows Preview for worthwhile independent argv tasks, dependency DAGs, supported platform entrypoints, or Apple-silicon operators. Use at task start or an execution boundary when structured local work may contain two or more worthwhile units; skip plain answers, one quick command, and work whose effects cannot be safely bounded.

日本語の概要は準備中です。原文の説明を表示しています。

hashgraph-online/awesome-codex-plugins1,2752026年10月11日 更新

add

無料

Register a deferred decision in the debt registry. Trigger by judgment, not a marker scan, whenever a future reader would ask "why this way?": an unmade decision, stub, loosened type, bypassed check, swallowed error, a default picked "for now", or a TODO/FIXME/HACK/XXX marker. Trigger immediately whenever you defer work, or when the user invokes $add. Over-register freely; the developer drops with "drop A", "drop A,C", or "drop all".

日本語の概要は準備中です。原文の説明を表示しています。

hashgraph-online/awesome-codex-plugins1,2752026年10月11日 更新

ADK 框架适配层。为 LangChain / EINO / AutoGen / AgentScope / CrewAI 提供框架特定的 代码模板、惯用模式、API 映射和项目结构,供 agent-dev-workshop Phase 5 代码生成使用。 每个框架 reference 文件标注 verified_date 用于版本锁定。

日本語の概要は準備中です。原文の説明を表示しています。

hashgraph-online/awesome-codex-plugins1,2752026年10月11日 更新

中文调试修复技能。用于报错、测试失败、页面异常、功能不符合预期、需要定位根因并做最小修复时。触发语包括"进入调试模式""帮我修问题""报错了""测试失败""页面坏了""找根因"。

日本語の概要は準備中です。原文の説明を表示しています。

hashgraph-online/awesome-codex-plugins1,2752026年10月11日 更新

交互式 AI Agent 开发工作坊:通过 6 阶段深度协作对话,引导用户完成 Agent 需求分析、架构设计、 工具定义、Prompt 与编排设计、代码生成、验证迭代,产出可直接运行的 Agent 项目。 框架无关设计优先,支持 LangChain / EINO / AutoGen / AgentScope / CrewAI 等 ADK 框架。

日本語の概要は準備中です。原文の説明を表示しています。

hashgraph-online/awesome-codex-plugins1,2752026年10月11日 更新

中文漂移审计技能。用于项目或学习过程变乱、上下文漂移、任务分叉、多个方案冲突、命名不一致、Codex 可能顺手改多了时。触发语包括"漂移检查""感觉跑偏了""项目变乱了""检查是否失控""分叉太多""上下文漂移"。

日本語の概要は準備中です。原文の説明を表示しています。

hashgraph-online/awesome-codex-plugins1,2752026年10月11日 更新

hashgraph-online のスキルをすべて見る

このスキルの問題を報告する