Review code for security problems; scan for vulnerabilities, secrets, dependency and prompt risks. Use when: asked whether code is safe to ship, even one small handler.
Purpose: Find and report security weaknesses in code, scripts, authorized binaries, and repo-managed prompt surfaces, with honest coverage.
Use this skill for a caller-requested security review of code, a repository scan, authorized binary assurance, dependency risk, secrets, or offline prompt-surface redteam.
Critical Constraints
Scan only repositories, binaries, and prompt surfaces the operator owns or is explicitly authorized to assess. Why: a security review does not grant access to third-party systems or proprietary material.
Keep collection read-only by default; do not exfiltrate secrets, execute destructive payloads, or mutate policy/baselines to manufacture green. Why: the assessment must not become the incident or erase its evidence.
Treat missing/error scanners as a coverage gap, never a clean finding; use --require-tools when complete tool coverage is required. Why: absent evidence is not evidence of absence.
Use the current agent and local shell; do not start another runtime or orchestration substrate unless explicitly requested. Why: repository scanning is a bounded operation, not permission to fan out.
Report findings and coverage gaps, then stop. Remediation, risk acceptance,
reruns, promotion, and any ship or merge call are caller decisions. Name each
finding's remediation class in a few words; do not write the patch, a plan,
an owner, or a priority.
What every review reports
Apply these to every review, scripted or manual. They are the rules most often skipped:
Fail-open paths. For every guard, check, timeout, and exception handler
on the surface, ask what happens when it errors or hangs. A control that
grants access, skips a check, or continues as success on error is a finding
even when its happy path is correct.
Borrowed identity. Trace the effective identity at each hop (user,
service, token, default, hook). A hop where identity is assumed, defaulted,
or inherited instead of verified is the borrowed identity failure mode
and a finding.
Per-class coverage ledger. Walk every applicable class in
the OWASP checklist (the attack pack for
prompt surfaces), plus fail-open and identity, and give each a result:
finding, clean, or not assessed. An unvisited class is a gap, never a clean.
Chasing one lead to the exclusion of the taxonomy is the first-scent
fixation failure mode.
Proven versus suspected. A finding is proven only when you ran a
concrete input, request, or command and observed the behavior; capture it.
A finding reasoned from the code is suspected, even with a candidate input;
give that input and rank it below proven findings.
target: <paths, endpoints, or binary>; authorization: <boundary>
findings: <id> <severity> <file:line> <class>: <what>
proven: <input run> | suspected: <candidate input, why not run>
fix class: <a few words>
coverage: <class> -> finding <ids> | clean | not assessed (<why>)
tools: <scanner or command> -> ran | missing | error
hunt: converged after <n> passes | unconverged | not run
Manual hunt
Code-level review and redteam passes work in any repository, with or without
AgentOps tooling. Walk the ledger against the full surface and probe fail-open
behavior where that is safe. Repeat full passes until one complete pass adds no
new finding and no new coverage gap; that quiet round is the stop condition. If
the budget ends first, report the hunt as unconverged. The quiet-round rule
applies only to the manual hunt.
Scripted scans
Each selected scan runs once per request; a rerun is a new caller decision.
Surface
Entry point
Location
Repository gate (quick or full)
scripts/security-gate.sh
AgentOps repository root only
Composable suite for authorized binaries
skills/security/scripts/security_suite.py
this skill's scripts/
Offline prompt-surface redteam
skills/security/scripts/prompt_redteam.py
this skill's scripts/
No gate script (any other repository): run the scanners the project
already uses, such as a dependency audit, secret scan, or static analyzer,
record each one that is absent as a coverage gap, and do the manual hunt.
Redteam pack: the bundled attack pack
targets AgentOps control surfaces. In another repository its cases fail with
"no files matched target globs"; that is a pack mismatch, not a finding.
Read the suite runbook before binary,
policy, baseline, or redteam work.
This is the canonical security runbook. Suite policy gating produces machine-consumable outputs, including policy/policy-verdict.json when a policy file is supplied.
Add --require-tools when skipped scanners would invalidate the assurance
claim. Checkpoint: preserve the exit code and verify the reported
security-gate-summary.json exists and parses before triage; report the result
as incomplete unless the selected artifact validator and process both succeed.
Scheduled automation runs the full gate against the intended branch and retains its artifact directory. A failing scheduled run creates actionable tracked work; AgentOps itself does not supply the scheduler.
Triage
Open the latest artifact and identify scanner, severity, file, and coverage gaps.
Reproduce the finding with the narrowest safe command; an unreproduced hit stays suspected.
Rank concrete findings and preserve coverage gaps.
Stop. Remediation, risk acceptance, and any later scan are new caller decisions. Do not downgrade, suppress, or update a baseline merely to pass.
Output Specification
Artifact directory: repository gates write ${SECURITY_GATE_OUTPUT_DIR:-${TMPDIR:-/tmp}/agentops-security}/<run-id>/; composable-suite and redteam runs use their explicit --out-dir.
Filename convention: repository gates require security-gate-summary.json (and raw summary.json); suite runs require suite-summary.json; redteam runs require redteam/redteam-results.json.
Serialization/schema format:security-gate-summary.json is JSON with nonempty mode, run_id, output_dir, and gate_status, numeric missing_tool_count, boolean require_tools, and object toolchain.
Validator command: with OUT=<security-gate-run-dir>, run jq -e '(.mode|type)=="string" and (.mode|length)>0 and (.run_id|type)=="string" and (.run_id|length)>0 and (.output_dir|type)=="string" and (.output_dir|length)>0 and .gate_status=="PASS" and (.missing_tool_count|type)=="number" and (.require_tools|type)=="boolean" and (.toolchain|type)=="object"' "$OUT/security-gate-summary.json" >/dev/null.
Output: the review report above; for scripted scans also the artifact
path, command/exit code, mode, and gate status. Do not add an owner, next
action, approval, release, ship, or retry decision.
Quality Checklist
Target and authorization boundary are explicit; collection stayed within them.
Every applicable class has a result; unvisited classes are listed as not assessed.
Scanner availability and skipped/error coverage are visible in the report.
Findings include severity, location, proven-or-suspected evidence, and a remediation class, with no patch, plan, owner, or priority.
Artifacts contain no newly exposed secrets or unredacted sensitive payloads.
The report distinguishes a passing scan from permission to promote, ship, or release.
Suppressions, policy changes, baselines, and risk acceptance require explicit judgment.
The report stops after evidence and contains no continuation decision.
Use AtomLane to compile and execute safe atomic parallel plans on macOS and native Windows Preview for worthwhile independent argv tasks, dependency DAGs, supported platform entrypoints, or Apple-silicon operators. Use at task start or an execution boundary when structured local work may contain two or more worthwhile units; skip plain answers, one quick command, and work whose effects cannot be safely bounded.
Register a deferred decision in the debt registry. Trigger by judgment, not a marker scan, whenever a future reader would ask "why this way?": an unmade decision, stub, loosened type, bypassed check, swallowed error, a default picked "for now", or a TODO/FIXME/HACK/XXX marker. Trigger immediately whenever you defer work, or when the user invokes $add. Over-register freely; the developer drops with "drop A", "drop A,C", or "drop all".