本文へ移動
cccskills
無料GitHub で公開

cost-guardrail

LLM and cloud cost awareness — model tiering, token budgets, right-sizing, and when a cheaper model suffices. Trigger before finalising any architecture that calls LLMs, before scaling a workload, or when a cost estimate is needed.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md4.4 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Cost Guardrail

The most expensive model is the one running on every request when it does not need to.

LLM cost is not a finance problem — it is an architecture problem. The design determines the bill. This skill enforces cost-awareness as a first-class design constraint, not an afterthought.

When to use

  • Designing any system that calls an LLM (directly or via an agent)
  • Before scaling a workload to higher volumes
  • When a cost estimate is required for a feature or release
  • When reviewing an architecture for unbounded cost vectors
  • When choosing between model tiers for a given task

Procedure

  1. Identify every LLM call in the system — list: which agent or component makes the call, the model tier used, the approximate input and output token counts, and the call frequency (per user action / per minute / per batch).

  2. Apply the model tiering test — for each LLM call, ask:

    • Does this task require deep reasoning, or is it classification / extraction / reformatting?
    • Can the task be completed with a smaller or faster model?
    • Is the model tier choice based on evidence (benchmark, A/B test) or assumption?

    General tiering principle (verify current pricing against your provider's documentation before relying on it):

    Task typeAppropriate tier
    Simple classification, extraction, summarisationSmall / fast model
    Complex reasoning, multi-step planning, code generationMid-tier model
    Deep analysis, architecture decisions, adversarial reviewHighest-tier model
  3. Identify unbounded cost vectors — flag any call pattern where the token count or call volume has no upper bound:

    • Loops that call an LLM until a condition is met (with no max-iteration guard)
    • User-triggered calls with no rate limiting
    • Context windows that grow unboundedly across a conversation
    • Batch jobs with no per-run budget ceiling
  4. Estimate the monthly cost envelope — for each LLM call:

    estimated monthly cost ≈ (input tokens × input price) + (output tokens × output price) × calls/month
    

    Use current published rates from your provider. Do not use rates from training data — they change.

  5. Add cost controls — for each unbounded vector:

    • Set a max-token budget per call (trim context if needed)
    • Add rate limiting at the application layer
    • Add a budget alert at the infrastructure layer
    • Consider caching repeated calls with identical or near-identical inputs
  6. Check for caching opportunities — LLM calls that return the same result for the same input are cacheable. Prompt caching (where supported by the provider) can reduce cost significantly on repeated prefixes.

  7. Document the cost model — in the ADR or design doc, record: model tiers chosen, rationale, estimated monthly cost at target scale, and the controls in place.

Outputs

  • LLM call inventory: component | model tier | input tokens (est.) | output tokens (est.) | frequency | monthly cost (est.)
  • Unbounded cost vectors flagged with mitigations
  • Monthly cost estimate at target scale
  • Recommended model tier per call with rationale

Guardrails

  • Never use pricing from training data. Rates change. Fetch current rates from the provider's documentation before estimating.
  • A call that "works" at low volume may be unaffordable at scale. Always estimate at the target scale, not the current scale.
  • Caching is not optional for high-frequency repeated calls. An uncached LLM call repeated thousands of times per day is a design flaw.
  • Token budgets are architecture decisions. Decide them explicitly; do not let the model decide by consuming whatever context is available.

Anti-rationalization table

ExcuseCounter
"It's only a few cents per call"At scale, cents become thousands of dollars. Estimate the monthly envelope.
"We'll optimise later"Cost optimisation is hardest after the architecture is set. Do it now.
"The big model gives better results"Verify with a test. Small models are often sufficient for structured tasks.
"We don't know the volume yet"Estimate a range. A 10x cost swing between low and high volume is a design risk.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Google ADK (Agent Development Kit) orchestration patterns — boundaries, agent composition, and tool seams. Trigger when designing or reviewing multi-agent systems built on ADK. Authoritative source: adk.dev.

日本語の概要は準備中です。原文の説明を表示しています。

jpantsjoha/ai-native-developer-experience122026年10月6日 更新

JP's signature red-team pass — "how would I break this?" Argue against your own approach before proceeding. Trigger on any high-stakes decision, architecture choice, or before marking work complete.

日本語の概要は準備中です。原文の説明を表示しています。

jpantsjoha/ai-native-developer-experience122026年10月6日 更新

Read-only SRE checkup of any GCP project: deterministic probes of the edge, Cloud Run services, 7-day error logs, Cloud Scheduler, alert policies and uptime checks, Secret Manager and IAM, the data stores and the machine's own scheduled jobs, audited into one fixed status table (LIVE / WARNING / RED / INCONCLUSIVE) with evidence, findings by severity, what could not be checked, and a single OVERALL line delivered as one notification. Parametrised by a per-project manifest, so the same routine runs on every project. Use when the operator says "cloud checkup", "SRE check", "is everything live", "what's healthy / warning / red", "any errors this week", "audit the infra", "weekly checkup", "set up the weekly checkup", before a deploy or demo, or after an incident. Cloud Run first; App Engine and GKE differ only in the serving probes.

日本語の概要は準備中です。原文の説明を表示しています。

jpantsjoha/ai-native-developer-experience122026年10月6日 更新

Cloud guardrails for any vendor workload — Google Cloud (GCP, Vertex AI, GKE), AWS (IAM, EKS, Bedrock), Azure (Entra ID, Policy, AKS), Alibaba Cloud (RAM, mainland/international residency). Enforces identity least-privilege, mechanical policy, data boundaries, residency, cost caps, network egress and observability, with official-source validation before any claim. Trigger on any cloud infrastructure design, review, Terraform plan, or LLM/agent deployment; the-architect routes here.

日本語の概要は準備中です。原文の説明を表示しています。

jpantsjoha/ai-native-developer-experience122026年10月6日 更新

Decompose an epic into atomic parallelizable tasks, route each to the right skill, and keep the four delivery records straight — issues, STATUS, ROADMAP, CHANGELOG. Use as a meta-router when several skills could apply, and as the baseline for how delivery state is recorded. Trigger at the start of any multi-track epic, when the skill count exceeds ~12, or when the records have drifted from reality.

日本語の概要は準備中です。原文の説明を表示しています。

jpantsjoha/ai-native-developer-experience122026年10月6日 更新

Validate agent output against declared domain rules and ground truth before trusting it downstream. Trigger after any agent produces output that will be used in a decision, stored persistently, or passed to another agent.

日本語の概要は準備中です。原文の説明を表示しています。

jpantsjoha/ai-native-developer-experience122026年10月6日 更新

jpantsjoha のスキルをすべて見る

このスキルの問題を報告する