本文へ移動
cccskills
無料GitHub で公開

ai-engineer

Build production-ready LLM applications, advanced RAG systems, and intelligent agents. Implements vector search, multimodal AI, agent orchestration, and enterprise AI integrations. Use PROACTIVELY for LLM features, chatbots, AI agents, or AI-powered applications.

インストール方法を見る

含まれるファイル(11)

  • SKILL.md17.0 KB
  • agents/openai.yaml1.2 KB
  • CHANGELOG.md2.2 KB
  • examples/expected-output.md5.9 KB
  • examples/sample-input.json1.0 KB
  • README.md1.8 KB
  • references/corpus-profile-evaluation.md916 B
  • references/INDEX.md132 B
  • schemas/ai-system-plan.schema.json5.5 KB
  • scripts/ai_system_audit.mjs11.3 KB
  • templates/output-template.md3.1 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

AI Engineer

Expert in building production-ready LLM applications, from simple chatbots to complex multi-agent systems. Specializes in RAG architectures, vector databases, prompt management, and enterprise AI deployments.

Decision Points

RAG Component Selection

Start with a corpus policy, not a vendor preset:
├── Define data class, egress authority, retention, and disclosure filters
├── Select retrieval roles and a versioned profile for each
├── Bind every vector to `spaceId` (model/config/preprocessing/dimensions/metric)
├── Evaluate lexical+dense hybrid retrieval and bounded reranking on held-out tasks
├── Calibrate threshold, top-k, and reranking against cost-quality curves
└── Promote only the profile that meets the stated corpus-specific operating envelope

Model Routing Strategy

Route through a versioned policy rather than query keywords, document count, or provider names. First apply corpus authority, data-class, retention, egress, tool, latency, and cost constraints. Then compare eligible profiles on the target corpus and harness: task outcome, grounding, latency, cost, and failure rate, with uncertainty. Keep the selected model/profile revision and the evaluation receipt with each promotion. A request that does not fit a declared operating envelope should be routed to an approved fallback or explicitly rejected; it must not silently become lexical-only retrieval.

Model-selection mechanics belong to llm-router; this skill specifies the evaluation and authorization inputs that a routing policy must consume.

Agent vs RAG Decision

Task Classification:
├── Static Knowledge Query → Pure RAG
├── Need External APIs → Agent with tools
├── Multi-step Reasoning → Agent with planning
├── Real-time Data Required → Agent with live tools
└── Simple Q&A → RAG with fallback to agent

Failure Modes

Semantic Mismatch Cascade

Symptoms: Good retrieval precision but poor answer relevance, users say "close but not quite right" Detection Rule: A held-out retrieval and answer-quality regression exceeds its pre-registered operating envelope. Root Cause: Query and document embeddings optimized for different semantic spaces Fix: Inspect authority filtering, query/document space identity, corpus coverage, and the calibrated profile before changing a model.

Context Window Overflow

Symptoms: Responses become generic, model ignores specific retrieved context, inconsistent answers Detection Rule: A pre-defined held-out probe shows that additional retrieved context lowers grounded-task quality or raises a measurable irrelevance/error rate. Record the actual probe and operating envelope; “generality score” is not a defined universal metric. Root Cause: Too many irrelevant chunks diluting relevant information Fix: Choose context under a measured token/quality budget; thresholds are corpus-specific estimates, never portable constants.

Tool Hallucination Loop

Symptoms: Agent makes up API calls, references non-existent functions, infinite retry cycles Detection Rule: A versioned tool-contract harness records unsupported calls, schema failures, or non-progressing recovery loops above the task's pre-registered operating envelope. Root Cause: Model trained on different tool schemas than implementation Fix: Add tool validation layer and explicit error handling in agent system prompt

Embedding Drift Degradation

Symptoms: Gradual decline in retrieval quality over time, seasonal performance drops Detection Rule: A versioned holdout shows a decision-relevant regression relative to its recorded baseline and uncertainty interval. Root Cause: Domain language evolves but embedding model remains static Fix: Verify profile identity and re-embed or recalibrate only with a receipted migration.

Response Latency Creep

Symptoms: P95 latency increases gradually, user complaints about slow responses Detection Rule: A monitored latency distribution crosses the corpus policy's calibrated SLO with enough observations to distinguish sustained regression from ordinary variation. Root Cause: Vector index degradation, context size inflation, or model endpoint saturation Fix: Implement index optimization schedule, context pruning, and multi-model load balancing

Worked Examples

Hypothetical example: customer-support retrieval system

The values and components below are placeholders for showing the evaluation sequence. They are not measured results, vendor recommendations, or profile defaults.

Initial requirements: “Answer authorized support questions from a versioned product corpus within a stated latency/cost envelope.”

Step 1: Architecture decision

  • Declare corpus authority, retention, and redaction policy before indexing.
  • Assign lexical, dense, and reranking roles with immutable compatible profile IDs.
  • Select a hybrid candidate set and reranker only after comparing alternatives on a development split; freeze a held-out split for promotion.

Step 2: Implementation walkthrough

const authorized = await authorityFilter(query, corpusPolicy);
const lexical = await lexicalRetriever.search(authorized);
const dense = await denseRetriever.search(authorized, { spaceId });
const candidates = reciprocalRankFuse(lexical, dense);
const reranked = await reranker.rank(authorized, candidates);
const selected = selectByCalibratedPolicy(reranked, evaluationProfile);

if (selected.needsEscalation) {
  return escalateWithEvidenceGap(selected);
}

Step 3: Performance decision

  • Measure each role separately on the frozen corpus/harness: retrieval outcome, grounded answer outcome, latency, and cost.
  • Change one profile or candidate-bound parameter at a time; compare paired outcomes with uncertainty against the declared operating envelope.
  • Promote only when the evidence supports the stated user outcome. Otherwise retain the current profile or narrow the approved task envelope.

Step 4: Evidence-gap handling

  • Detect questions whose authorized corpus lacks the required evidence.
  • Escalate with the missing evidence and query context; do not invent an answer.

Result: a receipted corpus-policy/profile/harness decision, with an explicit escalation path for evidence gaps. No performance result is implied by this example.

Anti-Patterns

Shipping RAG on Vibes

Novice: Ships retrieval after a handful of manual "looks good to me" spot-checks; no held-out evaluation set, no CI gate, no measured recall/precision. Expert: Stands up a repeatable eval harness (unit + retrieval + end-to-end + adversarial) before shipping, and re-runs it on every prompt/retrieval/model change. Detection: ai_system_audit.mjs returns no-eval-harness (critical) when evalHarness.exists is false, and eval-harness-thin (medium) when it exists but lacks end-to-end/adversarial coverage.

Grounding as a Prompt Suggestion, Not a Checked Property

Novice: Asks the model nicely to "cite your sources" and trusts that it will, with no retrieval measurement and no validation that citations actually match retrieved content. Expert: Measures retrieval@k recall/precision against a held-out set (never eyeballs "does this answer look right"), requires citations for every factual claim, and validates output against retrieved sources programmatically instead of trusting the prompt. Detection: ai_system_audit.mjs returns retrieval-never-measured (critical) when retrieval.used is true but recall/precision were never measured (the Semantic Mismatch Cascade failure mode above), no-grounding-requirement (critical) when the system makes factual claims but grounding.citationsRequired is false, and grounding-not-enforced (medium) when citations are required but sourceAttributionEnforced is false.

No Fallback, No Defense, No Ceiling

Novice: Ships an agent that always answers confidently (no low-confidence fallback), accepts raw user text and retrieved documents into the same context with no isolation (no injection defense), and has no per-request cost cap — a single adversarial or pathological request can run away. Expert: Adds a confidence threshold with an explicit fallback action, isolates untrusted content (user input, retrieved docs, tool output) from the system prompt, and enforces a per-request cost ceiling as part of the design — not as an afterthought infra control. Detection: ai_system_audit.mjs returns no-low-confidence-fallback (critical) when lowConfidenceFallback.exists is false, no-injection-defense (critical) when the system accepts untrusted input and promptInjectionDefense.exists is false, no-cost-ceiling (high) when costCeiling.enforced is false, and tool-hallucination-risk (critical) — the Tool Hallucination Loop failure mode above — when tools are used with no toolUse.validationLayer.

Quality Gates

  • Corpus policy filters authority before ranking and records selected profile/spaceId.
  • Held-out retrieval, grounding, latency, and cost outcomes have task-specific acceptance criteria and uncertainty.
  • Hybrid and reranked candidates are compared under the same corpus, model, and budget.
  • Thresholds are calibrated on a development split, then frozen for a held-out check.
  • Illustrative examples are not reported as measured production outcomes.
  • Untrusted retrieved content stays data, with an explicit fallback for insufficient evidence.

See references/corpus-profile-evaluation.md for the required evaluation receipt.

Machine-Checkable Audit

The build-quality subset of the Quality Gates above — the parts a JSON plan can state before a line of code ships — is machine-checkable. scripts/ai_system_audit.mjs exports auditAiSystem(plan), which scores a JSON AI-system plan and flags the failure modes most likely to ship a broken AI feature: no eval harness, unmeasured retrieval, unrequired/unenforced grounding, missing hallucination guardrails, no low-confidence fallback, missing streaming UX on an interactive system, no prompt injection defense on untrusted input, no enforced per-request cost ceiling, and unvalidated tool calls.

This is deliberately scoped to the AI system's own build quality — it does NOT re-check agentic-infrastructure-2026's infra_readiness.mjs gates (framework selection, MCP context overhead, observability wiring, organizational/adoption readiness). A plan can pass this audit and still fail that one (e.g. a well-built RAG pipeline with no chosen framework or kill switch), and vice versa.

  • schemas/ai-system-plan.schema.json — draft-07 shape of the plan the auditor consumes.
  • examples/sample-input.json — a complete plan that scores pass: true.
  • examples/expected-output.md — a "ships on vibes" plan audited, then the same plan fixed and passing.
node scripts/ai_system_audit.mjs --input examples/sample-input.json
# => { "pass": true, "score": 100, "findings": [], "recommendations": [...] }

Not-For Boundaries

Do NOT use this skill for:

Agent Infrastructure/Framework Selection → Use agentic-infrastructure-2026 instead

  • Choosing LangGraph/CrewAI/Semantic Kernel/MCP
  • Observability, evaluation-pipeline tooling, cost governance (kill switches, budget alerts, quotas)
  • Adoption strategy, ROI measurement, pilot scoping

Model-Routing Mechanics → Use llm-router instead

  • Building the routing layer that selects an approved, evaluated profile per request at runtime
  • Cost/latency-tiered dispatch across providers

Agentic App Shape Decisions → Use agentic-app-architecture instead

  • Interaction transparency, execution-substrate/side-effect isolation, overall memory/state shape (this skill builds what runs inside that shape, not the shape itself)

Memory Algorithm Internals → Use episodic-memory-algorithms instead

  • Vector-index internals (HNSW/IVF/PQ), forgetting curves, memory consolidation mechanics

Prompt Engineering Tasks → Use prompt-engineer instead

  • Optimizing prompt templates and instructions
  • A/B testing prompt variations
  • Chain-of-thought prompt design

ML Model Training/Fine-tuning → Out of scope for this skill; use dedicated model-training/fine-tuning tooling

  • Training custom embedding models
  • Fine-tuning LLMs on domain data
  • Model architecture research

Data Pipeline Engineering → Use data-pipeline-engineer instead

  • ETL processes for training data
  • Data validation and cleaning workflows
  • Batch processing systems

Infrastructure/DevOps → Out of scope for this skill; use your stack's infra/DevOps skill (e.g. cloudflare-worker-dev, devops-automator)

  • Kubernetes deployment strategies
  • Database optimization and sharding
  • Load balancer configuration

Analytics and Monitoring Setup → Use chatbot-analytics instead

  • Conversation flow analysis
  • User behavior tracking
  • Performance dashboard creation

Delegate When:

  • Task requires deep ML training/fine-tuning expertise → dedicated model-training tooling
  • Focus is on conversation design → prompt-engineer
  • Need infrastructure scaling → your stack's infra/DevOps skill
  • Want usage analytics → chatbot-analytics
  • Need agent infrastructure/framework/observability decisions → agentic-infrastructure-2026
  • Need model-routing dispatch mechanics → llm-router
  • Need the app's overall shape decided first → agentic-app-architecture
  • Building non-AI features → Relevant specialist skill

References

FileLoad When
templates/output-template.mdDrafting an AI system design and its Roadmap-Item trailer.
schemas/ai-system-plan.schema.jsonValidating an AI-system-plan JSON payload's shape before auditing it.
scripts/ai_system_audit.mjsNeed deterministic scoring of a plan's build-quality readiness.
examples/sample-input.jsonNeed a complete plan that scores pass: true.
examples/expected-output.mdNeed to see a "ships on vibes" system audited, then the same system fixed and passing.
agents/openai.yamlNeed a subagent descriptor for delegated AI-system design/build.
<!-- BEGIN BUNDLE INDEX (auto: index_references.py) -->

Skill Bundle Index

Every file in this skill, and when to open it. Auto-generated; run scripts/index_references.py --fix.

root

  • CHANGELOG.md — AI Engineer — Changelog — - Imported from the global jury_rig skill catalog (ai-engineer, SKILL.md-only) into the repo.
  • README.md — AI Engineer — Build production-ready LLM applications, RAG systems, and intelligent agents: retrieval component selection, model routing strategy, agent-v

agents/

examples/

schemas/

scripts/

templates/

<!-- END BUNDLE INDEX -->

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Expert in 2000s-era music visualization (Milkdrop, AVS, Geiss) and modern WebGL implementations. Specializes in Butterchurn integration, Web Audio API AnalyserNode FFT data, GLSL shaders for audio-reactive visuals, and psychedelic generative art. Activate on "Milkdrop", "music visualization", "WebGL visualizer", "Butterchurn", "audio reactive", "FFT visualization", "spectrum analyzer". NOT for simple bar charts/waveforms (use basic canvas), video editing, or non-audio visuals.

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/port-daddy22026年10月8日 更新

Expert legal research agent for finding and scraping expungement data state by state. Knows authoritative sources, URL patterns, Firecrawl configuration, and 2026 legal landscape.

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/port-daddy22026年10月8日 更新

Expert in 3D computer vision labeling tools, workflows, and AI-assisted annotation for LiDAR, point clouds, and sensor fusion. Covers SAM4D/Point-SAM, human-in-the-loop architectures, and vertical-specific training strategies. Activate on '3D labeling', 'point cloud annotation', 'LiDAR labeling', 'SAM 3D', 'SAM4D', 'sensor fusion annotation', '3D bounding box', 'semantic segmentation point cloud'. NOT for 2D image labeling (use clip-aware-embeddings), general ML training (use ml-engineer), video annotation without 3D (use computer-vision-pipeline), or VLM prompt engineering (use prompt-engineer).

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/port-daddy22026年10月8日 更新

Apply crisis decision-making research to agent routing, uncertainty triage, and coordination failure analysis in time-pressured systems. Use when diagnosing handoff failures, analytical paralysis, or expert judgment under incomplete information. NOT for routine coding, simple CRUD design, or static single-agent tasks with complete information.

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/port-daddy22026年10月8日 更新

Use for insight, reframing, contradiction, impasse, and anomaly-driven problem solving when execution effort no longer helps. NOT for routine optimization, error correction, or well-specified tasks with known solution paths.

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/port-daddy22026年10月8日 更新

Apply cognitive task analysis to expert work that depends on perceptual cues, branching judgment, and recurring monitoring loops. Use when decomposing expert capability into agent structure, simulation design, or validation interviews. NOT for ordinary step-by-step SOP capture or simple pipelines with no tacit cue layer.

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/port-daddy22026年10月8日 更新

curiositech のスキルをすべて見る

このスキルの問題を報告する