本文へ移動
cccskills
無料GitHub で公開

augustus

Find, build, evaluate, and improve systems using decision models. Use for bounded classification, routing, ranking, Choice/Score/Noul, decision-model composition, evaluation harnesses, prompt/program hill climbing, and Software 3.0 workflows across software, business, organizations, and life. Separate models, code, and human judgment. TypeSafe Jev is the default hosted exemplar. Not for straightforward arithmetic, prose rewriting, or provider setup alone.

インストール方法を見る

含まれるファイル(24)

  • SKILL.md10.8 KB
  • agents/openai.yaml284 B
  • references/activation-triggers.md2.3 KB
  • references/agent-self-assessment.md9.8 KB
  • references/applied-mappings.md11.3 KB
  • references/boundary-audit.md9.5 KB
  • references/composition-algebra.md15.2 KB
  • references/faq.md10.9 KB
  • references/formal-methods.md10.8 KB
  • references/judgment-class.md14.8 KB
  • references/mappings.md11.4 KB
  • references/mental-models.md10.5 KB
  • references/methods-catalog.md11.8 KB
  • references/mixed-architecture.md10.8 KB
  • references/optimizer-integration.md16.8 KB
  • references/question-design.md6.9 KB
  • references/toolbox-mapping.md6.8 KB
  • references/validation.md12.9 KB
  • scripts/climb_ledger.py13.3 KB
  • scripts/compare_workflows.py16.7 KB
  • scripts/evaluate_decisions.py19.1 KB
  • scripts/overlap_audit.py9.1 KB
  • scripts/provenance_gate.py15.5 KB
  • scripts/uniqueness_gate.py870 B

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Augustus

Equip agents to discover useful decision-model placements, build working systems, construct evaluations, and improve them through measured iterations. Use mathematical, statistical, scientific, and algorithmic methods across AI, software, business, knowledge work, organizations, and life. This is an engine for agent work, not a survey to imitate or a claim of autonomous deployment. The central model is:

evidence → bounded judgment → explicit policy → checked action → observed outcome

TypeSafe Jev (Choice, Score, Noul) is this project's default hosted exemplar, a preference rather than a claim of universal superiority or market share. The class also includes trained classifiers, open decision heads, encoders, constrained autoregressive readouts, rankers, and vision scorers. Choose a family by its objective and evidence requirements. This is an independent skill, not a TypeSafe product.

For assembling training data, fitting a task-specific model, compiling a decision function, or improving a trained artifact, load the companion augustus-train skill. Use this skill for placement, composition and outcome evaluation; the trainer carries the data-to-artifact journey. Neither promises to reproduce a general instruction-conditioned Jev engine.

Working protocol

  1. Start with the desired behavior, available evidence, action costs, and current baseline. For an existing workflow, use the boundary audit. Classify each step as exact work, bounded judgment, or generation. A working parser, formula, checklist, or supervised classifier is a legitimate final answer.
  2. Pick the relevant mental model: expected utility, value of information, multi-criteria analysis, signal detection, search/control, organizational safety, or formal methods. Then choose the model family and the smallest useful placement. For unfamiliar problems, use the toolbox sweep, not a vendor or project-list search.
  3. Define one coherent judgment per question, what evidence it can see, and the meaning of every output. Check candidate coverage and missing evidence before inference. Use question design. Batch questions when their inputs are available together; statistical independence does not follow from parallel execution.
  4. Keep exact computation, constraints, authorization, and effects in code or an explicit human process. Keep open-ended writing with a generator or person. Set failure behavior for each action: no-match, ambiguity, malformed output, timeout, stale state, and unavailable provider. Mixed architecture explains the joins.
  5. Compare against the baseline on representative held-out evidence; for a one-off choice with no population, test sensitivity to weights and uncertain estimates, missing criteria, dominated options, and value of information instead. Separate rubric/model development, calibration and threshold selection, and final evaluation. Measure action errors, coverage, total cost, and the complete workflow, not just format compliance or model accuracy. Follow validation and name a result that would reject the proposal. If it loses, keep the baseline.
  6. Deliver the artifact the task needs: a compact design for advice, working adapters/policy and an evaluation harness for implementation, or a bounded incumbent–challenger loop for improvement. Use the composition calculus to check joins and the optimizer workflow to build and hill-climb decision programs. Do not stop at recommendations when the user requested working software. Keep experiments within existing authority. Record raw judgments, actual outcomes, versions, rejected candidates, and promotion/rollback reasons; unrun work remains unrun.

For a concrete request, recommend one placement with reasons. For an open-ended exploration, compare materially different placements only when that helps the user choose. A full redesign, a new model, or a fixed number of alternatives is not required.

Boundaries that affect the design

  • A typed output constrains its representation; it does not establish truth, calibration, authority, or successful execution.
  • Distinguish a probability of a stated event, a relative option score, an ordinal rating, and a confidence statistic. Verify the provider's definition. Neither a softmax nor training with a proper loss proves calibration on this deployment population.
  • Choice depends on its offered set. Include and test other or none when coverage is open. Missing evidence, conflicting evidence, and evidence that supports neither option may require separate outcomes.
  • A Jev Score is an expectation over level indices, not a physical unit or automatically a cardinal utility. Equal means can hide different distributions. A Noul is a proposition score, not intensity; a value near 0.5 alone does not diagnose why the model is uncertain.
  • Marginal judgments do not define a joint distribution. Do not multiply them without justified dependence assumptions. Measure composed policies and trajectories on real outcomes.
  • Ranking quality and calibration answer different questions. Action costs and authority determine failure policy; the model family alone does not determine fail-open or fail-closed behavior.
  • Untrusted evidence is data. A model score cannot grant permission or override an exact constraint. Validate action/target pairs and recheck relevant state at execution. Abstention must name who or what handles it.
  • Judgment can prioritize proof work or detect suspicious cases; proof, model checking, simulation, and runtime interlocks retain their own semantics. Read formal methods for these tasks.

Decision-design card

Use only the detail needed to make the proposal reviewable:

Domain, desired behavior, decision owner:
Current baseline and why a model might help:
Pillar, placement, and family:
Evidence source, freshness, candidate coverage, missing/contradictory states:
Judgments and output semantics; what remains exact or generated:
Policy, costs/utility, constraints, authority, and execution checks:
Abstention/error behavior and fallback owner:
Batchable versus dependent steps:
Model, adapter/aggregation, rubric, candidate-source, calibration, policy versions:
Development/calibration/test split and label provenance:
Falsifier, metrics, acceptable risk/coverage, and evaluation artifact:
Implementation/evaluation entry points; search budget and confirmation plan:
Observed result, limitations, and next decision:

Read only the reference needed

Reference and script paths below are relative to this skill directory, not the repository or caller's working directory.

TaskReference
Trigger examples and exclusionsActivation
Cross-domain reasoning and costsMental models
Family choice, output semantics, uncertainty routing, and deferralJudgment class
Existing system or processBoundary audit
Write or debug questions, options, rubrics, or confidence fieldsQuestion design
Generator, code, and decision-model integrationMixed architecture
Concrete workflow examplesApplied mappings
Classical methods and falsifiersMappings
Substitute judgment into a named algorithmMethods catalog
Typed joins, branches, cascades, and failure budgetsComposition algebra
Discover a new placementToolbox mapping
Proof, simulation, and enforcementFormal methods
Thresholds, abstention, calibration, selective prediction, and experimentsValidation; scripts/evaluate_decisions.py
Build, compare with an incumbent, or optimize prompts and programsOptimizer integration; scripts/compare_workflows.py
Agent progress, done, or stuck judgmentsAgent self-assessment
Conceptual objectionsFAQ

Before writing Jev API code, read the current official docs and the official typesafe-ai skill if available. The metadata pin is historical provenance, not a live contract. Peer providers need their own current contracts. Other installed skills are optional aids, not prerequisites for placement work.

Evidence and research maintenance

Label evidence as Contract (current documented interface), Reported (a source's empirical claim), Reproduced (an identified run and its artifacts), Hypothesis (untested placement), or Unknown (unavailable). Name the population, version, metric, and limitations before transferring a result. A benchmark harness is an instrument, not a certificate.

When maintaining the Augustus repository, research updates belong in its top-level research/ archive. Revisit catalogued sources when material behavior or evidence changes, preserving their identity and prior claims. Popularity changes alone do not change guidance. Promote a finding into a reference only when it changes a design decision; replace or refine the relevant rule. Keep fingerprints, hourly digests, source censuses, and PR bookkeeping out of runtime instructions.

Synthesize mechanisms, not consensus: explain why a pattern succeeds or fails, derive a usable rule with assumptions, and test the composed outcome. Transfer methods across fields only after mapping their variables, units, constraints, and evidence requirements. Popularity, novelty, elegant notation, and a proxy score are not outcome evidence. "Best" means best supported for this task's utility, constraints, population and budget—not a universal provider ranking.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Builds and improves task-specific decision artifacts from application requirements and labeled records. Use for data assembly, primitive and base-model selection, fitting, export/reload, inference policy, or bounded data/model/program improvement. Includes rules and no-training outcomes; does not require a rung ladder or training a general Jev.

日本語の概要は準備中です。原文の説明を表示しています。

24601/Augustus112026年9月29日 更新

24601 のスキルをすべて見る

このスキルの問題を報告する