Forward Deployed Engineer delivery contract for AWS engagements, build-first artifacts, grounded cost estimates, Well-Architected review, evolution roadmap
日本語の概要は準備中です。原文の説明を表示しています。
Superforecasting with calibrated reasoning, Brier score tracking, and prediction ledger management
インストールする前に、エージェントに与えられる指示の中身を確認できます。
Make specific, falsifiable predictions with calibrated confidence levels. Track accuracy over time using Brier scores. Apply superforecasting methodology (Tetlock/Good Judgment Project) to any domain — technology trends, project outcomes, market shifts, competitive moves, risk assessment.
| Signal Type | Weight | Description |
|---|---|---|
| Leading indicator | High | Predicts before the event (e.g., job postings predict growth) |
| Lagging indicator | Medium | Confirms after the event (e.g., quarterly earnings) |
| Base rate | High | Historical frequency of similar events |
| Expert opinion | Medium | Domain expert assessment (weight by track record) |
| Data point | High | Quantitative measurement directly relevant |
| Anomaly | High | Deviation from expected pattern — investigate |
| Structural change | Very High | Rules of the game changing (regulation, technology shift) |
| Sentiment shift | Medium | Public/market mood change (often noise, sometimes signal) |
Signal strength:
| Probability | Meaning | Betting Odds |
|---|---|---|
| 5% | Almost certainly not | 19:1 against |
| 15% | Very unlikely | ~6:1 against |
| 25% | Unlikely but plausible | 3:1 against |
| 35% | Somewhat unlikely | ~2:1 against |
| 45% | Toss-up, leaning no | ~1.2:1 against |
| 55% | Toss-up, leaning yes | ~1.2:1 for |
| 65% | Somewhat likely | ~2:1 for |
| 75% | Likely | 3:1 for |
| 85% | Very likely | ~6:1 for |
| 95% | Almost certain | 19:1 for |
Adjustment rules: +/-5-15% per strong signal, +/-2-5% per moderate signal. If gut says 80% but analysis says 55%, trust the analysis.
Before finalizing ANY prediction, check against these 8 biases:
| Bias | Check | Fix |
|---|---|---|
| Anchoring | Am I stuck on the first number I thought of? | Re-derive from base rates |
| Availability | Am I overweighting recent/vivid examples? | Search for boring counterexamples |
| Confirmation | Am I only finding evidence that agrees? | Explicitly search for disconfirming evidence |
| Narrative | Am I constructing a compelling story that feels true? | Check: does the data support this without the story? |
| Overconfidence | Am I more certain than my evidence warrants? | Would I bet real money at these odds? |
| Scope insensitivity | Am I treating "some" and "a lot" as the same? | Quantify: how much exactly? |
| Recency | Am I overweighting what happened last? | Check 5-year and 10-year base rates |
| Status quo | Am I assuming things will stay the same? | What would need to change, and how likely is each change? |
For each prediction, construct:
Brier = (predicted_probability - actual_outcome)^2
Where actual_outcome is 0 (didn't happen) or 1 (happened).
| Score | Quality |
|---|---|
| < 0.10 | Excellent |
| 0.10 - 0.15 | Good |
| 0.15 - 0.25 | Average |
| 0.25 | Coin flip (no skill) |
| > 0.30 | Worse than guessing |
Track cumulative Brier score across all resolved predictions. Review monthly. If cumulative Brier > 0.25, recalibrate methodology.
When explicitly requested or when consensus confidence exceeds 85%:
| Domain | Priority Sources |
|---|---|
| Technology | GitHub trending, HN, arXiv, Crunchbase, job postings, patent filings |
| Finance | FRED, SEC filings, central bank statements, VIX, yield curves |
| Geopolitics | UN resolutions, RAND, think tank reports, diplomatic cables |
| Climate/Energy | IPCC, IEA, CDP, BloombergNEF, utility filings |
| AI/ML | arXiv, model benchmarks, API pricing trends, conference papers |
prediction_id: <PRED-YYYY-MM-DD-NNN>
created: <YYYY-MM-DD>
domain: <technology | finance | geopolitics | climate | ai_ml | general>
time_horizon: <1_week | 1_month | 3_months | 1_year>
prediction: <specific, falsifiable statement>
confidence: <probability 0.05-0.95>
reasoning_chain:
reference_class:
base_rate: <probability>
analogues:
- <historical analogue and outcome>
specific_evidence:
- signal: <description>
type: <leading | lagging | base_rate | expert | data | anomaly | structural | sentiment>
strength: <strong | moderate | weak>
adjustment: <+/- percentage>
synthesis: <narrative combining outside and inside views>
key_assumptions:
- assumption: <what must hold>
if_violated: <probability shift>
resolution:
date: <YYYY-MM-DD>
criteria: <exact observable condition>
data_source: <where to verify>
bias_check: <which biases were checked and adjustments made>
status: active | resolved | expired
updates:
- date: <YYYY-MM-DD>
old_confidence: <previous>
new_confidence: <updated>
reason: <what changed>
resolution_result:
date: <YYYY-MM-DD>
outcome: true | false
evidence: <what happened>
brier_score: <calculated score>
lesson: <what to learn from this>
src/genesis/learning/ — Outcome tracking for Brier score integrationsrc/genesis/identity/REFLECTION_STRATEGIC.md — Strategic reflection contextまだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Forward Deployed Engineer delivery contract for AWS engagements, build-first artifacts, grounded cost estimates, Well-Architected review, evolution roadmap
日本語の概要は準備中です。原文の説明を表示しています。
Canonical guide to Genesis browser automation - layers (Camoufox, Chromium, the user's Chrome over CDP, TinyFish, desktop), per-tool timeouts, safety gates, verify-after-act, what a click checks (scroll, hit test, covered targets), overlays, iframes, tabs, and failure diagnosis
日本語の概要は準備中です。原文の説明を表示しています。
Update Claude Code (the CC CLI / "clog code") to a new version, or bump the pinned CC version. Use when the user asks to update Claude Code, bump the CC pin, evaluate a new CC release, or says "clog code update". Routes to the canonical, standardized process in docs/reference/cc-compatibility.md — do NOT re-derive the update mechanism by grepping every time. Do NOT use for general "what changed in CC" trivia with no intent to update.
日本語の概要は準備中です。原文の説明を表示しています。
This skill should be used when a session's job is to DRIVE OPEN PRs TO MERGE rather than to write new code — "close out the open PRs", "review and fix the open PRs", "what's blocking our PRs", "which PRs are mergeable". It owns the In Review column: it reads each PR's gate status, verifies and fixes review findings on PRs OTHER sessions built, replies in-thread, and stops at the merge gate for the user's per-PR approval. Do NOT load it for building a feature and opening its PR — that is a build session (`genesis-development`).
日本語の概要は準備中です。原文の説明を表示しています。
Code understanding tool selection. Use when exploring architecture, finding definitions, tracing call chains, assessing blast radius of changes, or debugging code paths in the Genesis codebase.
日本語の概要は準備中です。原文の説明を表示しています。
End-to-end content creation and publishing. Takes a topic (or generates one), drafts in the user's voice, gets approval via Telegram, and publishes to Medium via browser automation. Invoke with "publish a post about X", "write and publish to Medium", "content-publish", or when an ego-dispatched session needs to create and distribute content.
日本語の概要は準備中です。原文の説明を表示しています。