本文へ移動
cccskills
無料GitHub で公開

maker-mdap

Knowledge base from "Solving a Million-Step LLM Task with Zero Errors" (Meyerson, Paolo, Dailey, Shahrzad, Francon, Hayes, Qiu, Hodjat, Miikkulainen — arXiv:2511.09030). Use when designing long-horizon, multi-step agentic LLM pipelines that need very high reliability, applying Maximal Agentic Decomposition (MAD), first-to-ahead-by-k voting, or red-flagging error correction, estimating cost/reliability scaling laws for multi-agent systems, or studying Massively Decomposed Agentic Processes (MDAPs).

インストール方法を見る

含まれるファイル(12)

  • SKILL.md7.5 KB
  • chapters/ch01-introduction.md3.5 KB
  • chapters/ch02-background.md4.4 KB
  • chapters/ch03-methods.md7.6 KB
  • chapters/ch04-experiments.md5.7 KB
  • chapters/ch05-discussion.md6.3 KB
  • chapters/ch06-appendix-derivations.md2.0 KB
  • chapters/ch07-appendix-prompts-parsers.md8.0 KB
  • cheatsheet.md4.7 KB
  • glossary.md4.4 KB
  • NOTICE.md586 B
  • patterns.md3.5 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

<!-- argument-hint: [topic, framework name, or chapter number] -->

Solving a Million-Step LLM Task with Zero Errors

Authors: Elliot Meyerson, Giuseppe Paolo, Roberto Dailey, Hormoz Shahrzad, Olivier Francon, Conor F. Hayes, Xin Qiu, Babak Hodjat, Risto Miikkulainen (Cognizant AI Lab / UT Austin) Source: arXiv:2511.09030 (Nov 2025) (Notice) | Pages: ~29 | Chapters: 7 | Generated: 2026-08-26

How to Use This Skill

  • Without arguments — load core frameworks below for reference
  • With a topic — ask about voting, red-flagging, cost scaling, decomposition, or another indexed topic; I find and read the relevant chapter
  • With a chapter — ask for ch03; I load that specific chapter
  • Browse — ask "what chapters do you have?" to see the full index

When you ask about a topic not covered in Core Frameworks below, I will read the relevant chapter file before answering.


Core Frameworks & Mental Models

The problem: LLMs have a persistent per-step error rate. A 1%-per-step error rate is expected to fail by step ~100. Standard benchmarks (independent examples, averaged accuracy) hide this — they don't expose what happens once you chain thousands or millions of dependent steps. This paper solves a task requiring over one million dependent LLM steps with zero errors, and argues the fix is architectural, not "wait for a smarter model."

MDAP (Massively Decomposed Agentic Process) — the general paradigm: decompose a task into the smallest possible subtasks, then apply subtask-level error correction. MAKER is this paper's concrete implementation: Maximal Agentic decomposition + first-to-ahead-by-K Error correction + Red-flagging.

1. Maximal Agentic Decomposition (MAD) — set subtask size m=1 (one action per LLM call). Each agent gets only (current state, fixed strategy) — never the accumulated history of past actions. This avoids the reliability degradation single agents suffer as their own context grows, and lets you use small, cheap, non-reasoning models. Use X when Y: use MAD whenever a task decomposes into a chain/recursion of minimal, independently-checkable actions.

2. First-to-ahead-by-k voting — resample a subtask until one candidate has been chosen k more times than every other candidate. Grounded in the Sequential Probability Ratio Test and the gambler's-ruin problem. Critically, k_min (the smallest k hitting a target success probability) grows only Θ(ln s) — logarithmically in total step count s.

3. Red-flagging — discard (don't repair) any response whose structure signals confusion: overlong output, or output that fails a strict format parser. Its main benefit is suppressing correlated errors (independent votes colliding on the same wrong answer), not just raising average per-step success rate.

The scaling-law punchline (the single most important equation): total expected cost is Θ(p^-m · c · s · ln s) — exponential in subtask size m, but only log-linear in total step count s. Corollary: decompose maximally (small m); don't fear long tasks (large s). Bundling steps "for efficiency" is the single biggest cost mistake this framework warns against.

Model/cost selection rule: rank candidate models by c/p (cost per token ÷ per-step success rate), not by raw price or "reasoning" reputation. Calibrate p on a small, cheap, ground-truth-known sample before committing to a large-scale run. In this paper's experiments, small non-reasoning models matched or beat reasoning models on cost-effectiveness.

Verify voting is actually working: don't trust average error rate alone — explicitly count collisions (steps whose first two independent votes are both wrong). A low average error rate can still hide dangerous correlated failures that only collision-counting exposes.

Scope: this paper covers execution (faithfully carrying out a given strategy) — not insight (generating the strategy/plan itself). Know which one your problem needs.


Chapter Index

#TitleKey Frameworks
ch01Introduction — Why Million-Step Tasks Break LLMsMDAP, MAKER, multi-agent advantage
ch02Background — Error Correction and the Hanoi TestbedLbAs, voting/ensembling, decomposition granularity
ch03Methods — MAD, Voting, Red-FlaggingMAD, first-to-ahead-by-k voting, red-flagging, cost/scaling equations
ch04Experiments — The 20-Disk Zero-Error RunCalibrate-then-scale, model/cost selection, convergence behavior
ch05Red-Flagging Impact, Discussion, Future WorkCollision counting, insight-vs-execution, microservices analogy, safety
ch06Appendix A/B — k_min DerivationFull derivation of k_min = Θ(ln s)
ch07Appendix C/D — Prompts, Parsers, SamplesReusable prompt template, repairing vs. red-flagging parsers

Topic Index

  • Calibrate-then-scale workflow → ch04
  • Collision counting / correlated errors → ch05
  • Cost scaling laws (Θ(s ln s), Θ(p^-m)) → ch03, ch04, ch06
  • Decomposition granularity (m) → ch02, ch03
  • First-to-ahead-by-k voting → ch03, ch06
  • Insight vs. execution → ch05
  • k_min derivation → ch03, ch06
  • MAD (Maximal Agentic Decomposition) → ch01, ch03
  • MAKER (the implementation) → ch01, ch03, ch04
  • MDAP (the general framework) → ch01
  • Microservices analogy → ch05
  • Model selection / cost projection → ch04
  • Multi-agent advantage → ch01
  • Parsers (repairing vs. red-flagging) → ch03, ch07
  • Prompt template (Towers of Hanoi agent) → ch07
  • Red-flagging → ch03, ch05, ch07
  • Safety implications of decomposition → ch05
  • Towers of Hanoi benchmark → ch02, ch04
  • Voting theory (SPRT, gambler's ruin) → ch02, ch03

Supporting Files

  • glossary.md — all key terms with definitions
  • patterns.md — MAD, voting, red-flagging, calibrate-then-scale, collision counting
  • cheatsheet.md — decision rules, thresholds, model-selection matrix

Scope & Limits

This skill covers this single paper's content only (arXiv:2511.09030, v1, Nov 2025). It addresses execution of long-horizon agentic tasks, not open-ended plan/strategy generation ("insight"). Numeric thresholds cited (e.g. the ~700-token red-flagging cutoff) are specific to this paper's Towers-of-Hanoi setup and prompt design — re-derive them empirically for any other domain rather than reusing them as universal constants. For applying these ideas to your own codebase or agent framework, combine with project-specific tools; for topics beyond this paper, ask directly.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Train and optimize AI agents using Microsoft's Agent Lightning framework with reinforcement learning. Use when setting up agent training, instrumenting agents with tracing, configuring LightningStore, implementing reward functions, or optimizing prompts with RL/APO algorithms.

日本語の概要は準備中です。原文の説明を表示しています。

coco-research/coco5462026年10月11日 更新

Post-run self-evaluation system that scores agent output on correctness, clarity, actionability, and conciseness. Use after /team runs, skill executions, or when explicitly asked to evaluate output quality.

日本語の概要は準備中です。原文の説明を表示しています。

coco-research/coco5462026年10月11日 更新

Create AI marketing videos for ads, promos, product launches, and brand content. Models: Veo, Seedance, Wan, FLUX for visuals, Kokoro for voiceover. Types: product demos, testimonials, explainers, social ads, brand videos. Use for: Facebook ads, YouTube ads, product launches, brand awareness. Triggers: marketing video, ad video, promo video, commercial, brand video, product video, explainer video, ad creative, video ad, facebook ad video, youtube ad, instagram ad, tiktok ad, promotional video, launch video

日本語の概要は準備中です。原文の説明を表示しています。

coco-research/coco5462026年10月11日 更新

Use when building AI features into a product: LLM integration, RAG pipelines, guardrails, streaming, AI UX, prompt engineering, or AI cost control. Treats prompts as code and validates every model output.

日本語の概要は準備中です。原文の説明を表示しています。

coco-research/coco5462026年10月11日 更新

Your AI research and engineering brain trust. 59 named personas across 8 cells covering frontier labs, applied product, model architecture, reasoning/RL/agents, alignment and interpretability, theory and science of DL, multimodal and…

日本語の概要は準備中です。原文の説明を表示しています。

coco-research/coco5462026年10月11日 更新

Use when designing a new REST or GraphQL API, reviewing an API spec before implementation, setting team API standards, or migrating REST to GraphQL. Covers resources, HTTP semantics, pagination, error handling, and pitfalls.

日本語の概要は準備中です。原文の説明を表示しています。

coco-research/coco5462026年10月11日 更新

coco-research のスキルをすべて見る

このスキルの問題を報告する