本文へ移動
cccskills
無料GitHub で公開

test-time-compute-optimizer

Expert guide for scaling test-time compute, dynamic reasoning token allocation (Gemini 4 Pro, Claude 5.5, GPT Astra 6), Monte Carlo Tree Search (MCTS), and Process Reward Models / Panduan ahli penskalaan komputasi waktu uji (test-time compute), alokasi dinamis budget penalaran, MCTS, dan Process Reward Models.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md11.8 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

test-time-compute-optimizer — vibes-plug Skill

English | Bahasa Indonesia


<a name="english"></a>

English

Orchestration & Integration

Connects and orchestrates with brainstorming, zero-to-prod-orchestrator, ai-llm-integration-expert, frontier-ai-models-expert, llm-finops-router, and speculative-multi-draft-synthesizer to dynamically balance reasoning depth, inference latency, and token expenditures across frontier models.

Description

Production guide for scaling test-time compute (inference-time reasoning) across frontier AI models including Google Gemini 4 Pro, Anthropic Claude 5.5 (Opus/Sonnet), and OpenAI GPT Astra 6 (o6). Provides concrete architectures for dynamic thinking budget allocation, tree-search exploration (MCTS, beam search), step-level Process Reward Model (PRM) verification, and self-consistency consensus loops.

Trigger Conditions

Activate this skill when:

  • Designing systems that tackle hard algorithmic, formal verification, or multi-step architectural problems.
  • Configuring model reasoning parameters (thinkingBudget, budget_tokens, reasoning_effort).
  • Implementing inference-time search algorithms (Monte Carlo Tree Search, Best-of-N with verification).
  • Optimizing cost-vs-accuracy trade-offs to prevent over-spending on trivial tasks while maximizing reasoning on critical bottlenecks.

Core Concepts & Patterns

1. Dynamic Reasoning Budget Matrix

Inference-time compute scaling operates across three distinct operational tiers based on prompt entropy and execution risk:

Complexity TierTarget Problem TypesGemini 4 Pro / 3.x (thinkingBudget)Claude 5.5 (budget_tokens)GPT Astra 6 (reasoning_effort)
Tier 1: Deterministic / CRUDSimple UI components, DTO mappings, regex conversions0 (Thinking disabled) or 10241024low
Tier 2: Business Logic & APIsMulti-table DB transactions, authentication flows, REST schemas4096 - 81924096 - 8192medium
Tier 3: Mission-Critical HardFinancial state machines, distributed consensus, cryptographic proofs, complex refactoring16384 - 3276816384 - 32768high

2. Entropy-Based Compute Dispatcher (TypeScript Implementation)

import { createGoogleGenerativeAI } from '@ai-sdk/google';
import { createAnthropic } from '@ai-sdk/anthropic';
import { generateText } from 'ai';

export type TaskComplexity = 'trivial' | 'standard' | 'complex' | 'extreme';

export interface ComputeAllocation {
  geminiThinkingBudget: number;
  claudeBudgetTokens: number;
  reasoningEffort: 'low' | 'medium' | 'high';
}

export function resolveComputeBudget(complexity: TaskComplexity): ComputeAllocation {
  switch (complexity) {
    case 'trivial':
      return { geminiThinkingBudget: 0, claudeBudgetTokens: 1024, reasoningEffort: 'low' };
    case 'standard':
      return { geminiThinkingBudget: 4096, claudeBudgetTokens: 4096, reasoningEffort: 'medium' };
    case 'complex':
      return { geminiThinkingBudget: 16384, claudeBudgetTokens: 16384, reasoningEffort: 'high' };
    case 'extreme':
      return { geminiThinkingBudget: 32768, claudeBudgetTokens: 32768, reasoningEffort: 'high' };
  }
}

export async function executeOptimizedReasoning(
  prompt: string,
  complexity: TaskComplexity,
  provider: 'gemini' | 'claude'
): Promise<string> {
  const budget = resolveComputeBudget(complexity);

  if (provider === 'gemini') {
    const google = createGoogleGenerativeAI({
      apiKey: process.env.GEMINI_API_KEY || ''
    });
    
    const result = await generateText({
      model: google('gemini-4-pro', {
        useSearchGrounding: false
      }),
      prompt,
      providerOptions: {
        google: {
          thinkingBudget: budget.geminiThinkingBudget
        }
      }
    });
    return result.text;
  }

  const anthropic = createAnthropic({
    apiKey: process.env.ANTHROPIC_API_KEY || ''
  });

  const result = await generateText({
    model: anthropic('claude-3-7-sonnet-20250219'),
    prompt,
    providerOptions: {
      anthropic: {
        thinking: {
          type: 'enabled',
          budgetTokens: budget.claudeBudgetTokens
        }
      }
    }
  });
  return result.text;
}

3. Best-of-N with External Verifier (Step-Level Scrutiny)

For non-deterministic algorithmic problems, test-time compute achieves superior results by sampling $N$ parallel candidate trajectories and scoring them against automated test suites or AST verifiers:

User Prompt ──┬──► Candidate 1 (Thinking = 8k) ──► Ephemeral Sandbox ──► Passed (Score 1.0)
              ├──► Candidate 2 (Thinking = 8k) ──► Ephemeral Sandbox ──► Failed Test 3
              └──► Candidate 3 (Thinking = 8k) ──► Ephemeral Sandbox ──► Syntax Error
                                                           │
                                                           ▼
                                               Synthesized Winning Node

Best Practices

  1. Never Hardcode Fixed Thinking Budgets: Calibrate reasoning tokens dynamically based on code AST complexity, cyclomatic complexity, or security criticality.
  2. Decouple Thinking from Output Stream: Ensure thinking tokens are captured in structured audit logs rather than directly streamed to end-user UI channels.
  3. Combine Extended Thinking with Verifiable Tests: Allocate maximum compute to generate both code and exhaustive property-based tests within the same execution loop.
  4. Enforce Hard Timeout Bounds: Set execution deadlines on reasoning calls to prevent tail latency explosions when models enter recursive chain-of-thought exploration.

Common Pitfalls to Avoid

Anti-PatternOperational ConsequenceEngineering Remedy
Max thinking tokens on simple string manipulation10x latency and 15x cost inflationRoute through lightweight classifier to set budget_tokens: 0
Zero thinking tokens on concurrent state machinesRace conditions and hallucinated thread safetyEnforce minimum Tier 3 compute budget (16384 tokens)
Discarding thinking logsLost insight into mathematical edge casesStore thinking blocks in persistent telemetry for model evaluation

Integration with Other Skills (MANDATORY)

  • frontier-ai-models-expert — Coordinate model parameters and feature flags across Gemini 4 Pro, Claude 5.5, and GPT Astra 6.
  • llm-finops-router — Calculate token costs and optimize token expenditure per tenant.
  • ephemeral-wasm-sandbox-executor — Serve as deterministic validation ground for Best-of-N search candidates.
  • speculative-multi-draft-synthesizer — Supply multi-candidate generations to the arbiter.

Referenced By Orchestrators (MANDATORY)

  • brainstorming — Added to "Integrasi AI & LLM" and "Frontier AI & Simulation" matrix rows.
  • zero-to-prod-orchestrator — Integrated in Phase 4 (Backend APIs, Microservices & AI Agents).

<a name="bahasa-indonesia"></a>

Bahasa Indonesia

Integrasi Orkestrasi

Terhubung dan mengorkestrasi dengan brainstorming, zero-to-prod-orchestrator, ai-llm-integration-expert, frontier-ai-models-expert, llm-finops-router, dan speculative-multi-draft-synthesizer untuk menyeimbangkan kedalaman penalaran, latensi inferensi, dan biaya token pada model frontier.

Deskripsi

Panduan produksi untuk penskalaan komputasi waktu uji (test-time compute / penalaran saat inferensi) pada model AI frontier termasuk Google Gemini 4 Pro, Anthropic Claude 5.5 (Opus/Sonnet), dan OpenAI GPT Astra 6 (o6). Menyediakan arsitektur konkret untuk alokasi dinamis budget penalaran (thinking budget), eksplorasi tree-search (MCTS, beam search), verifikasi bertahap dengan Process Reward Model (PRM), dan konsensus self-consistency.

Kondisi Pemicu

Aktifkan skill ini ketika:

  • Merancang sistem yang menyelesaikan masalah algoritma rumit, verifikasi formal, atau arsitektur multi-tahap.
  • Mengonfigurasi parameter penalaran model (thinkingBudget, budget_tokens, reasoning_effort).
  • Mengimplementasikan algoritma pencarian waktu inferensi (Monte Carlo Tree Search, Best-of-N dengan verifikator).
  • Mengoptimalkan trade-off biaya vs akurasi agar tidak boros pada tugas sepele namun tetap maksimal pada titik kritis.

Konsep Inti & Pola Praktik

1. Matriks Alokasi Budget Penalaran Dinamis

Penskalaan komputasi waktu inferensi dibagi menjadi tiga tingkatan operasional berdasarkan entropi instruksi dan risiko eksekusi:

  • Tingkat 1 (CRUD & UI Sederhana): Komponen antarmuka dasar, konversi tipe data, formatting string. Budget thinking dinonaktifkan (0) atau disetel minimal 1024 token.
  • Tingkat 2 (Logika Bisnis & API): Transaksi multi-tabel, alur autentikasi, skema REST/GraphQL. Budget penalaran disetel pada rentang 4096 hingga 8192 token.
  • Tingkat 3 (Kritis & Formal Proof): State machine finansial, konsensus terdistribusi, refaktor skala besar. Budget penalaran dialokasikan penuh antara 16384 hingga 32768 token.

2. Protokol Evaluasi Langkah (Process Reward Models)

Alih-alih hanya mengevaluasi hasil akhir kode (Outcome-based), terapkan evaluasi pada setiap langkah dekomposisi logika (Step-level):

  1. Dekomposisi Masalah: Model menguraikan bukti logika ke dalam n premis mandiri.
  2. Validasi Invarian: Setiap premis diuji terhadap invariant sistem menggunakan sandboxing instan.
  3. Penyatuan Solusi: Menggabungkan langkah-langkah yang terbukti valid ke dalam kode akhir.

Praktik Terbaik

  1. Hindari Menyetel Budget Statis: Sesuaikan alokasi token penalaran secara dinamis berdasarkan kompleksitas siklomatik dan risiko domain.
  2. Pisahkan Log Penalaran dari UI: Simpan token penalaran di dalam sistem telemetri atau log audit internal, bukan ditampilkan mentah ke pengguna akhir.
  3. Kombinasikan dengan Pengujian Otomatis: Alokasikan daya komputasi tinggi untuk menghasilkan kode beserta unit test berbasis properti secara simultan.
  4. Terapkan Batas Waktu Eksekusi: Tetapkan timeout yang ketat pada panggilan penalaran guna mencegah lonjakan latensi (tail latency).

Jebakan Umum yang Harus Dihindari

Praktik BurukDampak OperasionalSolusi Rekayasa
Menyetel token thinking maksimal untuk manipulasi string dasarLatensi 10x lebih lambat dan biaya token melonjak 15xLewatkan ke pengklasifikasi ringan untuk menyetel budget_tokens: 0
Menonaktifkan thinking untuk perancangan konkurensi databaseMunculnya race condition dan deadlock tersembunyiWajibkan alokasi Tingkat 3 (16384 token)
Mengabaikan log internal penalaran (thinking traces)Kehilangan visibilitas saat terjadi kesalahan logikaSimpan blok pemikiran ke penyimpanan telemetri

Integrasi dengan Skill Lain (WAJIB)

  • frontier-ai-models-expert — Sinkronisasi parameter model untuk Gemini 4 Pro, Claude 5.5, dan GPT Astra 6.
  • llm-finops-router — Menghitung biaya token dan mengoptimalkan anggaran inferensi.
  • ephemeral-wasm-sandbox-executor — Lingkungan validasi deterministik untuk kandidat Best-of-N.
  • speculative-multi-draft-synthesizer — Memberikan kandidat multi-draft untuk ditengahi oleh arbiter.

Direferensikan oleh Orchestrator (WAJIB)

  • brainstorming — Tambahkan ke baris "Integrasi AI & LLM" dan "Frontier AI & Simulation" pada Matriks Orkestrasi.
  • zero-to-prod-orchestrator — Tambahkan ke Fase 4 (Backend APIs, Microservices & AI Agents).

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Expert guide for automated and manual Web Accessibility (a11y) testing — axe-core, Pa11y, Playwright a11y, screen reader testing, and WCAG 2.2 Level AA/AAA compliance / Panduan ahli pengujian aksesibilitas web.

日本語の概要は準備中です。原文の説明を表示しています。

roedyrustam/vibes-plug752026年10月9日 更新

Expert guide for intelligent model cascading and routing — complexity-scored task routing from Flash/Haiku to Sonnet/Opus/Astra, dynamic escalation with quality gates, 40-60% token cost reduction while maintaining output quality / Panduan ahli untuk kaskade dan routing model cerdas — routing tugas berbasis skor kompleksitas dari Flash/Haiku ke Sonnet/Opus/Astra, eskalasi dinamis dengan gerbang kualitas, pengurangan biaya token 40-60% dengan kualitas output terjaga.

日本語の概要は準備中です。原文の説明を表示しています。

roedyrustam/vibes-plug752026年10月9日 更新

Expert guide for Affective Computing, emotional AI, and real-time sentiment analysis through native multimodal tokens (voice intonation and facial micro-expressions) / Panduan ahli komputasi afektif, AI emosional, dan analisis sentimen real-time melalui token multimodal native.

日本語の概要は準備中です。原文の説明を表示しています。

roedyrustam/vibes-plug752026年10月9日 更新

Expert guide for AI-assisted coding workflows — agentic code generation, multi-agent code swarms, self-healing CI/CD, automated PR review, spec-to-code pipelines, codebase knowledge graphs, and human-in-the-loop approval gates / Panduan ahli untuk workflow pengkodean berbasis AI — generasi kode agentic, code swarm multi-agen, CI/CD self-healing, review PR otomatis, pipeline spec-to-code, knowledge graph codebase, dan gate persetujuan human-in-the-loop.

日本語の概要は準備中です。原文の説明を表示しています。

roedyrustam/vibes-plug752026年10月9日 更新

Expert guide for long-term episodic memory integration (Mem0 v2, Letta/MemGPT, Zep v2), memory tier architecture, pgvector HNSW storage, and unified context management for autonomous AI agents / Panduan ahli untuk integrasi memori episodik jangka panjang (Mem0 v2, Letta/MemGPT, Zep v2), arsitektur tier memori, penyimpanan pgvector HNSW, dan manajemen konteks terpadu untuk agen AI otonom.

日本語の概要は準備中です。原文の説明を表示しています。

roedyrustam/vibes-plug752026年10月9日 更新

Expert guide for designing Machine-to-Machine (M2M) micro-economies, autonomous agent wallets, and swarm budget allocation / Panduan ahli merancang ekonomi mikro antar-agen (M2M), dompet agen otonom, dan alokasi anggaran swarm.

日本語の概要は準備中です。原文の説明を表示しています。

roedyrustam/vibes-plug752026年10月9日 更新

roedyrustam のスキルをすべて見る

このスキルの問題を報告する