Expert guide for automated and manual Web Accessibility (a11y) testing — axe-core, Pa11y, Playwright a11y, screen reader testing, and WCAG 2.2 Level AA/AAA compliance / Panduan ahli pengujian aksesibilitas web.
日本語の概要は準備中です。原文の説明を表示しています。
Expert guide for scaling test-time compute, dynamic reasoning token allocation (Gemini 4 Pro, Claude 5.5, GPT Astra 6), Monte Carlo Tree Search (MCTS), and Process Reward Models / Panduan ahli penskalaan komputasi waktu uji (test-time compute), alokasi dinamis budget penalaran, MCTS, dan Process Reward Models.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
<a name="english"></a>
Connects and orchestrates with brainstorming, zero-to-prod-orchestrator, ai-llm-integration-expert, frontier-ai-models-expert, llm-finops-router, and speculative-multi-draft-synthesizer to dynamically balance reasoning depth, inference latency, and token expenditures across frontier models.
Production guide for scaling test-time compute (inference-time reasoning) across frontier AI models including Google Gemini 4 Pro, Anthropic Claude 5.5 (Opus/Sonnet), and OpenAI GPT Astra 6 (o6). Provides concrete architectures for dynamic thinking budget allocation, tree-search exploration (MCTS, beam search), step-level Process Reward Model (PRM) verification, and self-consistency consensus loops.
Activate this skill when:
thinkingBudget, budget_tokens, reasoning_effort).Inference-time compute scaling operates across three distinct operational tiers based on prompt entropy and execution risk:
| Complexity Tier | Target Problem Types | Gemini 4 Pro / 3.x (thinkingBudget) | Claude 5.5 (budget_tokens) | GPT Astra 6 (reasoning_effort) |
|---|---|---|---|---|
| Tier 1: Deterministic / CRUD | Simple UI components, DTO mappings, regex conversions | 0 (Thinking disabled) or 1024 | 1024 | low |
| Tier 2: Business Logic & APIs | Multi-table DB transactions, authentication flows, REST schemas | 4096 - 8192 | 4096 - 8192 | medium |
| Tier 3: Mission-Critical Hard | Financial state machines, distributed consensus, cryptographic proofs, complex refactoring | 16384 - 32768 | 16384 - 32768 | high |
import { createGoogleGenerativeAI } from '@ai-sdk/google';
import { createAnthropic } from '@ai-sdk/anthropic';
import { generateText } from 'ai';
export type TaskComplexity = 'trivial' | 'standard' | 'complex' | 'extreme';
export interface ComputeAllocation {
geminiThinkingBudget: number;
claudeBudgetTokens: number;
reasoningEffort: 'low' | 'medium' | 'high';
}
export function resolveComputeBudget(complexity: TaskComplexity): ComputeAllocation {
switch (complexity) {
case 'trivial':
return { geminiThinkingBudget: 0, claudeBudgetTokens: 1024, reasoningEffort: 'low' };
case 'standard':
return { geminiThinkingBudget: 4096, claudeBudgetTokens: 4096, reasoningEffort: 'medium' };
case 'complex':
return { geminiThinkingBudget: 16384, claudeBudgetTokens: 16384, reasoningEffort: 'high' };
case 'extreme':
return { geminiThinkingBudget: 32768, claudeBudgetTokens: 32768, reasoningEffort: 'high' };
}
}
export async function executeOptimizedReasoning(
prompt: string,
complexity: TaskComplexity,
provider: 'gemini' | 'claude'
): Promise<string> {
const budget = resolveComputeBudget(complexity);
if (provider === 'gemini') {
const google = createGoogleGenerativeAI({
apiKey: process.env.GEMINI_API_KEY || ''
});
const result = await generateText({
model: google('gemini-4-pro', {
useSearchGrounding: false
}),
prompt,
providerOptions: {
google: {
thinkingBudget: budget.geminiThinkingBudget
}
}
});
return result.text;
}
const anthropic = createAnthropic({
apiKey: process.env.ANTHROPIC_API_KEY || ''
});
const result = await generateText({
model: anthropic('claude-3-7-sonnet-20250219'),
prompt,
providerOptions: {
anthropic: {
thinking: {
type: 'enabled',
budgetTokens: budget.claudeBudgetTokens
}
}
}
});
return result.text;
}
For non-deterministic algorithmic problems, test-time compute achieves superior results by sampling $N$ parallel candidate trajectories and scoring them against automated test suites or AST verifiers:
User Prompt ──┬──► Candidate 1 (Thinking = 8k) ──► Ephemeral Sandbox ──► Passed (Score 1.0)
├──► Candidate 2 (Thinking = 8k) ──► Ephemeral Sandbox ──► Failed Test 3
└──► Candidate 3 (Thinking = 8k) ──► Ephemeral Sandbox ──► Syntax Error
│
▼
Synthesized Winning Node
| Anti-Pattern | Operational Consequence | Engineering Remedy |
|---|---|---|
| Max thinking tokens on simple string manipulation | 10x latency and 15x cost inflation | Route through lightweight classifier to set budget_tokens: 0 |
| Zero thinking tokens on concurrent state machines | Race conditions and hallucinated thread safety | Enforce minimum Tier 3 compute budget (16384 tokens) |
| Discarding thinking logs | Lost insight into mathematical edge cases | Store thinking blocks in persistent telemetry for model evaluation |
frontier-ai-models-expert — Coordinate model parameters and feature flags across Gemini 4 Pro, Claude 5.5, and GPT Astra 6.llm-finops-router — Calculate token costs and optimize token expenditure per tenant.ephemeral-wasm-sandbox-executor — Serve as deterministic validation ground for Best-of-N search candidates.speculative-multi-draft-synthesizer — Supply multi-candidate generations to the arbiter.brainstorming — Added to "Integrasi AI & LLM" and "Frontier AI & Simulation" matrix rows.zero-to-prod-orchestrator — Integrated in Phase 4 (Backend APIs, Microservices & AI Agents).<a name="bahasa-indonesia"></a>
Terhubung dan mengorkestrasi dengan brainstorming, zero-to-prod-orchestrator, ai-llm-integration-expert, frontier-ai-models-expert, llm-finops-router, dan speculative-multi-draft-synthesizer untuk menyeimbangkan kedalaman penalaran, latensi inferensi, dan biaya token pada model frontier.
Panduan produksi untuk penskalaan komputasi waktu uji (test-time compute / penalaran saat inferensi) pada model AI frontier termasuk Google Gemini 4 Pro, Anthropic Claude 5.5 (Opus/Sonnet), dan OpenAI GPT Astra 6 (o6). Menyediakan arsitektur konkret untuk alokasi dinamis budget penalaran (thinking budget), eksplorasi tree-search (MCTS, beam search), verifikasi bertahap dengan Process Reward Model (PRM), dan konsensus self-consistency.
Aktifkan skill ini ketika:
thinkingBudget, budget_tokens, reasoning_effort).Penskalaan komputasi waktu inferensi dibagi menjadi tiga tingkatan operasional berdasarkan entropi instruksi dan risiko eksekusi:
0) atau disetel minimal 1024 token.4096 hingga 8192 token.16384 hingga 32768 token.Alih-alih hanya mengevaluasi hasil akhir kode (Outcome-based), terapkan evaluasi pada setiap langkah dekomposisi logika (Step-level):
| Praktik Buruk | Dampak Operasional | Solusi Rekayasa |
|---|---|---|
| Menyetel token thinking maksimal untuk manipulasi string dasar | Latensi 10x lebih lambat dan biaya token melonjak 15x | Lewatkan ke pengklasifikasi ringan untuk menyetel budget_tokens: 0 |
| Menonaktifkan thinking untuk perancangan konkurensi database | Munculnya race condition dan deadlock tersembunyi | Wajibkan alokasi Tingkat 3 (16384 token) |
| Mengabaikan log internal penalaran (thinking traces) | Kehilangan visibilitas saat terjadi kesalahan logika | Simpan blok pemikiran ke penyimpanan telemetri |
frontier-ai-models-expert — Sinkronisasi parameter model untuk Gemini 4 Pro, Claude 5.5, dan GPT Astra 6.llm-finops-router — Menghitung biaya token dan mengoptimalkan anggaran inferensi.ephemeral-wasm-sandbox-executor — Lingkungan validasi deterministik untuk kandidat Best-of-N.speculative-multi-draft-synthesizer — Memberikan kandidat multi-draft untuk ditengahi oleh arbiter.brainstorming — Tambahkan ke baris "Integrasi AI & LLM" dan "Frontier AI & Simulation" pada Matriks Orkestrasi.zero-to-prod-orchestrator — Tambahkan ke Fase 4 (Backend APIs, Microservices & AI Agents).まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Expert guide for automated and manual Web Accessibility (a11y) testing — axe-core, Pa11y, Playwright a11y, screen reader testing, and WCAG 2.2 Level AA/AAA compliance / Panduan ahli pengujian aksesibilitas web.
日本語の概要は準備中です。原文の説明を表示しています。
Expert guide for intelligent model cascading and routing — complexity-scored task routing from Flash/Haiku to Sonnet/Opus/Astra, dynamic escalation with quality gates, 40-60% token cost reduction while maintaining output quality / Panduan ahli untuk kaskade dan routing model cerdas — routing tugas berbasis skor kompleksitas dari Flash/Haiku ke Sonnet/Opus/Astra, eskalasi dinamis dengan gerbang kualitas, pengurangan biaya token 40-60% dengan kualitas output terjaga.
日本語の概要は準備中です。原文の説明を表示しています。
Expert guide for Affective Computing, emotional AI, and real-time sentiment analysis through native multimodal tokens (voice intonation and facial micro-expressions) / Panduan ahli komputasi afektif, AI emosional, dan analisis sentimen real-time melalui token multimodal native.
日本語の概要は準備中です。原文の説明を表示しています。
Expert guide for AI-assisted coding workflows — agentic code generation, multi-agent code swarms, self-healing CI/CD, automated PR review, spec-to-code pipelines, codebase knowledge graphs, and human-in-the-loop approval gates / Panduan ahli untuk workflow pengkodean berbasis AI — generasi kode agentic, code swarm multi-agen, CI/CD self-healing, review PR otomatis, pipeline spec-to-code, knowledge graph codebase, dan gate persetujuan human-in-the-loop.
日本語の概要は準備中です。原文の説明を表示しています。
Expert guide for long-term episodic memory integration (Mem0 v2, Letta/MemGPT, Zep v2), memory tier architecture, pgvector HNSW storage, and unified context management for autonomous AI agents / Panduan ahli untuk integrasi memori episodik jangka panjang (Mem0 v2, Letta/MemGPT, Zep v2), arsitektur tier memori, penyimpanan pgvector HNSW, dan manajemen konteks terpadu untuk agen AI otonom.
日本語の概要は準備中です。原文の説明を表示しています。
Expert guide for designing Machine-to-Machine (M2M) micro-economies, autonomous agent wallets, and swarm budget allocation / Panduan ahli merancang ekonomi mikro antar-agen (M2M), dompet agen otonom, dan alokasi anggaran swarm.
日本語の概要は準備中です。原文の説明を表示しています。