Audit whether a paper's baseline comparisons are COMPLETE, FAIR, and SIGNIFICANT: a required recent SOTA baseline is missing while 'best/SOTA' is claimed (HP-MISSING-BASELINE); a baseline is undertuned / given less compute-tuning-data, run at a mismatched config, or the equal-budget ablation-as-baseline is absent (HP-WEAK-BASELINE); 'outperforms' is asserted over overlapping error bars or with no variance/seeds (HP-SIG-OVERLAP); and a cross-row 'improves over baseline by X%' is arithmetically wrong (HP-DELTA-ERROR, cross-row form only). A versioned per-domain baseline profile + a live leaderboard/recency search are assembled by the EXECUTOR as structured facts; a fresh cross-model reviewer (gpt-5.6-sol xhigh, read-only, fresh thread per dimension) PROPOSES findings, each span-anchored to a ledger claim_id; tools/adjudicate_findings.py DECIDES the verdict. Works at L0 (stated comparisons) and deepens at L2 (configs/result files). A completeness question it cannot settle internally becomes needs_external_check, never a guessed missing baseline. Emits baseline-comparison-audit.findings.json; computes NO verdict. Detect-only. Triggers: "baseline audit", "missing baselines", "is the comparison fair", "weak baseline", "baseline 误报", "SOTA earned?".
日本語の概要は準備中です。原文の説明を表示しています。
wanshuiyin/Anti-Autoresearch☆ 1602026年10月7日 更新
Mint a new YYYY-MM-DD-exact-coding-baseline snapshot under research/workflow-dev/export/. Detects the current best correctness- oriented workflow from research/workflow-dev/workflow-construction.md (or takes an explicit source-workflow argument) and transforms it from a lab workflow into a consumer-ready one on three axes: strips lab-only measurement content, re-enables human-in-the-loop checkpoints, and converts auto-loading config into an explicitly-invoked skill. Supports both promoted EXACT Coding lines: Opus/Hybrid and SOL/Predictive TDD. Exports Claude Code, pi, OpenCode, cursor-agent and, for SOL, Copilot. Trigger when the user says "exact-coding baseline export", "neue exact-coding baseline", "exact-coding-baseline-export", or asks to refresh the baseline snapshot.
日本語の概要は準備中です。原文の説明を表示しています。
marcoemrich/agentic_coding_lab☆ 122026年10月6日 更新
Establish baseline → snapshot → compare → history monitoring for any KPI, config, or metric that can drift during PRD v0.8 Deployment & Ops. Triggers on requests to monitor drift, baseline a value for later comparison, or when user asks "how do we track if X changes?", "baseline this", "config drift", "performance regression", "metric drift", "compare to last week", "is this getting worse?". Outputs MON-DRIFT-* entries with baseline + comparison rules.
日本語の概要は準備中です。原文の説明を表示しています。
mattgierhart/PRD-driven-context-engineering☆ 1812026年8月31日 更新
Agent fallback for CDelayedCall2_CServerSideClientBase_ProcessBaselineAckCall_vtable (auto-generated, category: func). Locate CDelayedCall2_CServerSideClientBase_ProcessBaselineAckCall_vtable in the CS2 engine module via IDA Pro MCP and emit a fresh, minimal-unique artifact. The deterministic preprocessor could not resolve this symbol on the current gamever - your job is the re-sign. Trigger: CDelayedCall2_CServerSideClientBase_ProcessBaselineAckCall_vtable, CDelayedCall2_CServerSideClientBase_ProcessBaselineAckCall_vtable
日本語の概要は準備中です。原文の説明を表示しています。
mrc4tt/CS2_VibeSignatures☆ 32026年10月10日 更新
Captures, masks, compares, and updates screenshot baselines for web and mobile suites, with per-region thresholds and a reviewed-diff rule for every baseline change. Use when a screenshot assertion fails, a baseline needs updating, or visual checks are being added; not for deciding what to verify.
日本語の概要は準備中です。原文の説明を表示しています。
HoangNguyen0403/agent-skills-standard☆ 5732026年10月10日 更新
Establish and maintain longitudinal baselines of CLI binary contents across versions. Covers marker selection by category (API / identity / config / telemetry / flag / function), weighted scoring, threshold-based system-presence detection, and per-version baseline records. Use when tracking a feature's lifecycle across releases, when probing for dark-launched or removed capabilities, or when verifying that a scanning tool itself still catches known-good markers on old binaries.
日本語の概要は準備中です。原文の説明を表示しています。
pjt222/agent-almanac☆ 372026年10月10日 更新
Establishes behavioral baselines from historical telemetry (process, network, or logon events) so hunts can flag rare and first-seen activity instead of relying on static signatures. Activates for requests to build a telemetry baseline, find rare or first-seen activity, or compute frequency baselines for anomaly hunting.
日本語の概要は準備中です。原文の説明を表示しています。
meltedinhex/analyst-ai-pack☆ 212026年7月7日 更新
Dev-only check that a drafted spec for THIS baseline repo won't ship dev-tree references to consumer installs. Catches three failure modes — shipped SKILL.md prose that references paths under `src/`, `tests/`, `scripts/`, `obj/` as runtime invocations (in ```bash fences``` OR `inline backticks`, plus shipped `.mjs`/`.js`/`.sh`/`.py` helper-file imports); new Python helpers added under `.claude/skills/<slug>/` (shipped helpers must be `.sh` or `.mjs`/`.js` going forward); and imports of modules that aren't in `obj/template/.claude/manifest.json` (consumer won't have the file). The aggregate scanner (`scan-shipped-skills.mjs`) walks only baseline-owned skill dirs (via `owner: baseline` frontmatter) at top level — `references/` and other subdirs are documentation, not runtime. BLOCKER findings hard-block implementation entry; ADVISORY surfaces but doesn't block. Read-only — surfaces a punch list; maintainer edits the spec and re-runs until CLEAN.
日本語の概要は準備中です。原文の説明を表示しています。
friedbotstudio/baseline☆ 142026年9月9日 更新
新しい Web API / CSS の利用追加に対し、Baseline 状態とブラウザ互換性 / progressive enhancement の有無を suggestion で示す。
s977043/river-review☆ 42026年10月11日 更新
Webページの表示や操作、APIの応答、ビルド・テストの所要時間を計測するスキル。変更前後の結果を比較し、Gitで共有する基準値をもとに性能の悪化を確認します。
- コード変更前後の性能比較
- ページが遅いという報告の調査
- 公開前に性能目標を確認したいとき
affaan-m/ECC☆ 27.7万2026年10月10日 更新
SEO drift monitoring: capture baselines of SEO-critical elements, detect changes, and track regressions over time. Git for SEO: baseline, diff, and track changes to your on-page SEO. Use when user says "SEO drift", "baseline", "track changes", "did anything break", "SEO regression", "compare SEO", "before and after", "monitor SEO changes", or "deployment check".
日本語の概要は準備中です。原文の説明を表示しています。
AgriciDaniel/claude-seo☆ 1.9万2026年10月5日 更新
Diagnose MSBuild build performance bottlenecks using binary log analysis. USE FOR: identifying why builds are slow by analyzing binlog performance summaries, detecting ResolveAssemblyReference (RAR) taking >5s, Roslyn analyzers consuming >30% of Csc time, single targets dominating >50% of build time, node utilization below 80%, excessive Copy tasks, NuGet restore running every build. Covers timeline analysis, Target/Task Performance Summary interpretation, and 7 common bottleneck categories. Use after build-perf-baseline has established measurements. DO NOT USE FOR: establishing initial baselines (use build-perf-baseline first), fixing incremental build issues (use incremental-build), parallelism tuning (use build-parallelism), non-MSBuild build systems.
日本語の概要は準備中です。原文の説明を表示しています。
dotnet/skills☆ 5,6032026年10月11日 更新
SEO drift monitoring — snapshot a site's SEO state and detect regressions over time. Captures a baseline (rankings/positions, indexed page count, titles & meta descriptions, canonical/robots directives, schema presence, key on-page elements) and on later runs diffs against it to surface what changed: ranking drops, pages that fell out of the index, titles/metas that were accidentally overwritten (a CMS/redeploy classic), canonicals or noindex flipped, schema that disappeared. Use this skill when the user wants to monitor SEO over time, catch regressions after a site change / migration / redeploy, set a baseline, diff against a previous state, or asks "what changed on my site's SEO" or "did my redesign break SEO". Trigger on: "SEO drift", "SEO monitoring", "track SEO over time", "did my site change break SEO", "after migration SEO", "SEO regression", "baseline my SEO", "compare SEO to last month", "my titles changed", "pages fell out of the index". For a one-time full audit use /seo-analysis.
日本語の概要は準備中です。原文の説明を表示しています。
nowork-studio/notfair-plugin☆ 3,9172026年10月10日 更新
Separate a strategy return series into declared baseline exposure and residual edge with returns-based OLS attribution, HAC inference, rolling stability, alternate-baseline sensitivity, and regime breakdowns. Use when evaluating whether backtest, out-of-sample, or live returns contain independent alpha beyond market, equal-weight, momentum, sector, or user-supplied factor returns; when explaining whether a drawdown came from baseline exposure or strategy-specific behavior; or when a strategy needs an attribution quality gate after backtesting. Do not use for holdings-based Brinson attribution, feature-level Shapley explanations, or analysis from summary metrics without a dated return series.
日本語の概要は準備中です。原文の説明を表示しています。
tradermonty/claude-trading-skills☆ 2,9862026年10月11日 更新
Use when the user asks to "map what our surfaces say today", "inventory our current messaging", or "find the gap between what we say and what we mean"; produces the narrative baseline — a surface-by-surface inventory of what every owned touchpoint (homepage, pricing, docs, decks, social bios, email footers) claims RIGHT NOW, each line labeled Measured / User-provided / Estimated, plus a per-surface gap read vs the intended message and the drift-baseline snapshot the Evaluate phase measures future drift against. Not for authoring the canon — use message-system-architect; not for scoring the surfaces or running the vetoes — use narrative-quality-auditor. 现状叙事盘点/各触点口径/意图差距/漂移基线
日本語の概要は準備中です。原文の説明を表示しています。
aaron-he-zhu/aaron-marketing-skills☆ 2,9012026年10月11日 更新
Generates a meta-analysis baseline characteristics section (text + table) from raw data. Supports Chinese and English. Use when the user provides baseline data and wants a formatted results section.
日本語の概要は準備中です。原文の説明を表示しています。
aipoch/medical-research-skills☆ 1,9392026年9月17日 更新
Use when managing perf baselines, consolidating results, or comparing versions. Ensures one baseline JSON per version.
日本語の概要は準備中です。原文の説明を表示しています。
composio-community/awesome-claude-plugins☆ 1,9392026年7月26日 更新
SEO drift monitoring: capture baselines of SEO-critical elements, detect changes, and track regressions over time. Git for SEO: baseline, diff, and track changes to your on-page SEO. Use when user says "SEO drift", "baseline", "track changes", "did anything break", "SEO regression", "compare SEO", "before and after", "monitor SEO changes", or "deployment check".
日本語の概要は準備中です。原文の説明を表示しています。
hashgraph-online/awesome-codex-plugins☆ 1,2752026年10月11日 更新
Establish a security baseline for a website or web app. Use this skill when configuring HTTPS and TLS, setting security headers, planning secrets management, evaluating CSP policies, doing a basic security audit, or hardening a site before launch. Triggers on security headers, HTTPS, TLS, CSP, content security policy, HSTS, secrets management, vulnerability scan, security audit, harden, OWASP, security baseline. Also triggers when a security review is required for compliance or before going live.
日本語の概要は準備中です。原文の説明を表示しています。
rampstackco/claude-skills☆ 9462026年10月7日 更新
SEO drift monitoring: capture baselines of SEO-critical elements, detect changes, and track regressions over time. Git for SEO — baseline, diff, and track changes to your on-page SEO. Use when user says "SEO drift", "baseline", "track changes", "did anything break", "SEO regression", "compare SEO", "before and after", "monitor SEO changes", or "deployment check".
日本語の概要は準備中です。原文の説明を表示しています。
AgriciDaniel/codex-seo☆ 8022026年9月12日 更新
Diagnose MSBuild build performance bottlenecks using binary log analysis. USE FOR: identifying why builds are slow by analyzing binlog performance summaries, detecting ResolveAssemblyReference (RAR) taking >5s, Roslyn analyzers consuming >30% of Csc time, single targets dominating >50% of build time, node utilization below 80%, excessive Copy tasks, NuGet restore running every build. Covers timeline analysis, Target/Task Performance Summary interpretation, and 7 common bottleneck categories. Use after build-perf-baseline has established measurements. DO NOT USE FOR: establishing initial baselines (use build-perf-baseline first), fixing incremental build issues (use incremental-build), parallelism tuning (use build-parallelism), non-MSBuild build systems.
日本語の概要は準備中です。原文の説明を表示しています。
managedcode/dotnet-skills☆ 4852026年10月11日 更新
Apply stable-baselines3 in reproducible local data workflows with version-aware APIs, explicit assumptions, and validation. Use when the user chooses stable-baselines3 or its strengths fit the task.
日本語の概要は準備中です。原文の説明を表示しています。
sandbaseai/sandbase-skills☆ 2032026年9月26日 更新
Implement visual regression testing with Playwright screenshots, Chromatic, Percy, and Argos CI. Covers baseline management, diff threshold tuning, dynamic content masking, responsive viewport testing, and review/approval workflows. Use when: "visual test," "screenshot," "visual regression," "pixel diff," "snapshot diff," "update baselines," "Chromatic," "percy snapshot," "argos screenshot." Not for: bulk baseline regeneration after a redesign broke many tests — use selector-drift-recovery; cross-browser rendering matrices — use cross-browser-testing; general Playwright test structure — use playwright-automation. Related: playwright-automation, ci-cd-integration, cross-browser-testing.
日本語の概要は準備中です。原文の説明を表示しています。
petrkindlmann/qa-skills☆ 1722026年6月11日 更新
Learn the target framework from enabled knowledge tools and implement a baseline GPU kernel. Use this skill to understand compute semantics, determine the target platform and framework, search reference implementations, and produce a correct V0 baseline with performance records for later profile-driven optimization.
日本語の概要は準備中です。原文の説明を表示しています。
alibaba/atrex-kernel-agent☆ 1672026年10月10日 更新