Audit whether a paper's EVALUATION DESIGN actually measures what it claims and whether its reporting is complete — the validity layer family D (experiment-forensics) cannot reach. Three patterns: train/test leakage means the reported score may not measure generalization (HP-EVAL-LEAKAGE — adopts the Kapoor & Narayanan 8-type / 3-category leakage taxonomy; the illegitimate-proxy / sampling-bias / pretraining-contamination subtypes hand off as needs_external_check, naming but NEVER running Oren-2023 exchangeability / Shi-2023 Min-K% / Golchin-2023 Time-Travel / BIG-bench canary); a load-bearing LLM judge is conflicted (same model/family as a compared system) or unvalidated (no human-agreement, no bias control) (HP-JUDGE-VALIDITY); a declared condition/metric is dropped or switched to favor the method, or 'best' is chosen with no held-out set (HP-SELECTIVE-REPORTING). Verdict-bearing at L0/L1 from the DESCRIBED protocol — NOT repo-gated like experiment-forensics; L2 only CONFIRMS against split/preprocessing/result files. A fresh cross-model reviewer (gpt-5.6-sol xhigh, read-only, fresh thread per pass) PROPOSES findings, each span-anchored to a ledger claim_id; tools/adjudicate_findings.py DECIDES the verdict. Leakage and under-reporting are usually HONEST methodological errors — every finding describes a discrepancy to CHECK, never an accusation. An LLM generating GROUND-TRUTH labels is HP-FAKE-GT (experiment-forensics) — routed there, not here. Emits eval-design-forensics.findings.json; computes NO verdict. Detect-only. Triggers: "eval design audit", "evaluation validity", "train/test leakage", "data leakage", "is the score measuring generalization", "LLM judge bias", "is the judge validated", "selective reporting", "cherry-picked results", "评估设计审计", "评测有效性", "数据泄漏", "训练测试集泄漏", "裁判模型有没有验证", "选择性报告".
日本語の概要は準備中です。原文の説明を表示しています。
wanshuiyin/Anti-Autoresearch☆ 1602026年10月7日 更新
Supports work with Outpost Bio's open microbiome foundation models - the Waypoint checkpoints (Waypoint-6m, Waypoint-45m, Waypoint-170m), the Atlas pretraining corpus, the Compass eight-task benchmark, or the `waypoint` CLI from the `waypoint-bio` package. Covers embedding microbiome samples, fine-tuning on taxonomic abundance data, benchmarking a checkpoint on Compass, pretraining a GPT-2 model on taxonomic abundance profiles, and converting MetaPhlAn, Kraken2, QIIME 2, or MGnify abundance tables into waypoint format.
日本語の概要は準備中です。原文の説明を表示しています。
K-Dense-AI/scientific-agent-skills☆ 4.8万2026年10月5日 更新
重みが公開されたLLMの回答拒否に関わる仕組みを分析し、再学習なしで重みを変更する実験を支援。計算環境に合うモデルや手法の選択、変更前後の品質評価まで扱います。
- LLMの回答拒否の仕組みを調べたいとき
- 再学習せず重み変更を実験したいとき
- 計算環境に合うモデルと手法を選びたいとき
NousResearch/hermes-agent☆ 25.3万2026年10月11日 更新
Merge multiple fine-tuned models using mergekit to combine capabilities without retraining. Use when creating specialized models by blending domain-specific expertise (math + coding + chat), improving performance beyond single models, or experimenting rapidly with model variants. Covers SLERP, TIES-Merging, DARE, Task Arithmetic, linear merging, and production deployment strategies.
日本語の概要は準備中です。原文の説明を表示しています。
davila7/claude-code-templates☆ 3.3万2026年10月11日 更新
Merge multiple fine-tuned models using mergekit to combine capabilities without retraining. Use when creating specialized models by blending domain-specific expertise (math + coding + chat), improving performance beyond single models, or experimenting rapidly with model variants. Covers SLERP, TIES-Merging, DARE, Task Arithmetic, linear merging, and production deployment strategies.
日本語の概要は準備中です。原文の説明を表示しています。
Orchestra-Research/AI-Research-SKILLs☆ 1.3万2026年10月11日 更新
Plan, configure, and chain repo-native Nemotron customization steps into single-step or multi-step pipelines: curation, translation, SFT/PEFT (AutoModel or Megatron-Bridge), pretraining/CPT, RL alignment (DPO/RLVR/GRPO/RLHF), BYOB/MCQ benchmarks, checkpoint conversion, ModelOpt optimization, env profiles, and evaluation of trained checkpoints or existing/hosted endpoints. Use when a request names a Nemotron step or workflow, or asks to clean, translate, train, fine-tune, align, convert, optimize, evaluate, or compose these into a pipeline. Do NOT use for frontend/dashboard/visualization work, generic ML advice, billing/access, or non-Nemotron coding tasks.
日本語の概要は準備中です。原文の説明を表示しています。
NVIDIA/skills☆ 3,5602026年10月10日 更新
Use ML.NET to train, evaluate, or integrate machine-learning models into .NET applications with realistic data preparation, inference, and deployment expectations. USE FOR: ML.NET integration; local model training or retraining; inference pipelines, model loading, evaluation, and deployment review. DO NOT USE FOR: unrelated stacks; generic tasks that do not need this specific guidance. INVOKES: inspect the repository context, edit targeted files, and run relevant build, test, lint, or validation commands when changes are made.
日本語の概要は準備中です。原文の説明を表示しています。
managedcode/dotnet-skills☆ 4852026年10月11日 更新
Think and work like an expert Computer Vision Scientist. Use when a task calls for Computer Vision Scientist judgment. Reasons from image formation, projective geometry (pinhole intrinsics and extrinsics, epipolar/PnP/bundle adjustment, similarity-scale ambiguity), and COCO/LVIS AP mechanics through DINOv3/SigLIP 2 foundation baselines, RF-DETR/YOLO26/SAM 3 models, COLMAP 4/GLOMAP and VGGT geometry, pycocotools/TrackEval/BOP evaluation, and CVPR reporting and EU AI Act limits while treating train–test and pretraining leakage, AP evaluation-setting gaming, preprocessing mismatches (EXIF, BGR, aliased resizing), label noise, and camera-convention and scale errors as first-class failure modes.
日本語の概要は準備中です。原文の説明を表示しています。
K-Dense-AI/scientific-agents☆ 1992026年10月3日 更新
Run a Meta-Harness-style optimization loop NATIVELY — automatically search over the scaffolding around a FIXED base model (memory, retrieval, context construction, prompt templates, summarization, tool-selection logic) by proposing candidate variants, scoring each on a cheap deterministic eval, and keeping a Pareto frontier of quality vs cost — using native Agent / Workflow / loop tools instead of a standalone Python harness. Use this whenever the user wants to optimize, evolve, tune, distill, or search over a harness, scaffold, prompt system, memory or retrieval policy, context-assembly code, or summarizer while keeping the model fixed; whenever they mention Meta-Harness, harness optimization, scaffold evolution, automatic prompt/memory optimization, an evolutionary or Pareto search over candidate implementations, or "make the harness/agent better without retraining"; and whenever the gain must come from the code AROUND the model rather than the model weights. Reproduces the Meta-Harness paper's method natively, with no claude_wrapper.py and no metered solver API.
日本語の概要は準備中です。原文の説明を表示しています。
001TMF/harness-forge☆ 802026年6月15日 更新
Run a Meta-Harness-style optimization loop NATIVELY — automatically search over the scaffolding around a FIXED base model (memory, retrieval, context construction, prompt templates, summarization, tool-selection logic) by proposing candidate variants, scoring each on a cheap deterministic eval, and keeping a Pareto frontier of quality vs cost — using native Agent / Workflow / loop tools instead of a standalone Python harness. Use this whenever the user wants to optimize, evolve, tune, distill, or search over a harness, scaffold, prompt system, memory or retrieval policy, context-assembly code, or summarizer while keeping the model fixed; whenever they mention Meta-Harness, harness optimization, scaffold evolution, automatic prompt/memory optimization, an evolutionary or Pareto search over candidate implementations, or "make the harness/agent better without retraining"; and whenever the gain must come from the code AROUND the model rather than the model weights. Reproduces the Meta-Harness paper's method natively, with no claude_wrapper.py and no metered solver API.
日本語の概要は準備中です。原文の説明を表示しています。
gabrielmoreira/agent-skills-mirror☆ 192026年10月11日 更新
Merge multiple fine-tuned models using mergekit to combine capabilities without retraining. Use when creating specialized models by blending domain-specific expertise (math + coding + chat), improving performance beyond single models, or experimenting rapidly with model variants. Covers SLERP, TIES-Merging, DARE, Task Arithmetic, linear merging, and production deployment strategies.
日本語の概要は準備中です。原文の説明を表示しています。
huang-sh/DeepScience☆ 42026年7月15日 更新
Builds and troubleshoots TorchDrug 0.2.1 workflows for molecular graphs, property prediction, self-supervised pretraining, molecule generation, retrosynthesis, protein representation learning, and knowledge graph reasoning. Use when code imports torchdrug or needs its datasets, models, tasks, or Engine.
日本語の概要は準備中です。原文の説明を表示しています。
K-Dense-AI/scientific-agent-skills☆ 4.8万2026年10月5日 更新
Use when polishing, diagnosing, tailoring, or exporting resumes for LLM, RAG, Agent, Agentic RL, post-training, pretraining, AIGC, search/ranking, multimodal, AI backend, or LLM algorithm internships from raw resume text, a materials folder, and/or a target job description. Audits evidence, maps JD fit, enforces truth boundaries, writes polished and targeted resumes, generates interviewer-style grilling questions, answer cards, evidence-upgrade plans, and optional open-source project recommendations without fabricating experience.
日本語の概要は準備中です。原文の説明を表示しています。
wanyichen06/LLMInternSkill☆ 3262026年8月4日 更新
Maps query single-cell data to reference atlases using scArches transfer learning with scVI and scANVI models. Transfers cell type labels without retraining on combined data. Use when annotating new single-cell datasets using pre-trained reference models.
日本語の概要は準備中です。原文の説明を表示しています。
BioTender-max/awesome-bio-agent-skills☆ 2002026年7月2日 更新
Builds a GPT and BPE tokenizer from scratch. Use when implementing autograd, attention, a nanoGPT loop, muP hyperparameter transfer, WSD annealing, loss spikes, or mid-training.
日本語の概要は準備中です。原文の説明を表示しています。
vasilyu1983/AI-Agents-public☆ 912026年10月5日 更新
Somatic movement and whole-body retraining as Joseph Pilates designed it in the Contrology method and as Moshé Feldenkrais designed it in Awareness Through Movement and Functional Integration, with enough cross-reference to the broader somatics landscape (Alexander Technique, Hanna Somatics, Body-Mind Centering) that an agent can place a user into the right method. Covers the Pilates reformer and mat system, the six Pilates principles, the Feldenkrais ATM lesson structure, the nervous-system learning frame that distinguishes somatics from exercise, and the safety posture that matters for rehab populations. Use for queries about core training, rehab-adjacent movement, chronic pain patterns, and learning-based movement re-education.
日本語の概要は準備中です。原文の説明を表示しています。
Tibsfox/gsd-skill-creator☆ 712026年7月20日 更新
Mescle múltiplos modelos ajustados usando mergekit para combinar capacidades sem retreinar. Use ao criar modelos especializados misturando expertise específica de domínio (math + coding + chat), melhorando performance além de modelos únicos, ou experimentando rapidamente variantes de modelos. Cobre SLERP, TIES-Merging, DARE, Task Arithmetic, mesclagem linear e estratégias de deploy em produção.
日本語の概要は準備中です。原文の説明を表示しています。
artubss/SKILLS-CLAUDE-CODE☆ 112026年5月17日 更新