本文へ移動
cccskills

「preference optimization」の検索結果

29 件 ・ 関連度順

概要と使いどころ

Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.

日本語の概要は準備中です。原文の説明を表示しています。

wshobson/agents4万2026年10月5日 更新

Rebalances context across Memory, Custom Instructions, knowledge files, and User Preferences in Claude Projects. Audits Memory for redundancy, staleness, and misplacement; prescribes optimization including Codification of stable Memory patterns into explicit User Preferences or Project CI rules. Use when user says "optimize my memory," "what should be in my memory," "trim my knowledge files," "reduce my context usage," "rebalance my project," "what should be in my preferences vs memory," or asks whether something belongs in Memory, a knowledge file, or User Preferences. Also trigger on symptom-phrased: "my project feels bloated," "Claude keeps forgetting things," "my context window keeps hitting limits." Activate whenever Memory-layer balance is the primary concern. Do NOT use for full Project audits (use rootnode-project-audit if available), single-prompt evaluation (use rootnode-prompt-validation if available), or global-layer audits that don't touch Project Memory (use rootnode-global-audit if available).

日本語の概要は準備中です。原文の説明を表示しています。

drayline/rootnode-skills402026年9月14日 更新

Audits and optimizes the five global Claude layers (User Preferences, Styles, Global Memory, Skills, MCP Connectors) using the Global Layer Scorecard (six dimensions, anchored 1-5 rubrics). Detects eight cross-layer failure modes and produces evolutionary recommendations (Promotion, Demotion, Codification, Skill Extraction). Use when user says "audit my global setup," "optimize my preferences," "review my Claude configuration," "check my cross-project setup," "are my preferences working," "clean up my global memory," or "what should be in my preferences vs my project." Also use when a user has 3+ Projects and wants to improve their shared foundation. Do NOT use for single-Project audits, Project Memory optimization, or full-stack audits (use rootnode-project-audit, rootnode-memory-optimization, or rootnode-full-stack-audit respectively, if available). Run on Opus 5 or Sonnet 5 at `high` effort (both defaults); depth reduces on legacy models.

日本語の概要は準備中です。原文の説明を表示しています。

drayline/rootnode-skills402026年9月14日 更新

Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference model needed, more efficient than DPO. Use for preference alignment when want simpler, faster training than DPO/PPO.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference model needed, more efficient than DPO. Use for preference alignment when want simpler, faster training than DPO/PPO.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年10月11日 更新

Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference model needed, more efficient than DPO. Use for preference alignment when want simpler, faster training than DPO/PPO.

日本語の概要は準備中です。原文の説明を表示しています。

Lord1Egypt/awesome-skill-forge22026年6月10日 更新

Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年10月11日 更新

Simple Preference Optimization para alinhamento de LLMs. Alternativa sem modelo de referência ao DPO com melhor desempenho (+6.4 pontos no AlpacaEval 2.0). Sem modelo de referência necessário, mais eficiente que DPO. Use para alinhamento de preferências quando quer treinamento mais simples e rápido que DPO/PPO.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

dpo

無料

Set up and run Direct Preference Optimization (DPO) training on preference datasets using the Tinker API. Use when the user wants to train with preference data, chosen/rejected pairs, or DPO.

日本語の概要は準備中です。原文の説明を表示しています。

uiuc-kang-lab/rlvr_generalization_bounds52026年5月13日 更新

Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

Use this skill to create, configure, or tune Field Service Lightning scheduling policies — including work rules (pass/fail filters) and service objectives (weighted ranking criteria). Covers the four default policies, custom policy design, work rule type selection, and objective weighting strategy. NOT for running or tuning the Optimizer itself — use admin/fsl-scheduling-optimization-design. NOT for resource skills and preferences — use admin/fsl-resource-management. NOT for resource capacity records — use admin/fsl-capacity-planning.

日本語の概要は準備中です。原文の説明を表示しています。

PranavNagrecha/AwesomeSalesforceSkills192026年10月4日 更新

dpo-guide

無料

Align language models with human preferences using Direct Preference Optimization — no reward model, no RL loop.

日本語の概要は準備中です。原文の説明を表示しています。

aicodedecode/awesome-muse-skills132026年10月10日 更新

Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.

日本語の概要は準備中です。原文の説明を表示しています。

ibragimov-oasis/vibe-coder22026年6月24日 更新

Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference model needed, more efficient than DPO. Use for preference alignment when want simpler, faster training than DPO/PPO.

日本語の概要は準備中です。原文の説明を表示しています。

ibragimov-oasis/vibe-coder22026年6月24日 更新

Decide whether to fine-tune at all, and route to the right method (SFT, DPO/ORPO/KTO, GRPO/RLVR, continued pretraining) and base model. Use when starting any fine-tuning effort, when unsure whether RAG or prompting would suffice, or when choosing between preference-optimization and reinforcement methods.

日本語の概要は準備中です。原文の説明を表示しています。

wshobson/agents4万2026年10月5日 更新

Use when the user wants their Claude agent to self-improve from past usage, asks about a nightly/offline 'sleep' or 'dream' cycle, memory/skill consolidation, or says things like 'make my agent better the more I use it', 'review my past sessions', 'learn my preferences', 'consolidate what you learned', 'run the sleep cycle', or wants to schedule background self-optimization. Drives the skillopt_sleep engine: harvest past sessions -> mine recurring tasks -> replay through a selected backend -> consolidate validated CLAUDE.md/SKILL.md behind a held-out gate.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/SkillOpt1.8万2026年10月7日 更新

Use when the user wants the dsh agent to self-improve from past usage, asks about a nightly/offline 'sleep' or 'dream' cycle, skill/memory consolidation, or says things like 'make my agent better the more I use it', 'review my past sessions', 'learn my preferences', 'consolidate what you learned', 'run the sleep cycle', or wants to schedule background self-optimization. Drives the skillopt_sleep engine through the skillopt_* tools: harvest past sessions -> mine recurring tasks -> replay via a selected backend -> consolidate validated skills behind a held-out gate.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/SkillOpt1.8万2026年10月7日 更新

Use when the user wants Codex to self-improve from past usage, asks about a nightly/offline 'sleep' or 'dream' cycle, wants Codex to review past sessions, learn preferences, consolidate memory/skills, run dry-run/run/adopt/status for SkillOpt-Sleep, or schedule background self-optimization. Drives the skillopt_sleep engine: harvest past sessions -> mine recurring tasks -> replay through a selected backend -> consolidate validated memory + skills behind a held-out gate.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/SkillOpt1.8万2026年10月7日 更新

Fine-tune models on Microsoft Foundry using SFT (supervised), DPO (preference), or RFT (reinforcement with graders). Covers dataset preparation, training job submission, deployment, and evaluation. USE FOR: fine-tune, SFT, DPO, RFT, training data, grader, distillation, fine-tuned model, training job, large file upload, calibrate grader, deploy fine-tuned model, evaluate fine-tuned model. DO NOT USE FOR: general model deployment without fine-tuning (use deploy-model), agent creation (use agents), prompt optimization without training (use prompt-optimizer).

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/skills3,1012026年10月10日 更新

Fine-tune models on Microsoft Foundry using SFT (supervised), DPO (preference), or RFT (reinforcement with graders). Covers dataset preparation, training job submission, deployment, and evaluation. USE FOR: fine-tune, SFT, DPO, RFT, training data, grader, distillation, fine-tuned model, training job, large file upload, calibrate grader, deploy fine-tuned model, evaluate fine-tuned model. DO NOT USE FOR: general model deployment without fine-tuning (use deploy-model), agent creation (use agents), prompt optimization without training (use prompt-optimizer).

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/azure-skills1,5552026年10月10日 更新

Generates 3D conformer ensembles using RDKit ETKDGv3 with knowledge-enhanced distance geometry, MMFF94/UFF force-field optimization, CREST + GFN2-xTB semi-empirical refinement, and macrocycle-aware torsion preferences. Provides explicit decision rules for single vs ensemble conformer use, RMSD pruning, energy windows, conformer count, and force-field choice. Use when preparing 3D ligands for docking, generating descriptor input for 3D QSAR, or sampling macrocycle/peptide conformational ensembles.

日本語の概要は準備中です。原文の説明を表示しています。

GPTomics/bioSkills1,2192026年8月15日 更新

Generates 3D conformer ensembles using RDKit ETKDGv3 with knowledge-enhanced distance geometry, MMFF94/UFF force-field optimization, CREST + GFN2-xTB semi-empirical refinement, and macrocycle-aware torsion preferences. Provides explicit decision rules for single vs ensemble conformer use, RMSD pruning, energy windows, conformer count, and force-field choice. Use when preparing 3D ligands for docking, generating descriptor input for 3D QSAR, or sampling macrocycle/peptide conformational ensembles.

日本語の概要は準備中です。原文の説明を表示しています。

BioTender-max/awesome-bio-agent-skills2002026年7月2日 更新

Chooses DPO, GRPO/RLVR, reward models, and over-optimization controls. Use when adapting a pretrained or SFT model with preference or verifiable reward signals.

日本語の概要は準備中です。原文の説明を表示しています。

vasilyu1983/AI-Agents-public912026年10月5日 更新