本文へ移動
cccskills

「rlhf」の検索結果

31 件 ・ 関連度順

概要と使いどころ

trl-fine-tuning

無料日本語概要

言語モデルに指示への応答や好ましい回答を学習させるため、TRLによる追加学習を案内するスキル。データ準備、学習方式の選択、報酬モデルの学習や評価を扱います。

  • 回答例で指示に従うモデルを学習
  • 回答の好みをDPOで学習したいとき
  • 報酬モデルを使うRLHFの構築
NousResearch/hermes-agent25.3万2026年10月11日 更新

High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO training of large models (7B-70B+). Built on Ray, vLLM, ZeRO-3. 2× faster than DeepSpeedChat with distributed architecture and GPU resource sharing.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO training of large models (7B-70B+). Built on Ray, vLLM, ZeRO-3. 2× faster than DeepSpeedChat with distributed architecture and GPU resource sharing.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Framework RLHF de alta performance com aceleração Ray+vLLM. Use para treinamento PPO, GRPO, RLOO, DPO de modelos grandes (7B-70B+). Construído em Ray, vLLM, ZeRO-3. 2× mais rápido que DeepSpeedChat com arquitetura distribuída e compartilhamento de recursos GPU.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

openrlhf

無料

High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO training of large models (7B-70B+). Built on Ray, vLLM, ZeRO-3. 2× faster than DeepSpeedChat with distributed architecture and GPU resource sharing.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

Prepare high-quality datasets for LLM fine-tuning with filtering, deduplication, augmentation, and RLHF data formatting. Activate on: fine-tuning data, training data curation, RLHF dataset, data quality filtering, SFT dataset. NOT for: model training infrastructure (ai-engineer), prompt engineering without fine-tuning (prompt-engineer).

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/windags-skills132026年10月1日 更新

rlhf

無料

Set up and run the full RLHF pipeline (SFT, reward model training, RL from reward model) using the Tinker API. Use when the user wants to do RLHF, train a reward model, or run the full preference-based RL pipeline.

日本語の概要は準備中です。原文の説明を表示しています。

uiuc-kang-lab/rlvr_generalization_bounds52026年5月13日 更新

Prepare high-quality datasets for LLM fine-tuning with filtering, deduplication, augmentation, and RLHF data formatting. Activate on: fine-tuning data, training data curation, RLHF dataset, data quality filtering, SFT dataset. NOT for: model training infrastructure (ai-engineer), prompt engineering without fine-tuning (prompt-engineer).

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/port-daddy22026年10月8日 更新

Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Prepare for Anthropic-specific technical interviews covering Constitutional AI, RLHF, interpretability, scaling laws, MCP, agentic systems, and AI safety. Activate on "anthropic interview", "constitutional AI prep", "alignment interview", "safety interview", "RLHF opinion", "interpretability prep", "anthropic technical". NOT for general ML interview prep, system design interviews, coding challenges, or behavioral interview practice.

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/windags-skills132026年10月1日 更新

Ajuste fino de LLMs usando reinforcement learning com TRL - SFT para ajuste de instruções, DPO para alinhamento de preferências, PPO/GRPO para otimização de recompensas e treinamento de modelo de recompensas. Use quando precisar de RLHF, alinhar modelo com preferências ou treinar a partir de feedback humano. Funciona com HuggingFace Transformers.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

Fornece orientação para treinar LLMs com aprendizado por reforço usando verl (Volcano Engine RL). Use ao implementar RLHF, GRPO, PPO ou outros algoritmos de RL para pós-treinamento de LLMs em escala com backends de infraestrutura flexíveis.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

Prepare for Anthropic-specific technical interviews covering Constitutional AI, RLHF, interpretability, scaling laws, MCP, agentic systems, and AI safety. Activate on "anthropic interview", "constitutional AI prep", "alignment interview", "safety interview", "RLHF opinion", "interpretability prep", "anthropic technical". NOT for general ML interview prep, system design interviews, coding challenges, or behavioral interview practice.

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/port-daddy22026年10月8日 更新

为 RL 训练设计奖励函数或排查奖励黑客时使用——决定奖励来自规则、人类偏好还是模型评判,结果奖励还是过程奖励,标量、向量还是生成式诊断,以及如何用路径约束与 RLVP 惩罚违规动作;也用于多轮 Agent 的信用分配、隐藏测试与早停判定设计。触发词:奖励设计、reward hacking、reward seeking、RLVR、RLHF、GRM、ORM、PRM、RLVP、信用分配、隐藏测试、过程奖励、奖励塑形、组内差异。

日本語の概要は準備中です。原文の説明を表示しています。

bojieli/ai-agent-book5.3万2026年10月9日 更新

Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Use when fine-tuning LLMs, training custom models, or adapting foundation models for specific tasks. Invoke for configuring LoRA/QLoRA adapters, preparing JSONL training datasets, setting hyperparameters for fine-tuning runs, adapter training, transfer learning, finetuning with Hugging Face PEFT, OpenAI fine-tuning, instruction tuning, RLHF, DPO, or quantizing and deploying fine-tuned models. Trigger terms include: LoRA, QLoRA, PEFT, finetuning, fine-tuning, adapter tuning, LLM training, model training, custom model.

日本語の概要は準備中です。原文の説明を表示しています。

Jeffallan/claude-skills1.2万2026年10月4日 更新

Plan, configure, and chain repo-native Nemotron customization steps into single-step or multi-step pipelines: curation, translation, SFT/PEFT (AutoModel or Megatron-Bridge), pretraining/CPT, RL alignment (DPO/RLVR/GRPO/RLHF), BYOB/MCQ benchmarks, checkpoint conversion, ModelOpt optimization, env profiles, and evaluation of trained checkpoints or existing/hosted endpoints. Use when a request names a Nemotron step or workflow, or asks to clean, translate, train, fine-tune, align, convert, optimize, evaluate, or compose these into a pipeline. Do NOT use for frontend/dashboard/visualization work, generic ML advice, billing/access, or non-Nemotron coding tasks.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5602026年10月10日 更新

This skill should be used when picking or diagnosing a training move (SFT, LoRA, DPO/KTO/ORPO, RFT, GRPO/PPO/RLOO, RLHF), or when the user mentions fine-tuning, post-training, training recipe, reward design, or weight updates. Decision tree by reward shape, smoke-run gate, three failure diagnostics, five false-progress patterns. Provider recipes and I/O contract in references/.

日本語の概要は準備中です。原文の説明を表示しています。

evo-hq/evo1,4672026年10月5日 更新

Fine-tunes LLMs and trains custom models using LoRA/QLoRA adapters, JSONL training datasets, hyperparameter configuration, RLHF, DPO, and model quantization.

日本語の概要は準備中です。原文の説明を表示しています。

paperclipai/companies9192026年3月24日 更新