言語モデルに指示への応答や好ましい回答を学習させるため、TRLによる追加学習を案内するスキル。データ準備、学習方式の選択、報酬モデルの学習や評価を扱います。
- 回答例で指示に従うモデルを学習
- 回答の好みをDPOで学習したいとき
- 報酬モデルを使うRLHFの構築
31 件 ・ 関連度順
概要と使いどころ
言語モデルに指示への応答や好ましい回答を学習させるため、TRLによる追加学習を案内するスキル。データ準備、学習方式の選択、報酬モデルの学習や評価を扱います。
High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO training of large models (7B-70B+). Built on Ray, vLLM, ZeRO-3. 2× faster than DeepSpeedChat with distributed architecture and GPU resource sharing.
日本語の概要は準備中です。原文の説明を表示しています。
High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO training of large models (7B-70B+). Built on Ray, vLLM, ZeRO-3. 2× faster than DeepSpeedChat with distributed architecture and GPU resource sharing.
日本語の概要は準備中です。原文の説明を表示しています。
Framework RLHF de alta performance com aceleração Ray+vLLM. Use para treinamento PPO, GRPO, RLOO, DPO de modelos grandes (7B-70B+). Construído em Ray, vLLM, ZeRO-3. 2× mais rápido que DeepSpeedChat com arquitetura distribuída e compartilhamento de recursos GPU.
日本語の概要は準備中です。原文の説明を表示しています。
High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO training of large models (7B-70B+). Built on Ray, vLLM, ZeRO-3. 2× faster than DeepSpeedChat with distributed architecture and GPU resource sharing.
日本語の概要は準備中です。原文の説明を表示しています。
Prepare high-quality datasets for LLM fine-tuning with filtering, deduplication, augmentation, and RLHF data formatting. Activate on: fine-tuning data, training data curation, RLHF dataset, data quality filtering, SFT dataset. NOT for: model training infrastructure (ai-engineer), prompt engineering without fine-tuning (prompt-engineer).
日本語の概要は準備中です。原文の説明を表示しています。
Set up and run the full RLHF pipeline (SFT, reward model training, RL from reward model) using the Tinker API. Use when the user wants to do RLHF, train a reward model, or run the full preference-based RL pipeline.
日本語の概要は準備中です。原文の説明を表示しています。
Prepare high-quality datasets for LLM fine-tuning with filtering, deduplication, augmentation, and RLHF data formatting. Activate on: fine-tuning data, training data curation, RLHF dataset, data quality filtering, SFT dataset. NOT for: model training infrastructure (ai-engineer), prompt engineering without fine-tuning (prompt-engineer).
日本語の概要は準備中です。原文の説明を表示しています。
Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.
日本語の概要は準備中です。原文の説明を表示しています。
Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.
日本語の概要は準備中です。原文の説明を表示しています。
Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.
日本語の概要は準備中です。原文の説明を表示しています。
Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.
日本語の概要は準備中です。原文の説明を表示しています。
Prepare for Anthropic-specific technical interviews covering Constitutional AI, RLHF, interpretability, scaling laws, MCP, agentic systems, and AI safety. Activate on "anthropic interview", "constitutional AI prep", "alignment interview", "safety interview", "RLHF opinion", "interpretability prep", "anthropic technical". NOT for general ML interview prep, system design interviews, coding challenges, or behavioral interview practice.
日本語の概要は準備中です。原文の説明を表示しています。
Ajuste fino de LLMs usando reinforcement learning com TRL - SFT para ajuste de instruções, DPO para alinhamento de preferências, PPO/GRPO para otimização de recompensas e treinamento de modelo de recompensas. Use quando precisar de RLHF, alinhar modelo com preferências ou treinar a partir de feedback humano. Funciona com HuggingFace Transformers.
日本語の概要は準備中です。原文の説明を表示しています。
Fornece orientação para treinar LLMs com aprendizado por reforço usando verl (Volcano Engine RL). Use ao implementar RLHF, GRPO, PPO ou outros algoritmos de RL para pós-treinamento de LLMs em escala com backends de infraestrutura flexíveis.
日本語の概要は準備中です。原文の説明を表示しています。
Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.
日本語の概要は準備中です。原文の説明を表示しています。
Prepare for Anthropic-specific technical interviews covering Constitutional AI, RLHF, interpretability, scaling laws, MCP, agentic systems, and AI safety. Activate on "anthropic interview", "constitutional AI prep", "alignment interview", "safety interview", "RLHF opinion", "interpretability prep", "anthropic technical". NOT for general ML interview prep, system design interviews, coding challenges, or behavioral interview practice.
日本語の概要は準備中です。原文の説明を表示しています。
为 RL 训练设计奖励函数或排查奖励黑客时使用——决定奖励来自规则、人类偏好还是模型评判,结果奖励还是过程奖励,标量、向量还是生成式诊断,以及如何用路径约束与 RLVP 惩罚违规动作;也用于多轮 Agent 的信用分配、隐藏测试与早停判定设计。触发词:奖励设计、reward hacking、reward seeking、RLVR、RLHF、GRM、ORM、PRM、RLVP、信用分配、隐藏测试、过程奖励、奖励塑形、组内差异。
日本語の概要は準備中です。原文の説明を表示しています。
Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training
日本語の概要は準備中です。原文の説明を表示しています。
Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training
日本語の概要は準備中です。原文の説明を表示しています。
Use when fine-tuning LLMs, training custom models, or adapting foundation models for specific tasks. Invoke for configuring LoRA/QLoRA adapters, preparing JSONL training datasets, setting hyperparameters for fine-tuning runs, adapter training, transfer learning, finetuning with Hugging Face PEFT, OpenAI fine-tuning, instruction tuning, RLHF, DPO, or quantizing and deploying fine-tuned models. Trigger terms include: LoRA, QLoRA, PEFT, finetuning, fine-tuning, adapter tuning, LLM training, model training, custom model.
日本語の概要は準備中です。原文の説明を表示しています。
Plan, configure, and chain repo-native Nemotron customization steps into single-step or multi-step pipelines: curation, translation, SFT/PEFT (AutoModel or Megatron-Bridge), pretraining/CPT, RL alignment (DPO/RLVR/GRPO/RLHF), BYOB/MCQ benchmarks, checkpoint conversion, ModelOpt optimization, env profiles, and evaluation of trained checkpoints or existing/hosted endpoints. Use when a request names a Nemotron step or workflow, or asks to clean, translate, train, fine-tune, align, convert, optimize, evaluate, or compose these into a pipeline. Do NOT use for frontend/dashboard/visualization work, generic ML advice, billing/access, or non-Nemotron coding tasks.
日本語の概要は準備中です。原文の説明を表示しています。
This skill should be used when picking or diagnosing a training move (SFT, LoRA, DPO/KTO/ORPO, RFT, GRPO/PPO/RLOO, RLHF), or when the user mentions fine-tuning, post-training, training recipe, reward design, or weight updates. Decision tree by reward shape, smoke-run gate, three failure diagnostics, five false-progress patterns. Provider recipes and I/O contract in references/.
日本語の概要は準備中です。原文の説明を表示しています。
Fine-tunes LLMs and trains custom models using LoRA/QLoRA adapters, JSONL training datasets, hyperparameter configuration, RLHF, DPO, and model quantization.
日本語の概要は準備中です。原文の説明を表示しています。