PyTorchのモデルや学習コードを作成・点検し、データ読み込み、実験の再現性、学習状態の保存、GPUメモリや処理速度の改善を支援するスキル。
- PyTorchの学習コード作成・レビュー
- 実験結果が揃わない原因を調べたいとき
- 読み込み速度やGPUメモリの改善
25 件 ・ 関連度順
概要と使いどころ
PyTorchのモデルや学習コードを作成・点検し、データ読み込み、実験の再現性、学習状態の保存、GPUメモリや処理速度の改善を支援するスキル。
Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference model needed, more efficient than DPO. Use for preference alignment when want simpler, faster training than DPO/PPO.
日本語の概要は準備中です。原文の説明を表示しています。
Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference model needed, more efficient than DPO. Use for preference alignment when want simpler, faster training than DPO/PPO.
日本語の概要は準備中です。原文の説明を表示しています。
Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference model needed, more efficient than DPO. Use for preference alignment when want simpler, faster training than DPO/PPO.
日本語の概要は準備中です。原文の説明を表示しています。
Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace. Use when training large-scale models with limited compute (5× cost reduction vs dense models), implementing sparse architectures like Mixtral 8x7B or DeepSeek-V3, or scaling model capacity without proportional compute increase. Covers MoE architectures, routing mechanisms, load balancing, expert parallelism, and inference optimization.
日本語の概要は準備中です。原文の説明を表示しています。
Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace. Use when training large-scale models with limited compute (5× cost reduction vs dense models), implementing sparse architectures like Mixtral 8x7B or DeepSeek-V3, or scaling model capacity without proportional compute increase. Covers MoE architectures, routing mechanisms, load balancing, expert parallelism, and inference optimization.
日本語の概要は準備中です。原文の説明を表示しています。
Train GPT-2 scale models (~124M parameters) efficiently on a single GPU. Covers GPT-124M architecture, tokenized dataset loading (e.g., HuggingFace Hub shards), modern optimizers (Muon, AdamW), mixed precision training, and training loop implementation.
日本語の概要は準備中です。原文の説明を表示しています。
Simple Preference Optimization para alinhamento de LLMs. Alternativa sem modelo de referência ao DPO com melhor desempenho (+6.4 pontos no AlpacaEval 2.0). Sem modelo de referência necessário, mais eficiente que DPO. Use para alinhamento de preferências quando quer treinamento mais simples e rápido que DPO/PPO.
日本語の概要は準備中です。原文の説明を表示しています。
Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace. Use when training large-scale models with limited compute (5× cost reduction vs dense models), implementing sparse architectures like Mixtral 8x7B or DeepSeek-V3, or scaling model capacity without proportional compute increase. Covers MoE architectures, routing mechanisms, load balancing, expert parallelism, and inference optimization.
日本語の概要は準備中です。原文の説明を表示しています。
Expert guidance for fast fine-tuning with Unsloth - 2-5x faster training, 50-80% less memory, LoRA/QLoRA optimization
日本語の概要は準備中です。原文の説明を表示しています。
Expert guidance for fast fine-tuning with Unsloth - 2-5x faster training, 50-80% less memory, LoRA/QLoRA optimization
日本語の概要は準備中です。原文の説明を表示しています。
Expert guidance for fast fine-tuning with Unsloth - 2-5x faster training, 50-80% less memory, LoRA/QLoRA optimization
日本語の概要は準備中です。原文の説明を表示しています。
Expert guidance for fast fine-tuning with Unsloth - 2-5x faster training, 50-80% less memory, LoRA/QLoRA optimization
日本語の概要は準備中です。原文の説明を表示しています。
Treinar modelos de Mixture of Experts (MoE) usando DeepSpeed ou HuggingFace. Use ao treinar modelos em larga escala com computação limitada (redução de 5× em custos vs modelos densos), implementar arquiteturas esparsas como Mixtral 8x7B ou DeepSeek-V3, ou escalar capacidade de modelo sem aumento proporcional de computação. Cobre arquiteturas MoE, mecanismos de roteamento, balanceamento de carga, paralelismo de especialistas e otimização de inferência.
日本語の概要は準備中です。原文の説明を表示しています。
Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference model needed, more efficient than DPO. Use for preference alignment when want simpler, faster training than DPO/PPO.
日本語の概要は準備中です。原文の説明を表示しています。
Quantizes LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss. Use when GPU memory is limited, need to fit larger models, or want faster inference. Supports INT8, NF4, FP4 formats, QLoRA training, and 8-bit optimizers. Works with HuggingFace Transformers.
日本語の概要は準備中です。原文の説明を表示しています。
Quantizes LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss. Use when GPU memory is limited, need to fit larger models, or want faster inference. Supports INT8, NF4, FP4 formats, QLoRA training, and 8-bit optimizers. Works with HuggingFace Transformers.
日本語の概要は準備中です。原文の説明を表示しています。
This skill should be used at the start of any computationally intensive scientific task to detect and report available system resources (CPU cores, GPUs, memory, disk space). It creates a JSON file with resource information and strategic recommendations that inform computational approach decisions such as whether to use parallel processing (joblib, multiprocessing), out-of-core computing (Dask, Zarr), GPU acceleration (PyTorch, JAX), or memory-efficient strategies. Use this skill before running analyses, training models, processing large datasets, or any task where resource constraints matter.
日本語の概要は準備中です。原文の説明を表示しています。
This skill should be used at the start of any computationally intensive scientific task to detect and report available system resources (CPU cores, GPUs, memory, disk space). It creates a JSON file with resource information and strategic recommendations that inform computational approach decisions such as whether to use parallel processing (joblib, multiprocessing), out-of-core computing (Dask, Zarr), GPU acceleration (PyTorch, JAX), or memory-efficient strategies. Use this skill before running analyses, training models, processing large datasets, or any task where resource constraints matter.
日本語の概要は準備中です。原文の説明を表示しています。
Retrieve up-to-date library, framework, and project documentation using scripts + MCP tools (Context7, Context Hub) with intelligent fallback. Use this skill whenever the user or agent needs documentation for any library, framework, API, SDK, or internal project spec. Triggers on "docs for [X]", "how does [library] work", "find documentation", "API reference for", "look up [feature] in [library]", "latest docs", "what's the API for", "find our [internal spec]", or any request that requires current, accurate documentation rather than relying on training data. Always prefer this skill over raw WebSearch for documentation retrieval — it returns structured, context-efficient results.
日本語の概要は準備中です。原文の説明を表示しています。
Fine-tune large language models — data curation, SFT, parameter-efficient methods (LoRA), evaluation, and avoiding common training failures. Use when a base model needs task-specific behavior or domain knowledge.
日本語の概要は準備中です。原文の説明を表示しています。
Orientação especializada para fine-tuning rápido com Unsloth - treinamento 2-5x mais rápido, 50-80% menos memória, otimização LoRA/QLoRA
日本語の概要は準備中です。原文の説明を表示しています。
This skill should be used at the start of any computationally intensive scientific task to detect and report available system resources (CPU cores, GPUs, memory, disk space). It creates a JSON file with resource information and strategic recommendations that inform computational approach decisions such as whether to use parallel processing (joblib, multiprocessing), out-of-core computing (Dask, Zarr), GPU acceleration (PyTorch, JAX), or memory-efficient strategies. Use this skill before running analyses, training models, processing large datasets, or any task where resource constraints matter.
日本語の概要は準備中です。原文の説明を表示しています。
Quantizes LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss. Use when GPU memory is limited, need to fit larger models, or want faster inference. Supports INT8, NF4, FP4 formats, QLoRA training, and 8-bit optimizers. Works with HuggingFace Transformers.
日本語の概要は準備中です。原文の説明を表示しています。