PyTorchのモデルや学習コードを作成・点検し、データ読み込み、実験の再現性、学習状態の保存、GPUメモリや処理速度の改善を支援するスキル。
- PyTorchの学習コード作成・レビュー
- 実験結果が揃わない原因を調べたいとき
- 読み込み速度やGPUメモリの改善
209 件 ・ 関連度順
概要と使いどころ
PyTorchのモデルや学習コードを作成・点検し、データ読み込み、実験の再現性、学習状態の保存、GPUメモリや処理速度の改善を支援するスキル。
Generate C/C++ or CUDA code from an AI model (PyTorch, LiteRT) using MATLAB Coder or GPU Coder. Use when the user wants to integrate an AI model into an application with code generation as the end goal — generating MEX, CUDA MEX, static library, dynamic library, or executable — or using the model in Simulink for simulation and code generation. Covers PyTorch ExportedProgram (.pt2) via loadPyTorchExportedProgram and LiteRT (.tflite) via loadLiteRTModel (R2026a+). Keywords: PyTorch, torch, .pt2, ExportedProgram, loadPyTorchExportedProgram, invoke, codegen, MEX, CUDA, GPU, C, C++, deploy, AI model, deep learning model, LiteRT, TFLite, TensorFlow Lite, Simulink, slbuild, PyTorch ExportedProgram block, MATLAB Function block, dlosslib, loadLiteRTModel.
日本語の概要は準備中です。原文の説明を表示しています。
長い入力を扱うTransformerの学習・推論で、注意機構の計算を高速化し、GPUメモリ使用量を抑えるスキル。導入から性能測定、出力の比較まで案内します。
Convert existing plain or manual PyTorch training code into an NVFLARE federated job using Client API model exchange, local validation, and job export; use when the user names plain PyTorch or preliminary source inspection identifies one plain-PyTorch owner, and not for Lightning, other frameworks, deployment, POC/production lifecycle, or experiment workflows.
日本語の概要は準備中です。原文の説明を表示しています。
Convert existing PyTorch Lightning training code into an NVFLARE federated job using the Lightning Client API patch, local validation, and job export; use only when the request names federated/NVFLARE conversion or asks multiple sites to train collaboratively while keeping each site's data local, and either names PyTorch Lightning or preliminary source inspection identifies one Lightning owner; do not use for non-federated Lightning work such as DDP, profiling, inference serving, or training-loop changes, nor for plain PyTorch, TensorFlow/Keras, other frameworks, deployment, POC/production lifecycle, or experiment workflows.
日本語の概要は準備中です。原文の説明を表示しています。
Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder). Covers two workflow patterns: (1) MathWorks-native or imported models rebuilt as dlnetwork for lean hardware, (2) direct C/C++ code generation from PyTorch and LiteRT models. Both patterns support all targets (Cortex-M/A/R, x86, GPU). Trigger when: user wants to deploy AI to embedded targets; generate C/CUDA from neural networks; compress AI models for MCU; integrate AI in Simulink for system-level simulation; import PyTorch/ONNX/TensorFlow models for embedded deployment; optimize AI for resource-constrained hardware; or use loadPyTorchExportedProgram, loadLiteRTModel, importNetworkFromPyTorch, importNetworkFromONNX, importNetworkFromTensorFlow, importNetworkFromKeras, dlquantizer, exportNetworkToSimulink, or Embedded Coder with AI models.
日本語の概要は準備中です。原文の説明を表示しています。
Import PyTorch, ONNX, or Keras 3 / TensorFlow 2.16+ deep learning models into MATLAB as dlnetwork objects. Use when importing .pt2 exported programs, traced .pt files, .onnx models, or Keras 3 models via matlabsaver. Covers importNetworkFromPyTorch, importNetworkFromONNX, importNetworkFromKeras, importNetworkFromTensorFlow, torch.export.export, PyTorchInputSizes, InputDataFormats, matlabsaver, tf_keras downgrade, numeric validation against PyTorch or ONNX Runtime, and placeholder/custom layer implementation. Applies when user mentions any of these functions, file formats, or encounters import errors, unsupported operator warnings, 0 learnables, or uninitialized networks.
日本語の概要は準備中です。原文の説明を表示しています。
Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder). Covers two workflow patterns: (1) MathWorks-native or 3P-imported models rebuilt as dlnetwork for lean hardware (Cortex-M, DSP), (2) direct C/C++ code generation from PyTorch and LiteRT models for high-performance hardware (Cortex-A, x86, GPU). Trigger when: user wants to deploy AI to embedded targets; generate C/CUDA from neural networks; compress AI models for MCU/DSP; integrate AI in Simulink for system-level simulation; import PyTorch/ONNX/TensorFlow models for embedded deployment; optimize AI for resource-constrained hardware; or use loadPyTorchExportedProgram, importNetworkFromPyTorch, dlquantizer, exportNetworkToSimulink, or Embedded Coder with AI models.
日本語の概要は準備中です。原文の説明を表示しています。
面向 Ascend PyTorch Profiler / msprof DB(如 ascend_pytorch_profiler*.db、msprof_*.db)的 SQL 分析技能。将自然语言问题(算子耗时、通信、下发、调度、schema/table 查询)转为安全可执行 SQL,并按需从官方文档提取表结构详情。
日本語の概要は準備中です。原文の説明を表示しています。
Deep learning framework (PyTorch Lightning / lightning package). Organize PyTorch code into LightningModules, configure Trainers for multi-GPU/TPU, implement data pipelines, callbacks, logging (W&B, TensorBoard, MLflow), distributed training (DDP, FSDP, DeepSpeed), for scalable neural network training.
日本語の概要は準備中です。原文の説明を表示しています。
Expert guidance for Fully Sharded Data Parallel training with PyTorch FSDP - parameter sharding, mixed precision, CPU offloading, FSDP2
日本語の概要は準備中です。原文の説明を表示しています。
High-level PyTorch framework with Trainer class, automatic distributed training (DDP/FSDP/DeepSpeed), callbacks system, and minimal boilerplate. Scales from laptop to supercomputer with same code. Use when you want clean training loops with built-in best practices.
日本語の概要は準備中です。原文の説明を表示しています。
Fine-tune and serve Physical Intelligence OpenPI models (pi0, pi0-fast, pi0.5) using JAX or PyTorch backends for robot policy inference across ALOHA, DROID, and LIBERO environments. Use when adapting pi0 models to custom datasets, converting JAX checkpoints to PyTorch, running policy inference servers, or debugging norm stats and GPU memory issues.
日本語の概要は準備中です。原文の説明を表示しています。
High-level PyTorch framework with Trainer class, automatic distributed training (DDP/FSDP/DeepSpeed), callbacks system, and minimal boilerplate. Scales from laptop to supercomputer with same code. Use when you want clean training loops with built-in best practices.
日本語の概要は準備中です。原文の説明を表示しています。
Adds PyTorch FSDP2 (fully_shard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing. Use when models exceed single-GPU memory or when you need DTensor-based sharding with DeviceMesh.
日本語の概要は準備中です。原文の説明を表示しています。
Deep learning framework (PyTorch Lightning). Organize PyTorch code into LightningModules, configure Trainers for multi-GPU/TPU, implement data pipelines, callbacks, logging (W&B, TensorBoard), distributed training (DDP, FSDP, DeepSpeed), for scalable neural network training.
日本語の概要は準備中です。原文の説明を表示しています。
Write docstrings for PyTorch functions and methods following PyTorch conventions. Use when writing or updating docstrings in PyTorch code.
日本語の概要は準備中です。原文の説明を表示しています。
Route DRL-Pytorch reinforcement-learning workflows for standalone PyTorch Q-learning, DQN-family, PPO, DDPG, TD3, SAC, Atari DQN, and Actor-Sharer-Learner scripts.
日本語の概要は準備中です。原文の説明を表示しています。
Use denoising-diffusion-pytorch for PyTorch DDPM/DDIM image diffusion, 1D sequence diffusion, conditioning and guidance, and advanced diffusion variants.
日本語の概要は準備中です。原文の説明を表示しています。
Use audio-diffusion-pytorch for PyTorch waveform diffusion generators, text-conditioned audio generation, inpainting, upsampling, vocoding, and diffusion autoencoding.
日本語の概要は準備中です。原文の説明を表示しています。
Framework PyTorch de alto nível com classe Trainer, treinamento distribuído automático (DDP/FSDP/DeepSpeed), sistema de callbacks e boilerplate mínimo. Escala de laptop para supercomputador com o mesmo código. Use quando quiser loops de treinamento limpos com boas práticas integradas.
日本語の概要は準備中です。原文の説明を表示しています。
Orientação especializada para treinamento com Fully Sharded Data Parallel do PyTorch (FSDP) - sharding de parâmetros, precisão mista, CPU offloading, FSDP2
日本語の概要は準備中です。原文の説明を表示しています。
High-level PyTorch framework with Trainer class, automatic distributed training (DDP/FSDP/DeepSpeed), callbacks system, and minimal boilerplate. Scales from laptop to supercomputer with same code. Use when you want clean training loops with built-in best practices.
日本語の概要は準備中です。原文の説明を表示しています。
Expert guidance for Fully Sharded Data Parallel training with PyTorch FSDP - parameter sharding, mixed precision, CPU offloading, FSDP2
日本語の概要は準備中です。原文の説明を表示しています。