本文へ移動
cccskills
無料GitHub で公開

audio-diffusion-pytorch

Use audio-diffusion-pytorch for PyTorch waveform diffusion generators, text-conditioned audio generation, inpainting, upsampling, vocoding, and diffusion autoencoding.

インストール方法を見る

含まれるファイル(15)

  • SKILL.md3.7 KB
  • references/repo-provenance.md1.7 KB
  • references/repo-routing-metadata.json443 B
  • references/troubleshooting.md4.1 KB
  • scripts/check_install.py4.2 KB
  • sub-skills/conditioning/references/api-reference.md6.9 KB
  • sub-skills/conditioning/references/troubleshooting.md3.5 KB
  • sub-skills/conditioning/references/workflows.md4.2 KB
  • sub-skills/conditioning/scripts/tiny_conditioning_smoke.py5.6 KB
  • sub-skills/conditioning/SKILL.md1.5 KB
  • sub-skills/generation/references/api-reference.md4.1 KB
  • sub-skills/generation/references/troubleshooting.md3.2 KB
  • sub-skills/generation/references/workflows.md3.2 KB
  • sub-skills/generation/scripts/tiny_generation_smoke.py4.5 KB
  • sub-skills/generation/SKILL.md1.8 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

audio-diffusion-pytorch

Use this repo skill when a task involves the audio-diffusion-pytorch package or asks for PyTorch audio diffusion model setup, waveform generation, diffusion upsampling, mel vocoding, inpainting, or audio autoencoding.

This package provides building blocks and wrappers; it does not ship pretrained weights, ready-to-run checkpoints, or guaranteed Moûsai paper configs. Treat examples as model-construction and shape recipes unless the user supplies weights, data, or a training plan.

Install and quick checks

Install the public package:

pip install audio-diffusion-pytorch

Minimal import check:

python - <<'PY'
from importlib.metadata import version
import audio_diffusion_pytorch
print(version("audio-diffusion-pytorch"))
print("import ok")
PY

Optional dependencies:

  • Text conditioning uses the default a-unet T5 embedder and requires transformers. First use may consult Hugging Face cache or network.
  • README-style autoencoder examples may use audio_encoders_pytorch and auraloss, but the core DiffusionAE wrapper can also work with a local encoder object.
  • CUDA is optional for this package. CPU is enough for tiny smoke checks; use a CUDA-capable PyTorch install only when the user wants GPU execution.

Run scripts/check_install.py to report installed versions, optional modules, and public signatures. Use --check-cuda only when you want a tiny CUDA allocation.

Route map

  • Use sub-skills/generation/SKILL.md for DiffusionModel, UNetV0, VDiffusion, VSampler, text-conditioned generation, VInpainter, schedules, distributions, and expert DiffusionAR notes.
  • Use sub-skills/conditioning/SKILL.md for DiffusionUpsampler, DiffusionVocoder, DiffusionAE, EncoderBase, AdapterBase, mel spectrogram conditioning, and transform plugins.
  • Use references/troubleshooting.md for install/import issues, optional dependencies, backend questions, no-pretrained-weights expectations, and cross-cutting shape gotchas.
  • Use references/repo-provenance.md before deciding whether this skill is current for a checkout or should be refreshed.

Common decisions

  1. Identify the workflow family first:
    • new waveform generation, text prompts, sampler errors, or masks → generation;
    • lower-rate waveform conditioning, mel spectrograms, latents, encoders, adapters, or custom losses → conditioning.
  2. Keep smoke tests tiny: use channel counts and lengths from the bundled scripts before scaling to README-sized tensors.
  3. Set resnet_groups=1 for tiny channels, or use channel widths divisible by the default resnet_groups=8.
  4. Keep tensors on one device. Move the model and all inputs to CUDA only after CPU shape checks pass.
  5. Do not promise audio quality from random weights. Sampling APIs validate shape and execution, not trained generation quality.

Avoid this skill when

  • The user wants a pretrained audio diffusion checkpoint, dataset download, or benchmark reproduction and has not supplied the missing assets.
  • The task is about non-PyTorch audio libraries, ASR/TTS pipelines unrelated to diffusion/vocoding, or image diffusion.
  • The user is editing this repository's release workflow rather than using the package APIs.

Maintenance note

This skill is self-contained for operating use. If a future checkout changes package metadata, public constructors, README workflows, or source roots, run refresh-repo-skill instead of patching this skill ad hoc.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Routes 3D ResNets PyTorch video action-recognition workflows across training, inference, and data preparation.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3322026年9月3日 更新

3ddfa

無料

Guide 3DDFA Python inference, geometry rendering, training/evaluation, and optional C++ ONNX workflows for 3D dense face alignment.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3322026年9月3日 更新

3ddfa-v2

無料

Routes 3DDFA_V2 face-alignment setup, still-image demos, video tracking, and ONNX benchmarking workflows.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3322026年9月3日 更新

ab3dmot

無料

Operate AB3DMOT 3D multi-object tracking workflows for KITTI and nuScenes data, tracking, evaluation, and visualization.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3322026年9月3日 更新

Use Hugging Face Accelerate for PyTorch training-loop migration, distributed launch/configuration, DeepSpeed/FSDP/TPU backend setup, big-model inference/offload, checkpointing, tracking, and troubleshooting.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3322026年9月3日 更新

acme

無料

Route Acme reinforcement-learning framework tasks across core loops, replay/data, JAX agents, and TensorFlow/Sonnet agents.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3322026年9月3日 更新

VectorSpaceLab のスキルをすべて見る

このスキルの問題を報告する