本文へ移動
cccskills

「sglang」の検索結果

59 件 ・ 関連度順

概要と使いどころ

Workflow for upgrading/integrating cache-dit in SGLang diffusion (multimodal_gen): DBCache, DMD calibrator, TaylorSeer, SVDQuant DQ; porting upstream PRs and resolving conflicts against the per-request knob system; adding new cache knobs; building the sglang generate CLI test matrix; precision validation (PSNR / log evidence); troubleshooting environment issues (wheel ABI, svdq extension, flashinfer conflicts). Use when upgrading or integrating cache-dit in sglang diffusion, porting cache-dit PRs with conflicts, adding cache knobs, running the sglang generate CLI test matrix, or validating precision (PSNR) for DBCache/DMD/SVDQuant(DQ) paths.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Add or verify a model's chat prompt rendering and tokenization in rust/sglang-processor so it matches SGLang's Python serving path exactly (same prompt text, same token ids). Use when adding a model to sglang-processor, porting a model's rendering from sgl-router, bumping Dynamo's renderer, or debugging a processor-vs-Python prompt mismatch.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate. Use when adding, renaming, or reviewing any `SGLANG_*` environment variable (or migrating a legacy `SGL_*` alias), or when touching `python/sglang/srt/environ.py`.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

How SGLang's runtime configuration and process-global state are organized (RuntimeContext tiers, publish + namespace config bags, the pristine ServerArgs seed, override entry points, resource/stream/buffer leases, per-forward flags), the CI guardrails that enforce the design, and the idioms for developing and testing against it. Load this before touching server_args, model overrides, module-level state, or per-forward state in sglang.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Guide for writing SGLang CI/UT tests. Covers CustomTestCase, CI registration, server fixtures, model selection, mock testing, and test placement. Always read test/README.md for the full CI layout, how to run tests, and extra tips. Use when creating new tests, adding CI test cases, writing unit tests, or when the user asks to add tests for SGLang features.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Naming conventions for SGLang speculative decoding identifiers. Use when adding, renaming, or reviewing identifiers in speculative decoding code — anything under `python/sglang/srt/speculative/`, related attention backends, scheduler accumulators, IPC fields, observability metrics, or CLI flags.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Trigger the bot-cherry-pick workflow for a batch of merged PRs onto a release branch and monitor each run to completion. Use when an SGLang release manager asks to cherry-pick a list of PRs to a release branch.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Use when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Use when adding a new diffusion model or Diffusers pipeline to SGLang.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). Covers identifying hang locations via py-spy/watchdog/cuda coredump, per-rank logging to find state divergence, binary-search methodology for locating the first diverge point, and fix patterns. Use when a multi-GPU SGLang run hangs, freezes, or times out during collective operations.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Use when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Apply the SGLang kernels RFC when adding, moving, splitting, or reviewing kernel APIs, registry metadata, kernel tests, benchmarks, and model-specific implementations. Use with add-jit-kernel, add-sgl-kernel, and write-sglang-test for placement and migration checks.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Investigate consistently failing SGLang CI tests by extracting the failure signature from scheduled or rerun workflows, bisecting the passing/failing commit window, checking runner or hardware specificity, and optionally reproducing on a remote GPU host.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Replay-first debug flow for SGLang serving problems. Use when a live or recent server shows health-check failures, latency or throughput regressions, queue growth, timeouts, distributed stalls, crash dumps, wrong outputs after deploys, or PD/EP/HiCache issues, and the job is to turn the problem into a replay plus the right next debug tool.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Use when choosing the fastest SGLang Diffusion flags for a model, GPU, and VRAM budget.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Write, calibrate, and debug the prefill-vs-decode logprob (KL) consistency tests in sglang -- the two independent conditions a zero requires (every operator batch-invariant, and the two paths computing the same function), which helper separates them, how to pick a threshold once they hold, and how to localize a divergence to a single operator. Use when adding a KL test to a model, picking or defending a kl_div threshold, or investigating a KL number that is too high.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Code style for SGLang large classes `Scheduler`, `TokenizerManager`, and `ModelRunner`: frozen-code conventions and `__init__` orchestration style. Use when modifying any of these three classes or reviewing changes to them.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Unified LLM torch-profiler triage skill for `sglang`, `vllm`, `TensorRT-LLM`, and `TokenSpeed`. Use it to inspect an existing `trace.json(.gz)` or profile directory, or to drive live profiling against a running server when supported and return one three-table report with kernel, overlap-opportunity, and fuse-pattern tables.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Requirements for the SGLang scripted runtime, chiefly when to add (vs not add) a harness API. Use for anything related to the scripted runtime.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Audit SGLang startup logs, save evidence, and propose cleanup for user review. With no arguments, run Qwen3-8B at TP1 and TP2 plus gpt-oss-20b at TP1.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Review a pull request against the SGLang Cookbook (docs/, Mintlify) contribution checklist — the config-driven format (per-model config + benchmarks JSX consumed by the shared _deployment.jsx / _playground.jsx engines). Run with /cookbook-review-pr <PR number>.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Add a new model to the SGLang Cookbook (docs/, Mintlify), config-driven format — instantiate the model-agnostic template into a per-model config (+ benchmarks) JSX under src/snippets/configs/, an MDX page, the docs.json nav entry, NEW-tag hygiene, and the homepage vendor card. Interactive, multi-phase. Run with /cookbook-add-model.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head. Use when asked to monitor, babysit, retry, or fix PR CI for lint.yml, pr-test.yml, pr-test-extra.yml, AMD, or other named workflows; classify failures as PR-related versus flaky or infrastructural, auto-fix and push only small clean fixes, rerun failed jobs only up to 10 times, and ignore unselected workflows.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新