本文へ移動
cccskills

「vllm」の検索結果

109 件 ・ 関連度順

概要と使いどころ

Decision SOP for serving LLMs with vLLM. Covers PagedAttention mental model, quantization/parallelism/batching tradeoffs, OOM triage, and when NOT to use vLLM. Activates when a coder-agent is choosing or tuning an inference engine, debugging vLLM throughput/latency/OOM, or comparing vLLM against TGI/SGLang/TensorRT-LLM/llama.cpp.

日本語の概要は準備中です。原文の説明を表示しています。

agentsope/SkillAlchemy4372026年10月9日 更新

Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs. Verify dispatch, supported configurations, and GPU numerical behavior. Use for attention adaptations outside vllm_metax/patch/; do not apply monkey-patch headers or patch audit requirements.

日本語の概要は準備中です。原文の説明を表示しています。

MetaX-MACA/vLLM-metax1802026年10月10日 更新

Audit and adapt monkey patches in vllm_metax/patch/ against a target upstream revision. Decide whether to retain, update, migrate, or remove each patch; maintain patch headers and explanations; validate changes and record evidence. Use only for compatibility work on this directory, not standalone attention backend, model, kernel, or other adaptations elsewhere in vllm_metax.

日本語の概要は準備中です。原文の説明を表示しています。

MetaX-MACA/vLLM-metax1802026年10月10日 更新

Review and adapt vllm_metax/registry registrations, quantization configurations, CustomOps and kernel dispatch against a target vLLM revision and installed MetaX APIs. Verify registration, inherited state, weight layouts and actual backend selection. Use for registry compatibility work, not standalone attention algorithms or monkey patches.

日本語の概要は準備中です。原文の説明を表示しています。

MetaX-MACA/vLLM-metax1802026年10月10日 更新

High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO training of large models (7B-70B+). Built on Ray, vLLM, ZeRO-3. 2× faster than DeepSpeedChat with distributed architecture and GPU resource sharing.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO training of large models (7B-70B+). Built on Ray, vLLM, ZeRO-3. 2× faster than DeepSpeedChat with distributed architecture and GPU resource sharing.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Pick the right serving container for a SageMaker model deployment and find its current image URI. Use this skill whenever about to deploy a model to a SageMaker endpoint and an image URI needs to be chosen — including when the user says "deploy this LLM", "host this HuggingFace model", "serve this fine-tuned model", "deploy this embedding model", "host a reranker", "serve a sentence-transformers model", or when about to hardcode any container URI in deployment code. HuggingFace-curated Deep Learning Containers are ALWAYS preferred: HuggingFace vLLM (LLMs and generative rerankers), HuggingFace vLLM-Omni (multimodal), TEI (embeddings/cross-encoder rerankers), HF Inference Toolkit (other transformers). Generic images (AWS vLLM, DJL-LMI, SGLang) are used only when no HuggingFace image is compatible — never merely because they carry a newer version. Never hardcode a container URI from memory and never default to TGI. Prevents stale-image failures and wrong-region URIs.

日本語の概要は準備中です。原文の説明を表示しています。

huggingface/skills1.1万2026年10月9日 更新

Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5592026年10月10日 更新

Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels. Establish quantization and cache differences, preserve upstream structure, and distinguish shared upstream bugs from adaptation defects. Use for model support work, not standalone monkey-patch or registry audits.

日本語の概要は準備中です。原文の説明を表示しています。

MetaX-MACA/vLLM-metax1802026年10月10日 更新

Establish the shared environment, source/runtime correspondence, target confirmation and validation evidence for MetaX vLLM upgrades. Use as the common prerequisite of the patch, attention, registry and model upgrade skills, or for an explicitly requested MetaX environment preflight; it does not perform domain-specific adaptations.

日本語の概要は準備中です。原文の説明を表示しています。

MetaX-MACA/vLLM-metax1802026年10月10日 更新

环境安装与准备:把部署环境从零安装就绪——mindiesd 编译安装(本地昇腾直装 / SSH 推远端容器 / Docker 镜像直装)与三方推理框架全栈安装(vLLM-Omni 源码构建、DiffSynth-Engine 部署、 LightX2V editable 部署),并负责模型权重确认与下载(下载前先确认远端是否已存在)。不含特性使能与验证 (framework-integration)与 profiling(profiling-collect);SSH 工具由 remote-access 提供。 当用户需要安装 MindIE-SD、源码构建/直装 vLLM-Omni 或 LightX2V(editable + PLATFORM=ascend_npu)、 或确认/下载模型权重时使用此技能; 即使用户只提到"把代码推到服务器""在容器里装 vllm 全栈""准备 lightx2v 调优环境"而未说昇腾, 只要上下文涉及环境安装与准备都应触发。由 model-auto-optimization S0 与 dev-workflow 部署阶段指引加载。

日本語の概要は準備中です。原文の説明を表示しています。

Ascend/MindIE-SD152026年10月10日 更新

Framework RLHF de alta performance com aceleração Ray+vLLM. Use para treinamento PPO, GRPO, RLOO, DPO de modelos grandes (7B-70B+). Construído em Ray, vLLM, ZeRO-3. 2× mais rápido que DeepSpeedChat com arquitetura distribuída e compartilhamento de recursos GPU.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

Serve LLMs com alta throughput usando PagedAttention do vLLM e continuous batching. Use ao fazer deploy de APIs LLM em produção, otimizar latência/throughput de inferência, ou servir modelos com memória GPU limitada. Suporta endpoints compatíveis com OpenAI, quantização (GPTQ/AWQ/FP8) e tensor parallelism.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

Pick the right serving container for a SageMaker model deployment and find its current image URI. Use this skill whenever about to deploy a model to a SageMaker endpoint and an image URI needs to be chosen — including when the user says "deploy this LLM", "host this HuggingFace model", "serve this fine-tuned model", "deploy this embedding model", "host a reranker", "serve a sentence-transformers model", or when about to hardcode any container URI in deployment code. HuggingFace-curated Deep Learning Containers are ALWAYS preferred: HuggingFace vLLM (LLMs and generative rerankers), HuggingFace vLLM-Omni (multimodal), TEI (embeddings/cross-encoder rerankers), HF Inference Toolkit (other transformers). Generic images (AWS vLLM, DJL-LMI, SGLang) are used only when no HuggingFace image is compatible — never merely because they carry a newer version. Never hardcode a container URI from memory and never default to TGI. Prevents stale-image failures and wrong-region URIs.

日本語の概要は準備中です。原文の説明を表示しています。

bg-szy/TOP-SKILLS62026年9月8日 更新

openrlhf

無料

High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO training of large models (7B-70B+). Built on Ray, vLLM, ZeRO-3. 2× faster than DeepSpeedChat with distributed architecture and GPU resource sharing.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

vllm

無料

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

vLLM: high-throughput LLM serving, OpenAI API, quantization.

日本語の概要は準備中です。原文の説明を表示しています。

NousResearch/hermes-agent25.3万2026年10月11日 更新

outlines

無料

Guarantee valid JSON/XML/code structure during generation, use Pydantic models for type-safe outputs, support local models (Transformers, vLLM), and maximize inference speed with Outlines - dottxt.ai's structured generation library

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM. Use for provider attachment, endpoint policy, credential substitution, topology, and migration from the removed managed inference endpoint. Trigger keywords - debug inference, managed inference endpoint, local inference, ollama, lm studio, vllm, sglang, trtllm, NIM, inference failing, model server unreachable, credential_endpoint_mismatch, host.openshell.internal.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/OpenShell1.6万2026年10月10日 更新

outlines

無料

Guarantee valid JSON/XML/code structure during generation, use Pydantic models for type-safe outputs, support local models (Transformers, vLLM), and maximize inference speed with Outlines - dottxt.ai's structured generation library

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Deploy and manage vLLM for high-throughput LLM inference. Configure continuous batching, tensor parallelism, quantization, and OpenAI-compatible API endpoints for production LLM serving.

日本語の概要は準備中です。原文の説明を表示しています。

BagelHole/DevOps-Security-Agent-Skills1,1542026年5月22日 更新

Neutral, framework-agnostic decision tree for project kickoff: "which agent / RAG / LLM framework should I reach for?" Synthesizes the ecosystem sections of 7 landmark-project SOPs (LangGraph, LlamaIndex, DSPy, CrewAI, vLLM, Aider, Dify) into one layered rubric. Core stance: frameworks are LAYERS, not competitors — a real project usually combines DSPy (compile) + LlamaIndex (retrieve) + LangGraph (orchestrate) + vLLM (serve), and you choose ONE per layer, not one to rule all. Use when starting any LLM/agent/RAG project, or whenever the "which framework?" question is asked. Deliberately neutral — unlike vendor docs and the LangChain-biased `framework-selection` on skill.sh, this skill has no horse in the race.

日本語の概要は準備中です。原文の説明を表示しています。

agentsope/SkillAlchemy4372026年10月9日 更新