本文へ移動
cccskills

「inference optimization」の検索結果

55 件 ・ 関連度順

概要と使いどころ

Use when working with Baseten — baseten ML deployment platform management covering model inventory, deployment status, autoscaling configuration, inference call history, GPU allocation, environment management, and performance metrics. Use for comprehensive Baseten workspace assessment and ML serving optimization.

日本語の概要は準備中です。原文の説明を表示しています。

cloudthinker-ai/CloudSkills62026年4月5日 更新

Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio. Expert in quantization formats (GGUF, EXL2) and local AI privacy.

日本語の概要は準備中です。原文の説明を表示しています。

bouclem/skills62026年5月31日 更新

Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio. Expert in quantization formats (GGUF, EXL2) and local AI privacy.

日本語の概要は準備中です。原文の説明を表示しています。

hybridlabor-api/aos62026年10月8日 更新

hqq

無料

Half-Quadratic Quantization for LLMs without calibration data. Use when quantizing models to 4/3/2-bit precision without needing calibration datasets, for fast quantization workflows, or when deploying with vLLM or HuggingFace Transformers.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace. Use when training large-scale models with limited compute (5× cost reduction vs dense models), implementing sparse architectures like Mixtral 8x7B or DeepSeek-V3, or scaling model capacity without proportional compute increase. Covers MoE architectures, routing mechanisms, load balancing, expert parallelism, and inference optimization.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

Master TypeScript with advanced types, generics, and strict type safety. Handles complex type systems, decorators, and enterprise-grade patterns. Use PROACTIVELY for TypeScript architecture, type inference optimization, or advanced typing patterns.

日本語の概要は準備中です。原文の説明を表示しています。

itsimonfredlingjack/codex-dev-plugin22026年2月5日 更新

Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks. Implements model serving, feature engineering, A/B testing, and monitoring. Use PROACTIVELY for ML model deployment, inference optimization, or production ML infrastructure.

日本語の概要は準備中です。原文の説明を表示しています。

itsimonfredlingjack/codex-dev-plugin22026年2月5日 更新