Generalised autonomous optimisation loop — soft RLVR for any artifact a user can measure. Use this skill whenever a user wants to iteratively improve an artifact — code, prompts, documents, configs, designs, content — by running structured experiments, evaluating results against a multi-dimensional rubric, and learning from each attempt. Triggers include: "optimise this", "keep improving until it's good", "run experiments on", "autoresearch", "iterate on this overnight", "try different approaches and pick the best", or any request implying repeated evaluate-and-improve cycles. Also use when the user wants to improve a system prompt, a data pipeline, a writing style, or any artifact where quality can be decomposed into measurable tracks. For inference optimisation tasks (model latency, throughput, quantization, GPU deployment), a* delegates the low-level tuning to AITune while maintaining quality tracking and learning.
日本語の概要は準備中です。原文の説明を表示しています。
chrisvoncsefalvay/autostar☆ 392026年4月10日 更新
量化专家经验调优库:按「L1 通用专家调优意见 → L2 结构化专家调优意见 → L3 模型专属专家调优意见」三级递进, 输出量化精度调优的完整手段(离群值抑制 / 量化方法·粒度·对称性选型 / 校准集调整 / 敏感层回退 / 模型专属策略), 每条结论附专家原因分析与专家意见可信度等级。 本 Skill 是专家经验库,负责说明可选离群值抑制算法并询问用户选择哪些算法,以及回答「怎么调、哪些层需要回退及为什么」;不运行离群值抑制 processor,不修改 YAML、不执行量化、不做 EP 检查 / 服务化 / 任务评测。
日本語の概要は準備中です。原文の説明を表示しています。
kali20gakki/msAgent☆ 322026年10月10日 更新
为 msModelSlim 适配器实现逐层量化(按层加载/懒加载)能力。仅在用户明确要求逐层量化或基础适配因 CPU 内存不足无法全量加载权重时使用。该特性为高阶可选项,不是基础适配必需项。
日本語の概要は準備中です。原文の説明を表示しています。
kali20gakki/msAgent☆ 322026年10月10日 更新
Qdrant vector database -- collection management, point operations, payload filtering, named vectors, quantization, recommendations, snapshots
日本語の概要は準備中です。原文の説明を表示しています。
agents-inc/skills☆ 242026年9月8日 更新
Pulp musical/media time primitives, exact beat divisions, tempo and meter maps, transport-range grid projection, inline and order-preserving groove projection, coordinate randomness, streaming cursors, and quantization arithmetic.
日本語の概要は準備中です。原文の説明を表示しています。
Generous-Corp/pulp☆ 222026年10月12日 更新
Tune moflo's memory stack for speed, RAM, and index quality. Covers HNSW parameters (M, efConstruction, ef), vector quantization, batch operations, and common bottlenecks. Use when scaling past ~100k entries or when search latency regresses.
日本語の概要は準備中です。原文の説明を表示しています。
eric-cielo/moflo☆ 182026年10月1日 更新
Iterate on a GenAI notebook against the self-hosted stack via the genai-stack CLI (config dirs, auth, subdomains, quantization, GPU/VRAM). Arguments: <notebook|service> [--service comfyui|forge|vllm] [--quant int4|fp8] [--validate] [--bg]
日本語の概要は準備中です。原文の説明を表示しています。
jsboige/CoursIA☆ 162026年10月12日 更新
量化格式契约与位级对齐:从设备字节反推量化器的**编码公式、舍入模式、scale 粒度、退化块规则**, 并据此重实现/融合量化器(MXFP8 e8m0、int8 perblock、fp8 perchannel 等), 用「同进程同数据 + 独立参考实现 + 中点与退化输入全覆盖」做到**逐字节精确**。 当自研量化 kernel 与框架算子**对不上**、需要逆向某个量化器的数值契约、需要判断 「量化残差算不算可接受」、或要在融合内核里复现 aclnn 量化语义时使用此 skill。 即使用户只说「量化结果和参考不一致」「这个 scale 怎么算出来的」「融合量化器精度对不上」 而未提"契约",也应触发。**注意**:单纯选量化档位/配置(该不该开量化、开哪一档)走 dit-perf-opt;算子级性能调优与 DSL 选型走 operator-dev; 并行作用域导致的静默失效(DiT 侧「改了并行但不报错也没生效」的判定)走 dit-parallel-opt。
日本語の概要は準備中です。原文の説明を表示しています。
Ascend/MindIE-SD☆ 152026年10月11日 更新
DiT 计算模块(L3):把**已定位的 DiT 计算瓶颈**落成特性级选档与实施—— 量化档(W8A16 / W4A16 / W8A8 系列 / W4A4 / MXFP8 / FA 量化)、稀疏(rf_v2 / ada_bsa)、 缓存(DiTCache / AttentionCache / 时间步优化)、编译启用(MindieSDBackend / Pattern 融合 / ACLGraph) 的**开不开、开哪一档、怎么开、怎么复验**;依据是 `docs/zh/features/*`(特性真源)+ framework-integration/references/framework-support-matrix.md(支持状态)。 即使用户只说"这个模型怎么加速""量化/稀疏/Cache 怎么选怎么开""要不要开量化、开哪一档""这个档位开了有没有效果" 而未提 profiling,也应触发。 **入口条件**:瓶颈点已明确(用户带一句实测锚点,或编排层交付标签)时由域入口 `performance-optimization` 按标签分发到本技能;**瓶颈未明("怎么加速 / 跑通 / 采 profile")先走 `model-auto-optimization` 定位**,不在本技能内做占比分析。 near-miss:多卡并行形态 / 通信掩盖 / TP·offload 选型 → `dit-parallel-opt`;VAE 解码段与 host 固定开销 → 各自模块(VAE / host);单算子实现级实测选型(mindie_bench)→ `benchmark-dev`; 需要新增 pattern / 算子才能落地本档 → `pattern-dev` / `operator-dev`;框架侧开关与使能验证 → `framework-integration`;量化器位级契约与精度对齐(编码公式 / 舍入 / scale 粒度)→ `quantization-dev`;精度验收判据 → `accuracy-gate`;数字入库口径 → `perf-gate`。
日本語の概要は準備中です。原文の説明を表示しています。
Ascend/MindIE-SD☆ 152026年10月11日 更新
Optimize AgentDB performance with quantization (4-32x memory reduction), HNSW indexing (150x faster search), caching, and batch operations. Use when optimizing memory usage, improving search speed, or scaling to millions of vectors.
日本語の概要は準備中です。原文の説明を表示しています。
lilinji/GeneTind-Life-Skills☆ 142026年8月21日 更新
vLLM: high-throughput LLM serving, OpenAI API, quantization.
日本語の概要は準備中です。原文の説明を表示しています。
kevinnft/ai-agent-skills☆ 132026年8月1日 更新
Serve LLMs com alta throughput usando PagedAttention do vLLM e continuous batching. Use ao fazer deploy de APIs LLM em produção, otimizar latência/throughput de inferência, ou servir modelos com memória GPU limitada. Suporta endpoints compatíveis com OpenAI, quantização (GPTQ/AWQ/FP8) e tensor parallelism.
日本語の概要は準備中です。原文の説明を表示しています。
artubss/SKILLS-CLAUDE-CODE☆ 112026年5月17日 更新
Executa inferência de LLM em CPU, Apple Silicon e GPUs consumer sem hardware NVIDIA. Use para edge deployment, Macs M1/M2/M3, GPUs AMD/Intel ou quando CUDA não está disponível. Suporta quantização GGUF (1,5-8 bits) para redução de memória e aceleração de 4-10× vs PyTorch em CPU.
日本語の概要は準備中です。原文の説明を表示しています。
artubss/SKILLS-CLAUDE-CODE☆ 112026年5月17日 更新
Quantiza LLMs para 8-bit ou 4-bit com redução de memória de 50-75% e perda mínima de acurácia. Use quando a memória GPU é limitada, precisa ajustar modelos maiores ou quer inferência mais rápida. Suporta formatos INT8, NF4, FP4, treinamento QLoRA e otimizadores 8-bit. Funciona com HuggingFace Transformers.
日本語の概要は準備中です。原文の説明を表示しています。
artubss/SKILLS-CLAUDE-CODE☆ 112026年5月17日 更新
vLLM: high-throughput LLM serving, OpenAI API, quantization.
日本語の概要は準備中です。原文の説明を表示しています。
openamer/openamer☆ 62026年10月12日 更新
This work studies the post-training quantization of LLM parameters as a way to improve their runtime GAP: multi-agent orchestration: competitor signal
日本語の概要は準備中です。原文の説明を表示しています。
openamer/openamer☆ 62026年10月12日 更新
Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio. Expert in quantization formats (GGUF, EXL2) and local AI privacy.
日本語の概要は準備中です。原文の説明を表示しています。
bouclem/skills☆ 62026年5月31日 更新
Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio. Expert in quantization formats (GGUF, EXL2) and local AI privacy.
日本語の概要は準備中です。原文の説明を表示しています。
hybridlabor-api/aos☆ 62026年10月11日 更新
Integrate on-device AI using Foundation Models framework, Core ML, and open-source LLM runtimes on Apple Silicon. Covers Foundation Models (LanguageModelSession, @Generable, @Guide, SystemLanguageModel, structured output, tool calling), Core ML (coremltools, model conversion, quantization, palettization, pruning, Neural Engine, MLTensor), MLX Swift (transformer inference, unified memory), and llama.cpp (GGUF, cross-platform LLM). Use when building tool-calling AI features, working with guided generation schemas, converting models, or running on-device inference.
日本語の概要は準備中です。原文の説明を表示しています。
JordanCoin/ios-skills-collection☆ 62026年9月10日 更新
Optimize AgentDB performance with quantization (4-32x memory reduction), HNSW indexing (150x faster search), caching, and batch operations. Use when optimizing memory usage, improving search speed, or scaling to millions of vectors.
日本語の概要は準備中です。原文の説明を表示しています。
lucaspmarie-a11y/claude-skills-vault☆ 52026年6月7日 更新
Quantizes LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss. Use when GPU memory is limited, need to fit larger models, or want faster inference. Supports INT8, NF4, FP4 formats, QLoRA training, and 8-bit optimizers. Works with HuggingFace Transformers.
日本語の概要は準備中です。原文の説明を表示しています。
huang-sh/DeepScience☆ 42026年7月15日 更新
Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.
日本語の概要は準備中です。原文の説明を表示しています。
huang-sh/DeepScience☆ 42026年7月15日 更新
Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.
日本語の概要は準備中です。原文の説明を表示しています。
ibragimov-oasis/vibe-coder☆ 22026年6月24日 更新
Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. Use for edge deployment, M1/M2/M3 Macs, AMD/Intel GPUs, or when CUDA is unavailable. Supports GGUF quantization (1.5-8 bit) for reduced memory and 4-10× speedup vs PyTorch on CPU.
日本語の概要は準備中です。原文の説明を表示しています。
ibragimov-oasis/vibe-coder☆ 22026年6月24日 更新