本文へ移動
cccskills

「llama.cpp」の検索結果

35 件 ・ 関連度順

概要と使いどころ

GGUF format and llama.cpp quantization for efficient CPU/GPU inference. Use when deploying models on consumer hardware, Apple Silicon, or when needing flexible quantization from 2-8 bit without GPU requirements.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

GGUF format and llama.cpp quantization for efficient CPU/GPU inference. Use when deploying models on consumer hardware, Apple Silicon, or when needing flexible quantization from 2-8 bit without GPU requirements.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

llama-cpp

無料

Run GGUF models directly, load LoRA adapters, benchmark inference speed, and serve models via llama-server using llama.cpp. Includes Qwen 3.5 serve scripts (9B dense + F16, 35B MoE) with asymmetric KV cache and thinking mode. Secondary to Ollama; use when needing direct model control or LoRA hot-loading. Triggers on 'llama.cpp', 'GGUF', 'LoRA adapter', 'benchmark inference', 'llama-server'.

日本語の概要は準備中です。原文の説明を表示しています。

tdimino/claude-code-minoan412026年9月28日 更新

Analyze and tune a running llama.cpp server for local coding agents: propose config (ctx, prompt cache, MTP), benchmark before/after, report deltas. Use to optimize llama-server for Claude Code or Aider. Don't use for vLLM, ollama, or non-llama.cpp.

日本語の概要は準備中です。原文の説明を表示しています。

luongnv89/ccl392026年9月8日 更新

Formato GGUF e quantização llama.cpp para inferência eficiente em CPU/GPU. Use ao implantar modelos em hardware de consumidor, Apple Silicon, ou quando precisar de quantização flexível de 2-8 bits sem requisitos de GPU.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

gguf

無料

GGUF format and llama.cpp quantization for efficient CPU/GPU inference. Use when deploying models on consumer hardware, Apple Silicon, or when needing flexible quantization from 2-8 bit without GPU requirements.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

llama-cpp

無料

llama.cpp local GGUF inference + HF Hub model discovery.

日本語の概要は準備中です。原文の説明を表示しています。

NousResearch/hermes-agent25.3万2026年10月11日 更新

Cost per million tokens on hardware you own (Ollama, llama.cpp, vLLM, LM Studio) from watts, electricity price, hardware price and measured tokens/second, and the utilisation at which local beats a hosted model. Use for local-vs-API break-even questions.

日本語の概要は準備中です。原文の説明を表示しています。

ruvnet/ruflo7.4万2026年10月11日 更新

pi-agent

無料

Builds with and operates Pi, the minimal terminal coding harness. Use for installing Pi, configuring providers/models/settings/environment variables, creating Pi skills/extensions/packages/themes/prompt templates, embedding Pi through the SDK, integrating over RPC or JSON event streams, parsing sessions, running local models through the llama.cpp router, developing custom Pi providers and TUI components, or using ecosystem packages such as pi-subagents (delegation/orchestration), pi-mcp-adapter (MCP servers), pi-interview (interactive forms), and pi-web-access (web search, fetching, video understanding).

日本語の概要は準備中です。原文の説明を表示しています。

K-Dense-AI/scientific-agent-skills4.8万2026年10月5日 更新

llama-cpp

無料

Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. Use for edge deployment, M1/M2/M3 Macs, AMD/Intel GPUs, or when CUDA is unavailable. Supports GGUF quantization (1.5-8 bit) for reduced memory and 4-10× speedup vs PyTorch on CPU.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

llama-cpp

無料

Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. Use for edge deployment, M1/M2/M3 Macs, AMD/Intel GPUs, or when CUDA is unavailable. Supports GGUF quantization (1.5-8 bit) for reduced memory and 4-10× speedup vs PyTorch on CPU.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Use to select models to run locally with llama.cpp and GGUF on CPU, Mac Metal, CUDA, or ROCm. Covers finding GGUFs, quant selection, running servers, exact GGUF file lookup, conversion, and OpenAI-compatible local serving.

日本語の概要は準備中です。原文の説明を表示しています。

huggingface/skills1.1万2026年10月9日 更新

Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5582026年10月10日 更新

Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5582026年10月10日 更新

[omh] Self-hosted LLM serving on GPUs: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the endpoint with the standard TTFT/TPOT/goodput protocol. Use when the user says: inference-serving, inference serving, serve this model, serve the model, model serving, serving endpoint, vllm, llama.cpp.

日本語の概要は準備中です。原文の説明を表示しています。

rlaope/oh-my-hermes3,2562026年10月10日 更新

Delegate a coding task to Aider (`aider`) as a background implementer, then review its diff and land it yourself. Use this whenever the user wants to hand implementation work to Aider - phrasings like "have Aider do X", "delegate this to aider", "run it through Aider", or "use Aider to implement/fix/refactor" - or wants to run a queue of coding tasks through Aider while staying the reviewer. This includes asking Aider to drive a local or self-hosted OpenAI-compatible endpoint ("have Aider use my local model", "run Aider against llama.cpp / Ollama / vLLM / LM Studio"), which Aider reaches via `--api-base`. DO NOT USE for local-model or coding requests that do not name Aider, for tasks small enough to do inline, or when the user wants the code written directly without delegating.

日本語の概要は準備中です。原文の説明を表示しています。

amElnagdy/delegate-skills2,3552026年10月8日 更新

uninstall

無料

Uninstall or hard-clean OpenClaw Companion, Windows node, native Gateway/MXC, WSL Gateway, and managed llama.cpp/Local AI state for a clean retest. Choose the existing dev CLI uninstall or the reviewed hard-clean procedure, inventory exact targets, and obtain destructive confirmation before acting.

日本語の概要は準備中です。原文の説明を表示しています。

openclaw/openclaw-windows-node2,1452026年10月10日 更新

Build private, on-device AI features on iPhone, iPad, and Mac with Foundation Models, Core ML, MLX Swift, or llama.cpp. Use when choosing an Apple-local model runtime, building an Apple Intelligence chatbot or tool-calling feature, running an LLM on Apple Silicon, converting or compressing a Python model for Core ML, or comparing on-device inference backends. For Swift Core ML loading and prediction code, use the coreml skill.

日本語の概要は準備中です。原文の説明を表示しています。

dpearson2699/swift-ios-skills1,1852026年8月1日 更新

Use to select models to run locally with llama.cpp and GGUF on CPU, Mac Metal, CUDA, or ROCm. Covers finding GGUFs, quant selection, running servers, exact GGUF file lookup, conversion, and OpenAI-compatible local serving.

日本語の概要は準備中です。原文の説明を表示しています。

waybarrios/opencode-power-pack5352026年10月6日 更新

Decision SOP for serving LLMs with vLLM. Covers PagedAttention mental model, quantization/parallelism/batching tradeoffs, OOM triage, and when NOT to use vLLM. Activates when a coder-agent is choosing or tuning an inference engine, debugging vLLM throughput/latency/OOM, or comparing vLLM against TGI/SGLang/TensorRT-LLM/llama.cpp.

日本語の概要は準備中です。原文の説明を表示しています。

agentsope/SkillAlchemy4372026年10月9日 更新

Cross-engine decision rubric for self-hosting or recommending an LLM serving stack. Picks among vLLM, SGLang, TensorRT-LLM, TGI, llama.cpp, Ollama, and MLX as a function of (hardware × workload × constraint), not "which is fastest". Activates whenever a coder-agent must choose, defend, or migrate a serving runtime.

日本語の概要は準備中です。原文の説明を表示しています。

agentsope/SkillAlchemy4372026年10月9日 更新

在本机 Mac 或 Apple Silicon 上部署 Gemma 4 12B。本地安装/升级 llama.cpp,下载 GGUF 量化模型,用 llama-server 暴露 OpenAI-compatible API,或用 Ollama 暴露本地模型服务;按用户需求在默认 Q4_K_M、64K/128K 长上下文、QAT Q4_0 @ 256K、左右对比演示之间选择,配置 tmux 后台运行,验证健康检查、问答接口、资源占用和常见故障。当用户说部署 Gemma 4、Gemma 4 12B、本地大模型、长上下文、QAT、量化、llama-server、Ollama、GGUF、Mac 本地模型服务时使用。

日本語の概要は準備中です。原文の説明を表示しています。

majiayu000/spellbook2872026年10月8日 更新

Optimize Ollama configuration for the current machine's hardware. Use when asked to speed up Ollama, tune local LLM performance, or pick models that fit available GPU/RAM. Don't use for LM Studio, llama.cpp, vLLM, or hosted-API LLM providers.

日本語の概要は準備中です。原文の説明を表示しています。

luongnv89/skills1312026年10月10日 更新

Runs local and self-hosted LLMs with Ollama, LM Studio, MLX, llama.cpp, and Open WebUI. Use when choosing GGUF quants, sizing VRAM, or running models privately.

日本語の概要は準備中です。原文の説明を表示しています。

vasilyu1983/AI-Agents-public912026年10月5日 更新