本文へ移動
cccskills
無料GitHub で公開

model-benchmarker

Use when: benchmark model latency, throughput, or quality.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md1.8 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Model Benchmarker

Tests Model-Performance über mehrere Provider (OpenRouter, lokale Ollama-Instanzen).

Tests

TestBeschreibungRunsMetrik
LatenzZeit bis erster Token (kleiner Prompt)10Median in Sekunden
DurchsatzTokens pro Sekunde (4K Prompt)5Median tok/s
QualitätAntwort auf 4 Testfragen bewertet1Score 0-100%

CLI-Verwendung

# Einmaliger Test
python model-benchmarker.py --run openrouter/deepseek-v4-flash
python model-benchmarker.py --run "local/qwen3.5:9b"

# Alle Provider testen (Default-Modell pro Provider)
python model-benchmarker.py --all

# Vergleichstabelle aller getesteten Modelle
python model-benchmarker.py --compare
python model-benchmarker.py --compare --json   # JSON-Ausgabe

# Trend-Historie
python model-benchmarker.py --history openrouter
python model-benchmarker.py --history local

Provider-Konfiguration

Liest automatisch aus config.yaml:

  • model.default → OpenRouter-Modell
  • custom_providers → Lokale Ollama-Instanzen und deren Modelle
  • API-Keys aus .env (OPENROUTER_API_KEY)

Ergebnisverzeichnis

Alle Ergebnisse in ~/.model-benchmarks/results/<provider>-<model>.json mit vollständiger Historie (max 50 Einträge).

Cron-Job

Wöchentlicher --all Run ist als Cron-Job eingerichtet:

  • Führt alle Provider-Modelle durch
  • Sendet Ergebnisse an Terminal (lokal)
  • Kein Agent nötig (no_agent: true)

Abhängigkeiten

Keine — nur Python-Standardbibliothek (urllib, json, statistics, etc.)

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Build integrated IS/BS/CF financial workbooks in Excel.

日本語の概要は準備中です。原文の説明を表示しています。

openamer/openamer52026年10月11日 更新

a2a-swarm

無料

Run and grow the OpenAmer Agent-to-Agent swarm: identity, trust, node-to-node ask, signed skill/insight sharing, and the autonomous self-learning loop.

日本語の概要は準備中です。原文の説明を表示しています。

openamer/openamer52026年10月11日 更新

Use for A/B experiments on OpenAmer configs and skills.

日本語の概要は準備中です。原文の説明を表示しています。

openamer/openamer52026年10月11日 更新

Roleplay a hostile user to find and triage UX pain points.

日本語の概要は準備中です。原文の説明を表示しています。

openamer/openamer52026年10月11日 更新

Agent mesh: master/worker nodes, HTTP delegation, heartbeat.

日本語の概要は準備中です。原文の説明を表示しています。

openamer/openamer52026年10月11日 更新

Use for agentic commerce payments tasks. Candidate synthesized from trend radar; awaiting trial promotion.

日本語の概要は準備中です。原文の説明を表示しています。

openamer/openamer52026年10月11日 更新

openamer のスキルをすべて見る

このスキルの問題を報告する