本文へ移動
cccskills

「gsm8k」の検索結果

16 件 ・ 関連度順

概要と使いどころ

evaluating-llms-harness

無料日本語概要

大規模言語モデルの知識・推論・コード生成を共通の課題で評価するスキル。複数モデルの比較表や学習途中の成績推移を作り、評価結果と個別の回答を保存します。

  • MMLU・GSM8Kでのモデル評価
  • 複数モデルの比較表作成
  • 学習途中のモデルの成績追跡
NousResearch/hermes-agent25.3万2026年10月11日 更新

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

日本語の概要は準備中です。原文の説明を表示しています。

foryourhealth111-pixel/Vibe-Skills3,6522026年8月31日 更新

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

日本語の概要は準備中です。原文の説明を表示しています。

OpenRaiser/NanoResearch1,3402026年10月9日 更新

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

日本語の概要は準備中です。原文の説明を表示しています。

sangrokjung/claude-forge8522026年9月3日 更新

Avalia LLMs em mais de 60 benchmarks acadêmicos (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use quando benchmarking de qualidade de modelo, comparação de modelos, relatório de resultados acadêmicos ou rastreamento de progresso de treinamento. Padrão da indústria usado por EleutherAI, HuggingFace e laboratórios principais. Suporta HuggingFace, vLLM, APIs.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Build and govern a 50-200 example domain-specific held-out benchmark sampled from real traffic. Distinct from public benchmarks (MMLU/HumanEval/GSM8K via lm-evaluation-harness) which measure GENERAL capability. Only a held-out domain set predicts whether THIS system works on YOUR data. Collect real examples, label, hold out (never train/prompt on it), size 50-200, version it, refresh on drift.

日本語の概要は準備中です。原文の説明を表示しています。

agentsope/SkillAlchemy4412026年10月9日 更新

lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.).

日本語の概要は準備中です。原文の説明を表示しています。

kevinnft/ai-agent-skills132026年8月1日 更新

Avalia LLMs em 100+ benchmarks de 18+ harnesses (MMLU, HumanEval, GSM8K, segurança, VLM) com execução multi-backend. Use quando precisar de avaliação escalável em Docker local, HPC Slurm ou plataformas cloud. Plataforma enterprise-grade da NVIDIA com arquitetura container-first para benchmarking reproduzível.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.).

日本語の概要は準備中です。原文の説明を表示しています。

openamer/openamer62026年10月12日 更新

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

日本語の概要は準備中です。原文の説明を表示しています。

ibragimov-oasis/vibe-coder22026年6月24日 更新