Use when the user wants help with academic papers or citations but it's unclear which specific workflow fits — reviewing a paper, checking a BibTeX file for fake references, or benchmarking multiple LLMs on reference-recommendation accuracy. Also use when the user mentions paper review, peer review, BibTeX verification, citation checking, reference hallucination, or academic literature accuracy and hasn't specified which of those three tasks they mean. This skill is the entry router for the academic-eval suite: it asks one diagnostic question then routes to the right sub-skill.
日本語の概要は準備中です。原文の説明を表示しています。
agentscope-ai/OpenJudge☆ 8712026年9月11日 更新
Use when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom task, or a benchmark specifically about reference/citation hallucination rate. Also use when the user mentions model arena, agent arena, pairwise model comparison, win-rate ranking, or comparing models on a task and hasn't specified whether that task is generic or about citation accuracy. This skill is the entry router for the arena-eval suite: it asks one diagnostic question when needed, then recommends the workflow or workflows needed to cover the request.
日本語の概要は準備中です。原文の説明を表示しています。
agentscope-ai/OpenJudge☆ 8712026年9月11日 更新
Automatically evaluate and compare multiple AI models or agents without pre-existing test data. Generates test queries from a task description, collects responses from all target endpoints, auto-generates evaluation rubrics, runs pairwise comparisons via a judge model, and produces win-rate rankings with reports and charts. Supports checkpoint resume, incremental endpoint addition, and judge model hot-swap. Use when the user asks to compare, benchmark, or rank multiple models or agents on a custom task, or run an arena-style evaluation.
日本語の概要は準備中です。原文の説明を表示しています。
agentscope-ai/OpenJudge☆ 8712026年9月11日 更新
Build custom LLM evaluation pipelines using the OpenJudge framework. Covers selecting and configuring graders (LLM-based, function-based, agentic), running batch evaluations with GradingRunner, combining scores with aggregators, applying evaluation strategies (voting, average), auto-generating graders from data, and analyzing results (pairwise win rates, statistics, validation metrics). Use when the user wants to evaluate LLM outputs, compare multiple models, design scoring criteria, or build an automated evaluation system.
日本語の概要は準備中です。原文の説明を表示しています。
agentscope-ai/OpenJudge☆ 8712026年9月11日 更新
Review academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline. Supports PDF files and LaTeX source packages (.tar.gz/.zip). Covers 10 disciplines: cs, medicine, physics, chemistry, biology, economics, psychology, environmental_science, mathematics, social_sciences. Use when the user asks to review, evaluate, critique, or assess a research paper, check references, or verify a BibTeX file.
日本語の概要は準備中です。原文の説明を表示しています。
agentscope-ai/OpenJudge☆ 8712026年9月11日 更新
Verify a BibTeX file for hallucinated or fabricated references by cross-checking every entry against CrossRef, arXiv, and DBLP. Reports each reference as verified, suspect, or not found, with field-level mismatch details (title, authors, year, DOI). Use when the user wants to check a .bib file for fake citations, validate references in a paper, or audit bibliography entries for accuracy.
日本語の概要は準備中です。原文の説明を表示しています。
agentscope-ai/OpenJudge☆ 8712026年9月11日 更新