本文へ移動
cccskills
無料GitHub で公開

orc-evaluate

Set up LLM-as-judge evaluation for ORC workflows

インストール方法を見る

含まれるファイル(1)

  • SKILL.md2.8 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

ORC Evaluation (LLM-as-Judge)

Set up automated quality evaluation for ORC workflow outputs. Read docs/EVALUATION-COMPONENT.md for the full reference.

Require

(require '[ai.obney.orc.evaluation.interface :as eval])

Overview

The evaluation component provides LLM-as-judge scoring:

  • Grounding — Is the output faithful to the input?
  • Instruction-following — Does the output follow the instruction?
  • Reasoning — Is the reasoning sound?
  • Completeness — Does the output address all aspects?

Each judge returns a score (0.0-1.0) with structured feedback.

Built-in Judges

Evaluate one trace

(eval/evaluate-trace
  {:instruction "Summarize the article covering all main points"
   :inputs {:article "Long article text..."}
   :outputs {:summary "Short summary..."}}
  {:judges [:grounding :completeness]})
;; => ScoreWithFeedback with :score, :feedback, :dimensions, and—when
;;    model-backed—durable :model-provenance on the recorded score event

Workflow Judges (DSL Integration)

Attach judges to workflow nodes for automatic scoring during execution:

(orc/workflow "evaluated-pipeline"
  (orc/blackboard {:question :string :answer :string})

  ;; Define judge configurations
  (orc/judges
    {:grounding    {:type :grounding :weight 0.5}
     :completeness {:type :completeness :weight 0.5}})

  ;; Attach judges to a node
  (orc/llm "answer"
    :instruction "Answer the question thoroughly."
    :reads [:question]
    :writes [:answer]
    :judges [:grounding :completeness]))

When executed with tracing enabled, each judged node produces evaluation scores in the trace.

Batch Evaluation

Evaluate a workflow across multiple test cases:

(eval/evaluate-traces
  [{:instruction "Answer the question"
    :inputs {:question "What is AI?"}
    :outputs {:answer "Artificial intelligence..."}}
   {:instruction "Answer the question"
    :inputs {:question "What is ML?"}
    :outputs {:answer "Machine learning..."}}]
  {:judges [:grounding :completeness]})
;; => {:avg-score 0.82 :results [...] :min-score ... :max-score ...}

GEPA Integration

Evaluation judges feed directly into GEPA optimization. When GEPA proposes instruction candidates, judges score them — the scores drive Pareto selection.

;; GEPA uses judges automatically when configured
(gepa/optimize! ctx
  {:sheet-id sheet-id
   :node-name "answer"
   :judges {:grounding 0.5 :completeness 0.5}
   ...})

Reference

  • docs/EVALUATION-COMPONENT.md — Full evaluation guide
  • docs/SELF-IMPROVING-LOOP.md — Evaluation and continuous improvement
  • docs/GEPA-GUIDE.md — How judges integrate with optimization

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

allium

無料

Give your AI agents something more useful than a prompt. Velocity through clarity.

日本語の概要は準備中です。原文の説明を表示しています。

ObneyAI/orc212026年10月8日 更新

distill

無料

Extract an Allium specification from an existing codebase. Use when the user has existing code and wants to distil behaviour into a spec, reverse engineer a specification from implementation, generate a spec from code, turn implementation into a behavioural specification, or document what a codebase does in Allium terms.

日本語の概要は準備中です。原文の説明を表示しています。

ObneyAI/orc212026年10月8日 更新

elicit

無料

Run a structured discovery session to build an Allium specification through conversation. Use when the user wants to create a new spec from scratch, elicit or gather requirements, capture domain behaviour, specify a feature or system, define what a system should do, or is describing functionality and needs help shaping it into a specification.

日本語の概要は準備中です。原文の説明を表示しています。

ObneyAI/orc212026年10月8日 更新

Implement command handlers that validate state and emit events using the Grain framework

日本語の概要は準備中です。原文の説明を表示しています。

ObneyAI/orc212026年10月8日 更新

Create scheduled background jobs that run on cron schedules

日本語の概要は準備中です。原文の説明を表示しています。

ObneyAI/orc212026年10月8日 更新

Implement query handlers that read from event-sourced read models

日本語の概要は準備中です。原文の説明を表示しています。

ObneyAI/orc212026年10月8日 更新

ObneyAI のスキルをすべて見る

このスキルの問題を報告する