本文へ移動
cccskills
無料GitHub で公開

arize

Arize is an AI observability platform, and Phoenix is its open-source tracing and evaluation tool for LLM applications. Use this skill when the user wants to trace OpenAI or LangChain calls, run a local Phoenix server, score RAG answers for faithfulness, debug slow or wrong LLM responses, or send OpenTelemetry traces to Arize for production monitoring.

インストール方法を見る

含まれるファイル(2)

  • SKILL.md6.2 KB
  • _scores.json2.6 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Arize (Phoenix) — AI Observability Platform

Overview

Phoenix (arize-phoenix, source-available under the Elastic License 2.0, free to self-host) is an OpenTelemetry-native trace collector and UI for LLM apps: it records every model call, retrieval step and tool call, and lets you evaluate the results with an LLM judge. Arize is the hosted commercial platform from the same company for production monitoring, with the same OpenInference trace format. The workflow is: run Phoenix locally while developing, instrument with OpenInference, score traces with arize-phoenix-evals, and point the same instrumentation at Arize when you need dashboards and alerts.

Packages checked in October 2026: arize-phoenix 20.x, arize-phoenix-evals 3.x, arize (Arize SDK) 8.x, arize-otel 0.14.

Instructions

1. Install and start Phoenix

pip install arize-phoenix arize-phoenix-otel openinference-instrumentation-openai openai
phoenix serve            # UI and OTLP collector on http://localhost:6006

uvx arize-phoenix serve runs it without installing; a container is published as arizephoenix/phoenix. In a notebook, px.launch_app() still starts a session, but for applications run the server as a separate process. Set PHOENIX_TELEMETRY_ENABLED=false to opt out of usage analytics.

2. Trace an application

from phoenix.otel import register
from openinference.instrumentation.openai import OpenAIInstrumentor
import openai

tracer_provider = register(project_name="support-chatbot")   # sends to localhost:6006
OpenAIInstrumentor().instrument(tracer_provider=tracer_provider)

client = openai.OpenAI()
client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Explain CRDTs to a junior developer"}],
)

register(auto_instrument=True) instruments every installed openinference-instrumentation-* package. To send to another server, set PHOENIX_COLLECTOR_ENDPOINT (and PHOENIX_API_KEY for a protected instance) or pass endpoint=.

3. Pull spans and evaluate them

The current evals API is LLM plus metric evaluators run with evaluate_dataframe. The older OpenAIModel, run_evals and QAEvaluator classes are gone from the top-level API.

import pandas as pd
from phoenix.client import Client
from phoenix.client.types.spans import SpanQuery
from phoenix.evals import LLM, evaluate_dataframe
from phoenix.evals.metrics import FaithfulnessEvaluator, CorrectnessEvaluator

spans = Client().spans.get_spans_dataframe(
    query=SpanQuery().where("span_kind == 'LLM'"),
    project_name="support-chatbot",
    limit=200,
)

# Evaluators read columns named input, output (and context for faithfulness).
rows = pd.DataFrame({
    "input": spans["attributes.input.value"],
    "output": spans["attributes.output.value"],
    "context": "Refunds are issued within 14 days of purchase.",
})

judge = LLM(provider="openai", model="gpt-4o-mini")   # needs OPENAI_API_KEY
scores = evaluate_dataframe(rows, [FaithfulnessEvaluator(judge), CorrectnessEvaluator(judge)])

Each evaluator returns a label, a 0/1 score and an explanation. Other built-ins: HallucinationEvaluator (grounded in the conversation itself), RetrievalRelevanceEvaluator (replaces the deprecated DocumentRelevanceEvaluator), ToxicityEvaluator, RefusalEvaluator, PiiDetectionEvaluator, and tool-call evaluators. create_classifier builds a custom judge from your own prompt and labels. Judge models must support tool calling.

4. Send traces to Arize

pip install arize-otel
import os
from arize.otel import register

tracer_provider = register(
    space_id=os.environ["ARIZE_SPACE_ID"],
    api_key=os.environ["ARIZE_API_KEY"],
    project_name="support-chatbot",
    auto_instrument=True,
)

EU accounts pass endpoint=Endpoint.ARIZE_EUROPE (from arize.otel import Endpoint). The arize package (v8) provides ArizeClient(api_key=...) for datasets, experiments and spans over the REST API; the old arize.pandas.logger.Client and ModelTypes flow is not in the current SDK docs, so use tracing for LLM apps.

Examples

Example 1: "Why was this chatbot answer slow?"

phoenix serve &
python app.py                 # app calls register(project_name="support-chatbot")

Open http://localhost:6006, select the support-chatbot project, sort traces by latency, and expand the slowest one. The trace tree shows the retrieval span, the LLM span with prompt, completion and token counts, and which step took the time.

Example 2: "Check my RAG answers for hallucinations"

Run the step 3 script against 200 recent LLM spans. scores has one row per span with faithfulness_score, a label (faithful or unfaithful) and the judge's explanation. Filter for unfaithful rows, read the explanations, and fix the retrieval or prompt for those queries.

Guidelines

  • Phoenix stores data in a local SQLite file by default; for teams set PHOENIX_SQL_DATABASE_URL to PostgreSQL and run the container.
  • The LLM judge costs tokens: evaluate a sample, not every span, and use a cheaper model first.
  • The Phoenix client reads spans through phoenix.client; the old px.Client().get_spans_dataframe(filter_condition=...) call is replaced by SpanQuery().where(...).
  • Old px.Inferences / px.Schema embedding-drift sessions are no longer exported by the phoenix package; use Arize for embedding drift and clustering.
  • Keep API keys in environment variables; traces contain prompts and user data, so do not point a development app at a shared instance unintentionally.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Scripts and configures production rendering in Autodesk 3ds Max with the V-Ray and Corona renderers: output size and files, render elements, denoising, light mix, batch and command-line rendering, and network rendering. Use when a user asks to set up a production render, render several cameras in one batch, render from the command line or on a render farm, add render passes for compositing, or cut render time for archviz and product shots.

日本語の概要は準備中です。原文の説明を表示しています。

TerminalSkills/skills1632026年10月4日 更新

Covers scripting Autodesk 3ds Max, the 3D modeling and rendering application, with MAXScript and Python (pymxs): scene manipulation, object creation, material assignment, camera and light setup, batch operations, and file I/O. Use when tasks involve automating repetitive 3ds Max workflows, batch processing scenes, running scripts headless with 3dsmaxbatch, creating custom tools, or scripting scene setup for archviz, product visualization, or VFX.

日本語の概要は準備中です。原文の説明を表示しています。

TerminalSkills/skills1632026年10月4日 更新

3proxy

無料

3proxy is a small open-source proxy server that runs HTTP/HTTPS, SOCKS4/5, SNI and TCP/UDP port-mapping proxies from one config file. Use when a user asks to set up an HTTP or SOCKS5 proxy, add proxy users and passwords, write 3proxy access rules, chain or rotate upstream (parent) proxies, limit bandwidth, connections or monthly traffic per user, run 3proxy in Docker, or fix a 3proxy.cfg that will not start.

日本語の概要は準備中です。原文の説明を表示しています。

TerminalSkills/skills1632026年10月4日 更新

Builds Agent2Agent (A2A) servers and clients, the open protocol (originally from Google, now under the Linux Foundation) that lets AI agents from different frameworks call each other. Use when the user wants to create an A2A-compliant agent, build an Agent Card, implement task management, connect agents across frameworks, set up agent discovery, handle streaming responses, implement push notifications, or orchestrate multi-agent workflows. Trigger words: a2a, agent to agent, agent2agent, a2a protocol, a2a server, a2a client, agent card, agent interoperability, agent collaboration, multi-agent, agent discovery, a2a sdk, a2a task.

日本語の概要は準備中です。原文の説明を表示しています。

TerminalSkills/skills1632026年10月4日 更新

Plans a controlled experiment (A/B test) so its result can be trusted: writes the hypothesis, picks one primary metric and the guardrails, computes sample size and run time, specifies how visitors are assigned and when exposure is logged, and reads out the result with a confidence interval. Use when someone says "set up an A/B test", "split test this page", "how many visitors do I need", "how long should the experiment run", "is this result significant", "can I stop the test early", or wants to test a headline, price, layout or onboarding change against the current version.

日本語の概要は準備中です。原文の説明を表示しています。

TerminalSkills/skills1632026年10月4日 更新

ably

無料

Ably is a hosted realtime messaging service: clients publish and subscribe to named channels over WebSockets, see who is present, replay message history, and resume after a dropped connection. Use when a user asks to "add realtime updates", "push live notifications to the browser", "show who is online", "add a chat room with typing indicators", "publish from a serverless function", or "authenticate Ably clients without exposing the API key". Covers the ably 2.x JavaScript SDK (Realtime and REST), JWT token authentication, presence, history and rewind, batch publishing, and the @ably/chat 1.x SDK.

日本語の概要は準備中です。原文の説明を表示しています。

TerminalSkills/skills1632026年10月4日 更新

TerminalSkills のスキルをすべて見る

このスキルの問題を報告する