本文へ移動
cccskills
無料GitHub で公開

failproofai-sdk

Make a custom AI agent — Python or TypeScript/JavaScript, on a framework or hand-built — report what it did to Failproof AI, and run your own evaluator worker (the "eval pod") that scores those runs. Reach for it on vague phrasing too: "add observability to my agent", "why isn't my agent showing up?", "run an LLM judge on our own infra". Trigger when the user wants to: • plan an integration — which points in the agent loop to record; • instrument — add `failproofai-sdk` (Python) or `@failproofai/sdk` (Node, Bun, Deno, Next.js): turn on an adapter (LangChain/LangGraph, CrewAI, LlamaIndex, Pydantic AI, Vercel AI SDK, Mastra) or wire a hand-built loop; • verify — confirm events are written, or debug an integration that produces nothing; • evaluate — write, deploy or debug an Evaluator worker in Python or TypeScript. NOT for reading telemetry or scores that already landed (that's `fp-cloud-cli`), or deciding what is worth evaluating (that's `failproofai-eval-brainstorm`).

インストール方法を見る

含まれるファイル(8)

  • SKILL.md23.7 KB
  • agents/openai.yaml396 B
  • references/evaluator.md10.9 KB
  • references/events.md13.7 KB
  • references/frameworks.md16.2 KB
  • references/install.md4.7 KB
  • references/integration.md15.4 KB
  • references/typescript.md29.9 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Failproof AI SDK — Python and TypeScript

Two packages, one pipe: failproofai-sdk (Python, imported as failproofai_sdk) and @failproofai/sdk (TypeScript/JavaScript). They write the same 15 events in the same wire format into the same spool directory. Everything in this file is language-neutral unless it says otherwise, and code blocks are Python. For a TypeScript or JavaScript agent, read references/typescript.md alongside it: it has the camelCase names, the adapters, bundler and Next.js setup, the no-framework wiring, shutdown and verification. Both packages also ship the evaluator worker that scores finished sessions on your own infrastructure (§7, references/evaluator.md).

The SDK records what your agent did, from inside your agent. You call it at points you choose; it appends structured events to local .jsonl files. A separate collector ships those files to the platform.

your agent calls failproofai_sdk.event.*
  → SDK queues it in memory
  → flush thread writes <base_dir>/events/event-<timestamp>.jsonl
  → collector picks the file up and ships it
  → visible as sessions / events / errors / evals

The SDK's job ends at the file. That boundary is the most useful thing to know about it: everything up to the .jsonl is yours to get right and yours to verify, and it is verifiable on a laptop with no server, no API key, and no network.

The API is small — 15 event methods, all keyword-only. The hard parts are deciding where to call them and knowing which silences are bugs, because this SDK does not raise when you get it wrong. Sections 1-3 are the plan, 4 is the code, 5-6 are the proof, 7 is scoring the runs.

1. Install it

TypeScript/JavaScript: npm install @failproofai/sdk — zero dependencies, Node ≥ 20.9, Bun or Deno, ESM and CommonJS. The npm name has no lookalike trap; the rest of this section is Python. See references/typescript.md.

pip install failproofai-sdk        # or: uv add failproofai-sdk

The distribution is failproofai-sdk and the import is failproofai_sdk. Public PyPI, no token, no dependencies.

One command to never run: pip install agenteye. That name belongs to a stranded release of an old CLI — a different product that shipped under it before moving to fp-cloud-cli. PyPI versions cannot be withdrawn, so the name still resolves to that build forever. You get the CLI, import failproofai_sdk raises ModuleNotFoundError, and on a codebase still using the pre-rename SDK (which published under agenteye too) pip treats it as an upgrade and removes the SDK.

Tell: if a coding agent proposes pip install agenteye to install the SDK, this skill never loaded. Stop and re-read it.

The CLI is a fine thing to want — it is what reads the telemetry back. Install it separately, never with pip into your agent's environment:

pipx install fp-cloud-cli        # the command is `fp`

Confirm what you actually have before writing a line of instrumentation:

python -c "import failproofai_sdk; print(failproofai_sdk.__version__)"

A version like 0.0.1b1 is the SDK. ModuleNotFoundError means it is not installed — check pip show agenteye, which returning anything means the wrong name was installed. references/install.md covers migrating an existing import agenteye integration.

2. Plan before you instrument

Instrumentation lands in code that already exists and already works. Read it first, then decide. Two questions settle most of the design, and only the user can answer the first:

What is one run of this agent? That is your session_id — one value for the whole run, generated by you at the point the run starts. A chat turn, a job, a request, a workflow execution. If the agent handles concurrent runs, this must be per-run, not per-process.

What are the distinguishable actors in a run? That is your agent_id — a stable label, not a unique id. "planner", "researcher", "main". It is how the platform tells sub-agents apart, so reuse the same string across runs.

Get these two named and agreed before writing code. They are the axes every surface groups by, and changing them later splits the history: old runs keep the old labels and the trends break.

The two events everything else hangs off

Most of the catalog is optional and incremental. These two are not:

EventWithout it
agent_startThe session does not exist. No row on Sessions, no timeline, no evaluation — while every other event you emit still lands fine and shows up in the event stream.
agent_endThe run never closes, and it is not handed to the evaluator at the normal time.

That first row is the single most common integration failure, and it is completely silent: a run emitting 500 tool calls and no agent_start produces a busy event stream and zero sessions. Sessions are defined as "something that emitted agent_start". So:

Emit agent_start at the top of the run and agent_end at every exit, and get those two working end-to-end before you instrument anything else. One event at each end proves the whole path — install, identity, base dir, collector — with almost no code to be wrong. Add tools, models, and hooks after that path is green.

Then map the rest onto the agent's shape

Walk the agent loop and pick the points that exist in this codebase. Skip what doesn't apply; there is no requirement to emit every type.

In the codeEmitBuys you
every exit path of a run — success, exception, early returnagent_start / agent_endthe session itself
the tool dispatcher, both sides of the calltool_use / tool_resultwhat ran, in what order, how long
the LLM client wrapper, both sidesmodel_request / model_responsemodel mix, token spend, stop reasons
your except blockserrorthe Errors surface
a policy/guard/middleware layerhook_triggered / hook_completedhook behaviour
an approval gate or human handoffhuman_wait / human_input, human_pause, human_interruptwhere runs sit waiting on people
a run that suspends and resumes — waiting for a human, throttled, user-pausedagent_pause / agent_resumea real "paused" state: the agent isn't ended, the resume isn't a new agent, and wait time is excluded from active work

If the codebase has one tool dispatcher and one LLM wrapper, you have two edit sites for the bulk of the value. If tool calls are scattered inline across the codebase, say so — a wrapper (§4) is worth more than 40 call sites.

Full field-by-field catalog: references/events.md.

3. The contract

Work with these; none of them raise, so none of them show up in testing.

(TypeScript: the same contract with camelCase spellings, the same "dev" default, the same reserved names and the same duration_ms rule, but a different shutdown recipe — references/typescript.md → The contract, in TypeScript and Shutdown.)

  • There IS an ambient session, and it is the ergonomic path. session(), agent() and tool_call() bind identity on contextvars, so session_id and agent_id are optional on all 15 event methods — omitted, they resolve from the enclosing scope. current() reads it; propagate(fn) carries it into a new thread, which contextvars do NOT do on their own.

    This section said the opposite until the scopes existed, and the reference integration shipped a contextvars wrapper as markdown for customers to paste into their own code. That is now in the package.

    Nothing bound and nothing passed raises TypeError naming the fix — never a silent emit, because ingest skips an event with no session and answers 200.

    Two more shapes raise, for the same reason:

    You passRaisesWhy it cannot be allowed through
    A non-str idTypeErrorIngest skips the event and still answers 200
    "" or " "ValueErrorWorse — ingest accepts it, and every event merges under one blank id
  • configure() is optional, and every call restates all of it. It is keyword-only with exactly three settings:

    argdefault resolution
    base_dir~/.failproofai/custom-agents (honours $FAILPROOFAI_HOME)
    environment$AGENTEYE_ENVIRONMENT, else "dev"
    flush_interval0.5 (seconds)

    No environment variable can move the spool out of the umbrella. $FAILPROOFAI_HOME relocates the umbrella itself, but custom-agents is appended unconditionally, so the spool is always inside it. base_dir is the only way to write anywhere else, and it is an explicit argument at the call site rather than something inherited from the environment.

    The default root moved here from ~/.agenteye. failproofaid watches both, so on a host running it nothing changes but the directory name, and batches already in ~/.agenteye/events still get collected. On a host running the older agenteye-collector — which resolves $AGENTEYE_HOME or ~/.agenteye and nothing else — point the collector at this SDK with AGENTEYE_HOME=~/.failproofai/custom-agents, or pass base_dir here. Setting AGENTEYE_HOME no longer moves the SDK: it used to, which meant exporting it for the collector silently relocated the SDK too. AGENTEYE_SPOOL_TO_FAILPROOFAI is retired; it required a directory nothing created, so it never fired.

    Each call sets all three — omitted arguments are reset to default resolution, not left alone. So a later configure(flush_interval=1.0) silently moves your events back to the default directory and re-resolves the environment. An explicit configure(environment=...) beats the env var; omit it and the env var applies again. Call it once, at startup, before the first event, passing every argument you care about.

  • environment defaults to "dev". An unconfigured production agent reports its runs as dev and they are invisible wherever the team filters on production. Set it explicitly via configure(environment=...) or the AGENTEYE_ENVIRONMENT env var. This is a favourite: everything works, in the wrong bucket.

    A comma raises ValueError. configure(environment="prod,eu") is rejected at the call site: ingest splits this field on commas to build filter facets, so a comma would discard the whole event server-side with nothing said.

  • Non-JSON payload leaves are stringified. Events are serialized on a background thread. Ordinary structured JSON retains its types; unsupported leaves such as datetime, UUID, Decimal, set, bytes, or a Pydantic model are converted with str(value) so one awkward tool result cannot stop recording. Prefer plain JSON values when downstream queries need their structure; use explicit custom serialization when a string would be ambiguous.

  • Field names are unvalidated — but only the optional ones. Every method takes arbitrary **fields and stores them as-is, so a typo'd optional name (inpt= for input=) is not an error, it is a new field, and nothing will tell you. Typos in required names raise TypeError (they're real parameters), and five reserved names — timestamp, session_id, agent_id, type, environment — raise ValueError.

  • outcome="failed", not "failure". A run counts as failed only when outcome (or status) is one of error, failed, timeout, rejected (case-insensitive). "failure" is the natural antonym of the "success" in every example — and it silently counts as not a failure. The run shows green.

  • You own correlation, and ids are scoped per session AND per kind. Pending spans are keyed tool:<session_id>:<tool_call_id> and hook:<session_id>:<hook_id>, so a hook_completed(hook_id="x") cannot pair with a pending tool_use(tool_call_id="x"), and two concurrent sessions both using call_1 cannot cross-pair either. input_id and pause_id are scoped the same way.

    What must still be unique is an id within one session, for one kind. The pending map is a plain assignment, so emitting tool_use(tool_call_id="call_1") twice in one session overwrites the first entry and the first tool_result measures from the wrong start. Reusing your framework's id is always safe (Anthropic and OpenAI ids are globally unique); a per-run counter is safe only if you do not reset it inside a session.

    If you have read older guidance describing one flat, process-wide map shared between tools and hooks: that was true, and is not any more.

  • duration_ms is computed for you on four methods only — tool_result, hook_completed, human_input, agent_resume — from the matching earlier event. Passing it to those four raises ValueError. Passing it to any of the other eleven is silently accepted as a custom field.

  • Events are fire-and-forget. event.* queues in memory and returns; a daemon thread writes every 0.5s, plus once at interpreter exit. A clean exit flushes. A hard kill (SIGKILL, os._exit, a container OOM) drops whatever is queued, silently.

    SIGTERM deserves its own line, because it is not exotic — it is every rolling deploy, every docker stop, every Kubernetes eviction, and every plain kill. CPython installs no handler for it: signal.getsignal(SIGTERM) is SIG_DFL, the OS terminates the process where it stands, and atexit does not run. Whatever is queued is gone — and what is in flight at shutdown is disproportionately agent_end, so runs never close and never reach the evaluator. The 0.5s flush interval is what bounds the loss, not the exit path.

    So if your process can receive SIGTERM, handle it — the SDK will not install a handler in your process behind your back:

    import signal, sys, failproofai_sdk
    
    def _flush_and_exit(signum, frame):
        failproofai_sdk._writer.flush_now()
        sys.exit(128 + signum)
    
    signal.signal(signal.SIGTERM, _flush_and_exit)
    

    (sys.exit here rather than os._exit: it unwinds, so any agent() scope still open emits its agent_end before the flush. That scope closes outcome="failed" with an error naming SystemExit, because an evicted run did not finish — which is the thing you want to be able to see.) SIGKILL, os._exit and a container OOM cannot be handled by anything, and drop the queue silently.

4. Write it

Threading session_id and agent_id through every call site by hand is the thing that makes integrations ugly and abandoned. Don't. Bind identity once per run and let the call sites read it.

references/frameworks.md covers the four Python adapters; references/typescript.md covers the four TypeScript ones (LangChain.js/LangGraph.js, Vercel AI SDK, Mastra, LlamaIndex.TS), Next.js, bundlers, and the three-edit-site wiring for a hand-built TypeScript loop. references/integration.md has the hand-written wrapper — one small module, correct under asyncio and threads, adaptable to any codebase — plus worked shapes for a tool dispatcher, an LLM client wrapper, and framework-specific callback layers. Read it before writing your own; the naive version (a module global, or a plain attribute) breaks the moment two runs overlap, and it breaks by mixing two runs' events together rather than by failing.

Match the codebase you're in. If it's async, the wrapper is async. If it already has a request context or a trace id, bind to that instead of inventing one.

5. Verify — watch the files

(TypeScript: same directory, same checklist; the commands are in references/typescript.md → Verify.)

This is the whole point of the file boundary: you can prove the integration without a server. Run the agent and look.

Resolve the spool the way the SDK does, rather than guessing at a path:

python -c "import failproofai_sdk._resolver as r; print(r.get_base_dir() / 'events')"

That prints ~/.failproofai/custom-agents/events unless the application called configure(base_dir=...). $FAILPROOFAI_HOME moves the ~/.failproofai part and nothing else. $AGENTEYE_HOME does not affect it — that variable belongs to the older agenteye-collector, which reads it to decide what to WATCH.

ls -la ~/.failproofai/custom-agents/events/

You are looking for event-<UTC timestamp>-<pid>-<seq>.jsonl files — the pid and sequence number are what keep two processes flushing in the same millisecond from overwriting each other. Each line is one event. Read them with a JSON parser, not grep — the exact spacing is not a contract, and a grep for "type":"agent_start" returns nothing on a perfectly healthy integration:

cat ~/.failproofai/custom-agents/events/*.jsonl | python -m json.tool --json-lines | head -20

Then check, in this order — the first failure explains everything downstream:

  1. Any files at all — or do they stop mid-run? (Python) Look at stderr for Exception in thread failproofai-sdk-flush. This is the first thing to check and the worst thing to miss: one non-JSON-serializable value killed the writer, and everything after it — including the at-exit flush — is gone (§3). The tell is that events stop for every type at once, and nothing raised. If instead there were never any files: did import failproofai_sdk succeed (§1)? Is the base dir writable? Did the process die hard (SIGKILL, docker stop, an OOM) before a flush?
  2. Is agent_start there, once per run? No → you will see events on the platform and no sessions, and you will spend an afternoon on it (§2).
  3. Sessions but no tool or model events? Your emit path is dropping them before the SDK ever sees them — nearly always because they're emitted from a thread the identity never reached. See references/integration.md → "Threads will drop your events". The SDK is silent here; only your own wrapper can warn.
  4. Is environment what you expect? It is "dev" unless you set it (§3).
  5. Is outcome on agent_end a word that counts? failed/error/timeout/ rejected — not "failure" (§3). Failed runs showing green is this, every time.
  6. Run two overlapping runs. Confirm two session_ids with no events crossing between them. Do not check this with one run: a single run passes even when identity is a module global, and mixing only appears once two runs overlap — which is production, not your laptop (§4).
  7. Do tool_use and tool_result share a tool_call_id? Unpaired means no duration. Also confirm no id repeats within a session for the same kind — a collision pairs the wrong two events and reports a confident wrong duration (§3).

A test-mode loop that costs nothing:

export FAILPROOFAI_HOME=/tmp/failproofai-sdk-test
rm -rf /tmp/failproofai-sdk-test && python your_agent.py
cat /tmp/failproofai-sdk-test/custom-agents/events/*.jsonl | python -m json.tool --json-lines

FAILPROOFAI_HOME sends events somewhere disposable, so you can iterate on the integration without touching the real directory or shipping test runs to the platform. Note the custom-agents segment in the read path — the SDK appends it unconditionally. Note too that the SDK reads the variable late, per flush, so set it before you start the process, not halfway through.

AGENTEYE_HOME used to do this job and no longer does anything to the SDK; using it here would write to your REAL spool while you read an empty temp directory.

Do not verify by installing the CLI into your agent's environment. It will uninstall the SDK you just integrated (§1). Reading back what landed on the platform is the fp-cloud-cli skill's job, from a separate environment.

6. Production — the collector has to agree with you

The SDK writes files. It never talks to the network, so from its point of view a completely unshipped integration looks perfect.

In production, the collector must be running and reading the same directory the SDK is writing to. That is the whole contract, and both halves fail silently:

  • Collector not running → files pile up in events/ forever. The SDK is fine.
  • Collector reading a different base dir than the agent writes to — the collector's own AGENTEYE_HOME pointing somewhere else, a different user's ~, a container path that isn't mounted → files pile up in a directory nobody reads. The SDK is fine. (failproofaid watches both ~/.failproofai/custom-agents/events and ~/.agenteye/events, so it is the half of this pair least likely to be misconfigured.)

So when events are on disk but not on the platform, the SDK is not the suspect. Compare the two paths first: print the directory your agent is actually writing to (python -c "import failproofai_sdk._resolver as r; print(r.get_base_dir())" in the agent's own environment, with the agent's own env vars) and check the collector is running and pointed at the same one. A .jsonl count that only grows is the tell. The reverse is healthy: with failproofaid running, the directory empties seconds after each flush, because the daemon ships each file and deletes it — so an empty real spool proves nothing either way. Verify content with the throwaway FAILPROOFAI_HOME loop in §5, and arrival with fp-cloud-cli.

Confirming events arrived on the platform is deliberately not this skill's job — that is the fp-cloud-cli skill, from a separate environment (§1). Collector setup and deployment are your platform's own documentation.

If the files look right (§5) and the collector is running against the same directory, the integration is done.

7. Score the runs — your own evaluator worker

Sessions that land can be scored. Hosted evaluations are written in the dashboard (Analyze → eval authoring) and run on Failproof AI's managed evaluator: that is the default. When an evaluation needs your own model keys, packages, secrets, private network or heavy compute, run it in your own worker — the eval pod.

It ships in the same packages: failproofai_sdk.evaluator and @failproofai/sdk/evaluator. You declare an Evaluator, register versioned evaluations with an optional when condition, and start it with an evaluations:run key in FAILPROOFAI_EVALUATOR_TOKEN. It only calls out over HTTPS — claim finished sessions, score them, submit — so a pod needs egress and a secret, and no ingress. Its results carry the customer tag.

Two things to settle before writing one: the agents must already produce finished sessions (§2 — nothing scores a run with no agent_end), and bump an evaluation's version whenever its logic changes. Everything else — the API in both languages, env vars, a Dockerfile, SIGTERM drain vs. the pod's grace period, scaling, and a debugging order for a worker that scores nothing — is in references/evaluator.md.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

The way to answer "how are my production AI agents doing?" and to run the team's agent-observability deployment — reach for it even on casual phrasing that names no tool. Trigger when the user wants to: • inspect agent telemetry — did agents error/fail/go flaky; sessions, events, latency, token usage, slowest models; eval/quality scores and whether quality dropped; • operate the deployment — ack/assign/resolve/mute/dismiss issues (alerts, reports, and audit findings) with notes; run and triage audits; see who has access and change roles (e.g. read-only); create or scope API keys (e.g. a push-only CI key); change settings; run saved or ad-hoc ClickHouse queries. Served by the `fp` CLI against FailproofAI Cloud. NOT for writing or designing an evaluator service / scoring logic (that's `agenteye-evaluator`), adding SDK/instrumentation to your app (that's `failproofai-sdk`, imported as `failproofai_sdk`), debugging the collector/daemon, or unrelated dev work (why a build/CI run failed, rotating non-FailproofAI Cloud secrets).

日本語の概要は準備中です。原文の説明を表示しています。

FailproofAI/failproofai5,2722026年10月11日 更新

mintlify

無料

Build and maintain documentation sites with Mintlify. Use when creating docs pages, configuring navigation, adding components, or setting up API references.

日本語の概要は準備中です。原文の説明を表示しています。

FailproofAI/failproofai5,2722026年10月11日 更新

FailproofAI のスキルをすべて見る

このスキルの問題を報告する