本文へ移動
cccskills
無料GitHub で公開

write-fixtures

Use when writing test fixtures for @copilotkit/aimock — mock LLM responses, tool call sequences, error injection, multi-turn agent loops, embeddings, structured output, sequential responses, or debugging fixture mismatches

インストール方法を見る

含まれるファイル(1)

  • SKILL.md72.9 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Writing aimock Test Fixtures

What aimock Is

aimock is a zero-dependency mock infrastructure for AI apps. Fixture-driven. Multi-provider (OpenAI, Anthropic, Gemini, Gemini Interactions, AWS Bedrock, Azure OpenAI, Vertex AI, Ollama, Cohere, OpenRouter). Multimedia endpoints (image generation, text-to-speech, audio transcription, video generation). MCP, A2A, AG-UI, and vector DB mocking. Runs a real HTTP server on a real port — works across processes, unlike MSW-style interceptors. WebSocket support for OpenAI Responses/Realtime and Gemini Live APIs. Record-and-replay for all endpoints including multimedia. Chaos testing and Prometheus metrics.

Core Mental Model

  • Fixtures = match criteria + response
  • First-match-wins — order matters
  • All providers share one fixture pool (provider adapters normalize to ChatCompletionRequest)
  • Fixtures are live — mutations after start() take effect immediately
  • Sequential responses are supported via sequenceIndex (match count tracked per fixture)

Match Field Reference

FieldTypeMatches Against
userMessagestringSubstring of last role: "user" message text
userMessageRegExpPattern test on last role: "user" message text
systemMessagestringSubstring of the concatenated text of every role: "system" message in the request. Use to gate a fixture on host-supplied context (persona, agent-context entries) so changes to that context cause the fixture to fall through instead of returning a stale baked response
systemMessagestring[]Array of substrings — ALL must be present in the joined system text (AND semantics). Use when the gate must combine multiple non-adjacent tokens whose serialisation order isn't stable
systemMessageRegExpPattern test on the concatenated system-message text
inputTextstringSubstring of embedding input text (concatenated if multiple inputs)
inputTextRegExpPattern test on embedding input text
toolNamestringExact match on any tool in request's tools[] array (by function.name). On the OpenAI Responses API, with responsesTools: "extended", this also includes tools inside a namespace tool, custom tools (by name, in req.customTools), and tools in an additional_tools or tool_search_output input item; by default only top-level function tools, as in 1.44.0. Other Responses tool types (web_search, mcp, ...) are not visible
toolNamespacestringExact OpenAI Responses tool namespace; needs responsesTools: "extended" (ignored otherwise). Alone: any offered tool sits in it. With toolName: one offered tool must carry both (Codex routes by this exact pair). No default namespace is assumed. Chat Completions passes request tools through, so a client-sent top-level namespace there also matches. With "extended" (aimock validate --responses-tools extended), aimock validate and --validate-on-load warn about a file fixture with toolNamespace and an endpoint other than "chat". A Responses namespace tool with an empty or missing name is dropped. Codex names an MCP server's namespace mcp__<server> (no trailing __). aimock-pytest needs aimock 1.45.0+, its default pin from 0.7.0
toolCallIdstringExact match on tool_call_id of last role: "tool" message
toolResultContainsstringSubstring of the last tool message's text content, gated on that message being the request's LAST message (same rule as toolCallId). Discriminates resume paths that share a tool_call_id and differ only inside the tool-result payload (e.g. approve {"chosen_time": …} vs cancel {"cancelled": true})
modelstringExact match on req.model
modelRegExpPattern test on req.model
responseFormatstringExact match on req.response_format.type ("json_object", "json_schema")
sequenceIndexnumberMatches only when this fixture's match count equals the given index (0-based)
turnIndexnumberStateless conversation-depth matching. Counts role: "assistant" messages in the request; matches when that count equals the value. turnIndex: 0 = first turn (no prior assistant messages). Use instead of sequenceIndex for shared/deployed instances where stateful counters break under concurrency
hasToolResultbooleanStateless tool-message presence matching, scoped to the CURRENT turn (messages after the last role: "user" message). true matches when a role: "tool" message appears after the last user message; false matches when none does. (If the request has no user message, the whole conversation is scanned.) Provider-consistent across all aimock handlers (OpenAI, Claude, Gemini, Bedrock, Ollama, Cohere)
endpointstringRestrict to endpoint type: "chat", "image", "speech", "transcription", "video", "embedding"
predicate(req: ChatCompletionRequest) => booleanCustom function — full access to request

AND logic: all specified fields must match. Empty match {} = catch-all.

Multi-part content (e.g., [{type: "text", text: "hello"}]) is automatically extracted — userMessage matching works regardless of content format.

When to Use Each Multi-turn Matching Approach

ApproachStateless?Best For
turnIndexYesShared/deployed instances; matches on conversation depth (count of assistant messages in request)
hasToolResultYesSimplest option for 2-step tool flows — boolean: does the current turn (after the last user message) carry a tool result?
sequenceIndexNoSingle-client unit tests with repeated identical requests (server-side counter, breaks under concurrency)
toolCallIdYesMatching specific tool result IDs in the conversation history
toolResultContainsYesSame tool call id, different outcomes — match on the tool-result payload (approve vs cancel legs)

Prefer stateless approaches (turnIndex, hasToolResult, toolResultContains) for shared aimock instances (deployed via Docker, used by multiple test runners). Use sequenceIndex only in isolated single-client unit tests where the counter won't be corrupted by concurrent requests.

Multi-turn fixture examples

// 2-step HITL with turnIndex
{"match": {"userMessage": "trip to mars", "turnIndex": 0}, "response": {"toolCalls": [{"id": "call_001", "name": "generate_steps", "arguments": "{}"}]}}
{"match": {"userMessage": "trip to mars", "turnIndex": 1}, "response": {"content": "Great choices! Proceeding."}}

// Same thing with hasToolResult (simpler for 2-step)
{"match": {"userMessage": "trip to mars", "hasToolResult": false}, "response": {"toolCalls": [{"id": "call_001", "name": "generate_steps", "arguments": "{}"}]}}
{"match": {"userMessage": "trip to mars", "hasToolResult": true}, "response": {"content": "Great choices!"}}

// HITL suspend tool where approve and cancel resume with the SAME tool call id —
// discriminate on the tool-result payload; put the cancel leg first (first match wins)
{"match": {"toolCallId": "call_001", "toolResultContains": "\"cancelled\""}, "response": {"content": "No problem — nothing was booked."}}
{"match": {"toolCallId": "call_001"}, "response": {"content": "Booked: Monday 9:00 AM confirmed."}}

Response Types

Text

{
  content: "Hello!";
}

Tool Calls

// Preferred: object form (auto-stringified by the fixture loader)
{
  toolCalls: [{ name: "get_weather", arguments: { city: "SF" } }];
}

// Also accepted: JSON string form (backward compatible)
{
  toolCalls: [{ name: "get_weather", arguments: '{"city":"SF"}' }];
}

Both object and string forms are accepted for arguments. The fixture loader auto-stringifies objects via JSON.stringify(). Object form is preferred for readability.

OpenAI Responses API only — namespaced and custom tool calls (used by Codex for MCP tools, sub-agents and apply_patch). Start the server with responsesTools: "extended" (--responses-tools extended): it makes namespaced, custom and additional_tools tools visible to toolName and predicates, counts custom tool rounds for hasToolResult / toolCallId / turnIndex, emits a toolCalls entry's namespace, and records namespaces and custom calls. The default ("legacy") behaves exactly like 1.44.0. toolNamespace, customToolCalls and responsesBlocks also need "extended"; the default ignores them.

const response = {
  toolCalls: [
    // function call to a tool offered inside a `namespace` tool
    {
      name: "spawn_agent",
      namespace: "collaboration",
      arguments: JSON.stringify({ message: "..." }),
    },
  ],
  // custom (freeform) tool calls: free-text `input`, never parsed or stringified,
  // emitted after the function calls. A custom-only turn sets `toolCalls: []`.
  customToolCalls: [{ name: "apply_patch", input: "*** Begin Patch\n...\n*** End Patch\n" }],
};

A toolCalls entry is always a function call; a type or input on it is ignored (as in 1.44.0). To interleave custom calls with text or function calls, use responsesBlocks (the blocks entries plus { "type": "customToolCall", "name", "input", "id"?, "namespace"? }); a fixture may not set both blocks and responsesBlocks. Every other API (Chat Completions, Anthropic Messages, Gemini, Gemini Interactions, Bedrock, Ollama, Cohere, Realtime, Gemini Live) ignores namespace and rejects a fixture with custom calls before any content with code aimock_unsupported_tool_call (Ollama /api/generate keeps its HTTP 400 for every tool-call fixture); a responsesBlocks without custom calls is served there as blocks. A malformed customToolCalls entry or responsesBlocks tool block fails validation (aimock validate, --validate-on-load, addFixturesFromJSON(), POST /__aimock/fixtures); in a programmatic, factory or unvalidated file fixture it fails the Responses request with code aimock_invalid_fixture_tool_call. Fixtures that 1.44.0 accepted get no new validation finding and no new error. Where the code goes: error.code (OpenAI Chat / Responses, Azure, Cohere, Anthropic, Gemini Interactions); error.metadata.reason (OpenRouter Chat); a google.rpc.ErrorInfo detail (Gemini, Vertex AI, Gemini Live error frame with code 13); reason (Bedrock); code next to error (Ollama /api/chat); an error event's error.code (Responses WebSocket); response.done with status: "failed" and status_details.error.code (Realtime). For Codex, gate the turn-0 fixture with hasToolResult: false and the follow-up with hasToolResult: true, or the turn-0 fixture matches again and the agent loops. A fixture that matches only on toolName / toolNamespace loops the same way, because Codex offers the same tools on every request.

Blocks (ordered text / tool-call streaming)

The optional blocks array expresses an explicit, ordered sequence of stream entries — something plain content + toolCalls cannot, since those imply text-then-tools. Each entry is { "type": "text", "text": "..." }, { "type": "toolCall", "name": "...", "arguments": "...", "id"?: "...", "namespace"?: "..." }, streamed in array order (OpenAI Responses responsesBlocks also take customToolCall blocks). This enables tool-first ordering (a tool call before any text) and interleaved text/tool ordering.

// Tool-first: tool call streams before the text
{
  blocks: [
    { type: "toolCall", name: "get_weather", arguments: { city: "SF" } },
    { type: "text", text: "Checking the weather for you…" },
  ];
}

When blocks is present it takes precedence over content/toolCalls for stream order; when absent, legacy behavior is unchanged. blocks-only fixtures are first-class — a response may be just { blocks: [...] } with no content and no toolCalls, and builders derive the aggregate content/tool_calls from the blocks. A toolCall block's arguments may be a JSON object or a string (objects auto-stringify), exactly like top-level toolCalls.

Replay caveat: block order is observable on some providers and not others — see the per-provider observability matrix.

Embedding

{
  embedding: [0.1, 0.2, 0.3, -0.5, 0.8];
}

The embedding vector is returned for each input in the request. If no embedding fixture matches, deterministic embeddings are auto-generated from the input text hash — you only need fixtures when you want specific vectors.

Image

<!-- prettier-ignore -->
// Single image
{
  image: {
    url: "https://example.com/generated.png"
  }
}
// Multiple images
{
  images: [{ url: "https://example.com/1.png" }, { b64Json: "iVBOR..." }]
}

Use match: { endpoint: "image" } to prevent cross-matching with chat fixtures.

Speech (TTS)

{ audio: "base64-encoded-audio-data" }
// With explicit format (default: mp3)
{ audio: "base64-data", format: "opus" }

Transcription

// Simple
{ transcription: { text: "Hello world" } }
// Verbose with timestamps
{ transcription: { text: "Hello world", language: "en", duration: 2.5, words: [...], segments: [...] } }

Video

{ video: { id: "vid-1", status: "completed", url: "https://example.com/video.mp4" } }

Video uses async polling — POST /v1/videos creates, GET /v1/videos/{id} checks status.

Error

{ error: { message: "Rate limited", type: "rate_limit_error" }, status: 429 }

Chaos (Failure Injection)

The optional chaos field on a fixture enables probabilistic failure injection:

{
  chaos?: {
    dropRate?: number;      // Probability (0-1) of returning a 500 error
    malformedRate?: number; // Probability (0-1) of returning malformed JSON
    disconnectRate?: number; // Probability (0-1) of disconnecting mid-stream
  }
}

Rates are evaluated per-request. When triggered, the chaos failure replaces the normal response.

Bad model output and retry recovery

Use misbehavior to change a valid fixture response. Keep chaos for transport failures. Semantic fault support varies by wire, fault, and mode; consult the support matrix before authoring a fixture. The example below uses openai-chat, including Azure and OpenRouter chat endpoints. Local support does not establish native fidelity; preserve the modeled contracts and native capture limits.

Save this fixture as bad-tool.json. Its first matching request receives invalid tool JSON; the next receives valid arguments.

{
  "fixtures": [
    {
      "match": { "userMessage": "weather", "endpoint": "chat" },
      "response": {
        "toolCalls": [{ "id": "call_weather", "name": "weather", "arguments": { "city": "Paris" } }]
      },
      "misbehavior": {
        "seed": 42,
        "faults": [
          {
            "fault": "tool-args-invalid-json",
            "style": "trailing-comma",
            "times": 1,
            "providers": ["openai-chat"]
          }
        ]
      }
    }
  ]
}

Check application parsing and recovery explicitly. An HTTP 200 with bad arguments does not trigger the SDK's HTTP retry logic. For Vitest, use the real local server and official SDK:

import OpenAI from "openai";
import { LLMock } from "@copilotkit/aimock";
import { expect } from "vitest";

const mock = new LLMock();
mock.loadFixtureFile("bad-tool.json");
await mock.start();
try {
  const client = new OpenAI({
    apiKey: "local",
    baseURL: `${mock.url}/v1`,
    maxRetries: 0,
    defaultHeaders: { "X-Test-Id": "retry-tool" },
  });
  const request = {
    model: "gpt-4",
    messages: [{ role: "user", content: "weather" }],
  } as const;
  const first = await client.chat.completions.create({
    ...request,
    messages: [...request.messages],
  });
  const bad = first.choices[0].message.tool_calls?.[0].function.arguments;
  expect(bad).toBeDefined();
  expect(() => JSON.parse(bad!)).toThrow(SyntaxError);
  const second = await client.chat.completions.create({
    ...request,
    messages: [...request.messages],
  });
  expect(JSON.parse(second.choices[0].message.tool_calls![0].function.arguments)).toEqual({
    city: "Paris",
  });
  expect(mock.journal.getAll()[0].response.misbehavior).toMatchObject({
    applied: true,
    source: "fixture",
    wire: "openai-chat",
    fault: "tool-args-invalid-json",
  });
} finally {
  await mock.stop();
}

Keep the same test ID for both attempts. times counts firings per source, entry and test ID, not matching requests. Use mock.resetMatchCounts("retry-tool") to restart that test's counters. Clearing journal entries alone preserves fault counters.

A configuration accepts a fault ID string or { seed, faults }. Use { faults: [] } for an explicit opt-out. Only POST /__aimock/misbehavior accepts bare {} as an opt-out. Each entry accepts rate from 0 to 1, positive integer times, a tool selector and a providers wire list. The default rate is 1. An omitted times has no firing limit. The default seed is 0; "random" selects one process seed.

FaultFault-specific configuration
tool-args-invalid-jsonstyle: truncated (default), trailing-comma, single-quotes
tool-args-schema-violationviolation: missing-required (default), wrong-type, extra-property, enum-mismatch, not-object; optional property
tool-unknown-nameOptional name
tool-call-id-duplicateNo extra parameters; requires at least one call on an ID-emitting wire; duplicates the sole call when only one exists
stop-length-mid-toolat strictly between 0 and 1; default 0.5; emits a clean length stop with partial arguments
empty-responseNo extra parameters
refusalOptional message; keep category null or omitted on OpenAI Chat
content-filterNo extra parameters
reasoning-onlyOptional reasoning; OpenRouter emits reasoning with reasoning_details; native OpenAI exposes no reasoning

For schema violations, send the tool schema in the request. Use direct schema constraints; do not assume $ref or combinator support. Selection order is HTTP header, fixture, runtime scope, then server baseline. Selected configurations replace one another; they do not merge. Fixture and header faults that cannot apply fail loudly. Runtime and server faults that cannot apply are skipped. Bad fixture configuration throws FixtureLoadError with a misbehavior/ rule, including CLI loads without --validate-on-load. Bad server configuration or setter input throws a rule-prefixed TypeError. Check response.misbehavior.evaluations in the journal for skips, exhausted budgets and provider exclusions. Check actual SDK output too: servedToolCalls describes the prepared response, not proof of network delivery. See the control API for runtime scope, validation errors and header grammar.

Common Patterns

Basic text fixture

mock.onMessage("hello", { content: "Hi there!" });

Tool call → tool result → final response (3-step agent loop)

The most common pattern. Fixture 1 triggers the tool call, fixture 2 handles the tool result.

// Step 1: User asks about weather → LLM calls tool
mock.onMessage("weather", {
  toolCalls: [{ name: "get_weather", arguments: { city: "SF" } }],
});

// Step 2: Tool result comes back → LLM responds with text
mock.addFixture({
  match: { predicate: (req) => req.messages.at(-1)?.role === "tool" },
  response: { content: "It's 72°F in San Francisco." },
});

Why predicate, not userMessage? After a tool call, the client replays the same conversation with the tool result appended. The user message hasn't changed — userMessage: "weather" would match the SAME fixture again, creating an infinite loop.

Embedding fixture

// Match specific input text
mock.onEmbedding("search query", {
  embedding: [0.1, 0.2, 0.3, 0.4, 0.5],
});

// Match with regex
mock.onEmbedding(/product.*description/, {
  embedding: [0.9, -0.1, 0.5, 0.3, 0.2],
});

Structured output / JSON mode

// onJsonOutput auto-sets responseFormat: "json_object" and stringifies objects
mock.onJsonOutput("extract entities", {
  entities: [
    { name: "Acme Corp", type: "company" },
    { name: "Jane Doe", type: "person" },
  ],
});

// Equivalent manual form:
mock.addFixture({
  match: { userMessage: "extract entities", responseFormat: "json_object" },
  response: { content: '{"entities":[...]}' },
});

Sequential responses (same match, different responses)

// First call returns tool call, second returns text
mock.on(
  { userMessage: "status", sequenceIndex: 0 },
  { toolCalls: [{ name: "check_status", arguments: {} }] },
);
mock.on({ userMessage: "status", sequenceIndex: 1 }, { content: "All systems operational." });

Match counts are tracked per fixture group. Use resetMatchCounts() between tests to reset counts while keeping loaded fixtures. reset() also clears the fixture pool, so avoid it between tests that share a loaded fixture set.

Streaming physics (realistic timing)

mock.onMessage(
  "tell me a story",
  { content: "Once upon a time..." },
  {
    streamingProfile: {
      ttft: 200, // 200ms before first token
      tps: 30, // 30 tokens per second after that
      jitter: 0.1, // ±10% random variance
    },
  },
);

Predicate-based routing (same user message, different context)

Common in supervisor/orchestrator patterns where the system prompt changes:

mock.addFixture({
  match: {
    predicate: (req) => {
      const sys = req.messages.find((m) => m.role === "system")?.content ?? "";
      return typeof sys === "string" && sys.includes("Flights found: false");
    },
  },
  response: { toolCalls: [{ name: "search_flights", arguments: {} }] },
});

Catch-all (always add one)

Prevents unmatched requests from returning 404 and crashing the test:

mock.addFixture({
  match: { predicate: () => true },
  response: { content: "I understand. How can I help?" },
});

Tool result catch-all with prependFixture

Must go at the front so it matches before substring-based fixtures:

mock.prependFixture({
  match: { predicate: (req) => req.messages.at(-1)?.role === "tool" },
  response: { content: "Done!" },
});

Stream interruption simulation (v1.3.0+)

mock.onMessage(
  "long response",
  { content: "This will be cut short..." },
  {
    truncateAfterChunks: 3, // Stop after 3 SSE chunks
    disconnectAfterMs: 500, // Or disconnect after 500ms
  },
);

Chaos testing (probabilistic failures)

mock.addFixture({
  match: { userMessage: "flaky" },
  response: { content: "Sometimes works!" },
  chaos: { dropRate: 0.3 },
});

30% of requests matching this fixture will get a 500 error instead of the response. Can also use malformedRate (garbled JSON) or disconnectRate (connection dropped mid-stream).

Server-level chaos applies to ALL requests:

mock.setChaos({ dropRate: 0.1 }); // 10% of all requests fail
mock.clearChaos(); // Remove server-level chaos

Error injection (one-shot)

mock.nextRequestError(429, { message: "Rate limited", type: "rate_limit_error" });
// Next request gets 429, then fixture auto-removes itself

JSON fixture files

{
  "fixtures": [
    {
      "match": { "userMessage": "hello" },
      "response": { "content": "Hi!" }
    },
    {
      "match": { "userMessage": "weather" },
      "response": {
        "toolCalls": [
          {
            "name": "get_weather",
            "arguments": { "city": "SF", "units": "fahrenheit" }
          }
        ]
      }
    },
    {
      "match": { "inputText": "search query" },
      "response": { "embedding": [0.1, 0.2, 0.3] }
    },
    {
      "match": { "userMessage": "status", "sequenceIndex": 0 },
      "response": { "content": "First response" }
    }
  ]
}

JSON auto-stringify: In JSON fixture files, arguments and content can be objects — the loader auto-stringifies them with JSON.stringify(). This also applies to a blocks entry's arguments — object form auto-stringifies just like top-level toolCalls. The escaped-string form ("{\"city\":\"SF\"}") still works but objects are preferred for readability.

JSON files cannot use RegExp or predicate — those are code-only features. streamingProfile is supported in JSON fixture files.

Load with mock.loadFixtureFile("./fixtures/greetings.json") or mock.loadFixtureDir("./fixtures/").

API Endpoints

All providers share the same fixture pool — write fixtures once, they work for any endpoint.

EndpointProviderProtocol
POST /v1/chat/completionsOpenAIHTTP
POST /v1/responsesOpenAIHTTP + WS
POST /v1/messagesAnthropicHTTP
POST /v1/embeddingsOpenAIHTTP
POST /v1beta/models/{model}:{method}Google GeminiHTTP
POST /model/{modelId}/invokeAWS BedrockHTTP
POST /openai/deployments/{id}/chat/completionsAzure OpenAIHTTP
POST /openai/deployments/{id}/embeddingsAzure OpenAIHTTP
GET /health—HTTP
GET /ready—HTTP
POST /model/{modelId}/invoke-with-response-streamAWS BedrockHTTP
POST /model/{modelId}/converseAWS BedrockHTTP
POST /model/{modelId}/converse-streamAWS BedrockHTTP
POST /v1/projects/{p}/locations/{l}/publishers/google/models/{m}:generateContentVertex AIHTTP
POST /v1/projects/{p}/locations/{l}/publishers/google/models/{m}:streamGenerateContentVertex AIHTTP
POST /api/chatOllamaHTTP
POST /api/generateOllamaHTTP
GET /api/tagsOllamaHTTP
POST /v2/chatCohereHTTP
POST /api/v1/chat/completionsOpenRouterHTTP
GET /api/v1/models · /api/v1/key · /api/v1/creditsOpenRouterHTTP
GET /metrics—HTTP
GET /v1/modelsOpenAI-compatHTTP
WS /v1/responsesOpenAIWebSocket
WS /v1/realtimeOpenAIWebSocket
WS /ws/google.ai...BidiGenerateContentGemini LiveWebSocket
POST /v1/images/generationsOpenAIHTTP
POST /v1beta/models/{model}:predictGemini ImagenHTTP
POST /v1/audio/speechOpenAIHTTP
POST /v1/audio/transcriptionsOpenAIHTTP
POST /v1/videosOpenAIHTTP
GET /v1/videos/{id}OpenAIHTTP

Response Template Overrides

Fixture responses can include optional override fields to control auto-generated envelope values. These are merged into the provider-specific response format (OpenAI, Claude, Gemini, Responses API).

FieldTypeDefaultDescription
idstringauto-generatedOverride response ID (e.g., chatcmpl-custom)
creatednumberDate.now()/1000Override Unix timestamp
modelstringechoes requestOverride model name in response
usageobjectzeroedOverride token counts: { prompt_tokens, completion_tokens, total_tokens }. OpenAI Chat includes usage in response body; Responses API uses response.usage. When omitted, auto-computed from content length
finishReasonstring"stop" / "tool_calls"Override finish reason. Mappings: stop -> end_turn (Claude), STOP (Gemini); tool_calls -> tool_use (Claude), FUNCTION_CALL (Gemini); length -> max_tokens (Claude), MAX_TOKENS (Gemini); content_filter -> SAFETY (Gemini), failed (Responses API)
rolestring"assistant"Override message role
systemFingerprintstring(omitted)Add system_fingerprint to response
providerstringslug authorOpenRouter only: top-level serving-provider display name (default = the winning model slug's author). Override to assert who served the request
nativeFinishReasonstringmirrors finishReasonOpenRouter only: the raw upstream native_finish_reason alongside the normalized finish_reason
usage.costnumber(omitted)OpenRouter only: per-request usage.cost (scriptable — powers budget-guard tests). When set, usage.cost_details is emitted too. Never fabricated when omitted
usage.is_byokbool(omitted)OpenRouter only: emit usage.is_byok. Also usage.prompt_tokens_details, usage.completion_tokens_details — emitted only when set

Example

mock.onMessage("hello", {
  content: "Hi!",
  model: "gpt-4-turbo-2024-04-09",
  usage: { prompt_tokens: 10, completion_tokens: 5, total_tokens: 15 },
  systemFingerprint: "fp_abc123",
});

In JSON fixtures

{
  "match": { "userMessage": "hello" },
  "response": {
    "content": "Hi!",
    "model": "gpt-4-turbo-2024-04-09",
    "usage": { "prompt_tokens": 10, "completion_tokens": 5, "total_tokens": 15 },
    "systemFingerprint": "fp_abc123"
  }
}

These fields map correctly across all provider formats — for example, finishReason: "stop" becomes finish_reason: "stop" in OpenAI, stop_reason: "end_turn" in Claude, and finishReason: "STOP" in Gemini.

OpenRouter (chat / router)

A request whose path starts with /api/v1/ (point the OpenAI SDK at a baseURL ending /api/v1) is shaped as OpenRouter: gen- id, top-level provider, per-choice native_finish_reason, always-present system_fingerprint/service_tier (null by default), an always-present message.reasoning (null unless a fixture supplies reasoning and the model is reasoning-capable), and a rich usage. Requests on the plain /v1/... base are untouched OpenAI. Same fixture pool — the fields above are the only additions.

  • Scriptable cost / provider / finish reason: set provider, nativeFinishReason, and usage.cost on the response (see the overrides table). cost/cost_details are emitted only when a fixture supplies cost — aimock never fabricates a cost.
  • models[] fallback (router failover): when the request body carries models: [m1, m2, ...], aimock walks [model, ...models] in order and serves the first fixture that returns a NON-error response. A 429/503 error fixture on a candidate simulates a RUNTIME provider failure and falls through to the next candidate; the winning slug is echoed back as the top-level model (assert failover via response.model). Model the primary's "failure" as a 429/503 — an unknown/invalid model is just a fixture miss (aimock does not replicate OpenRouter's up-front invalid-model 400).
  • Terminal (non-failover) error class — fallthrough: false: real OpenRouter fails over inconsistently by error class (a 403 budget-exceeded / generic "provider returned error" is served as terminal and does NOT advance to the next candidate, while 429/503 usually do). Set fallthrough: false on an error fixture to make it terminal: the fallback loop stops and serves that error even when a good candidate follows. Absent / true keeps the default fall-through. Composes with provider.allow_fallbacks: fall-through happens only when both allow it (if either says stop, the error is terminal). Use it to reproduce the exact provider error a dev's app must handle itself. { error: { message: "budget exceeded" }, status: 403, fallthrough: false }
  • Keepalive: set the fixture option openRouterProcessing: true to emit one : OPENROUTER PROCESSING SSE comment before the first data frame (opt-in, default off).
  • Error envelope: OpenRouter errors are { error: { message, code } } (numeric code == HTTP status), with optional free-form metadata.
// primary is a runtime 429, fallback answers
mock.on(
  { model: "openai/gpt-4o", userMessage: "route" },
  { error: { message: "rate limited" }, status: 429 },
);
mock.on(
  { model: "anthropic/claude-3.5-sonnet", userMessage: "route" },
  {
    content: "served by the fallback",
    provider: "Anthropic",
    usage: { cost: 0.0021 },
  },
);
// POST /api/v1/chat/completions { model: "openai/gpt-4o", models: [...], ... }
//   → response.model === "anthropic/claude-3.5-sonnet"

Provider Support Matrix

FeatureOpenAI ChatOpenAI ResponsesClaudeGeminiGemini Int.BedrockAzureOllamaCohereOpenRouter
TextYesYesYesYesYesYesYesYesYesYes
Tool CallsYesYesYesYesYesYesYesYesYesYes
Content + Tool CallsYesYesYesYesYesYesYesYesYesYes
StreamingSSESSESSESSESSEBinarySSENDJSONSSESSE
ReasoningYesYesYesYes--YesYes----Yes
Web Searches--Yes----------------
Response OverridesYesYesYesYesYes--Yes----Yes
Custom Tool CallsErrorYesErrorErrorErrorErrorErrorErrorErrorError

Custom tool calls (customToolCalls / customToolCall blocks in responsesBlocks): "Error" means the request fails before any content with code aimock_unsupported_tool_call. The Ollama column is /api/chat; /api/generate rejects every tool-call fixture, custom calls included, with HTTP 400 (invalid_request_error) and no aimock code. The Azure column is the deployment route (/openai/deployments/{id}/chat/completions) and the OpenRouter column is /api/v1/chat/completions. aimock serves /openai/v1/responses and OpenRouter's /api/v1/responses with its OpenAI Responses handler, so custom tool calls work there. Realtime and Gemini Live WebSockets also reject them.

Critical Gotchas

  1. Order matters — first match wins. Specific fixtures before general ones. Use prependFixture() to force priority.

  2. arguments accepts both objects and strings — "arguments": {"key":"value"} (preferred, auto-stringified) or "arguments": "{\"key\":\"value\"}" (legacy). The same applies to content fields that contain JSON. The fixture loader detects typeof === "object" and calls JSON.stringify() automatically.

  3. Latency is per-chunk, not total — latency: 100 means 100ms between each SSE chunk, not 100ms total response time. Similarly, truncateAfterChunks and disconnectAfterMs are for simulating stream interruptions (added in v1.3.0).

  4. streamingProfile takes precedence over latency — when both are set on a fixture, streamingProfile controls timing. Use one or the other.

  5. Tool result messages don't change the user message — after a tool call, the client sends the same conversation + tool result. Matching on userMessage will hit the SAME fixture again → infinite loop. Always use predicate checking role === "tool" for tool results. Note: a whole-conversation role === "tool" check (e.g. req.messages.some((m) => m.role === "tool")) diverges from hasToolResult's current-turn scoping in multi-turn flows — the built-in hasToolResult matcher only looks after the last user message, so a later turn whose history carries an earlier tool result still reads false.

  6. clearFixtures() preserves the array reference — uses .length = 0, not reassignment. The running server reads the same array object.

  7. Journal records everything — including 404 "no match" responses. Use mock.getLastRequest() to debug mismatches.

  8. All providers share fixtures — a fixture matching "hello" works whether the request comes via /v1/chat/completions (OpenAI), /v1/messages (Anthropic), Gemini, Bedrock, or Azure endpoints.

  9. WebSocket uses the same fixture pool — no special setup needed for WebSocket-based APIs (OpenAI Responses WS, Realtime, Gemini Live).

  10. Embeddings auto-generate if no fixture matches — deterministic vectors are generated from the input text hash. You don't need a catch-all for embedding requests.

  11. Sequential response counts are tracked per fixture — use resetMatchCounts() between tests to reset counts while keeping loaded fixtures; reset() also clears the fixture pool, so don't use it between tests that share a loaded fixture set. The count increments after each match of that fixture group (all fixtures sharing the same non-sequenceIndex match fields).

  12. Bedrock uses Anthropic Messages format internally — the adapter normalizes Bedrock requests to ChatCompletionRequest, so the same fixtures work. Bedrock supports both non-streaming (/invoke, /converse) and streaming (/invoke-with-response-stream, /converse-stream) endpoints.

  13. Azure OpenAI routes through the same handlers — /openai/deployments/{id}/chat/completions maps to the completions handler, /openai/deployments/{id}/embeddings maps to the embeddings handler. Fixtures work unchanged.

  14. Ollama defaults to streaming — opposite of OpenAI. Set stream: false explicitly in the request for non-streaming responses.

  15. Ollama tool call arguments is an object, not a JSON string — unlike OpenAI where arguments is a JSON string, Ollama sends and expects a plain object.

  16. Bedrock streaming uses binary Event Stream format — not SSE. The invoke-with-response-stream and converse-stream endpoints use AWS Event Stream binary encoding.

  17. Vertex AI routes to the same handler as consumer Gemini — the same fixtures work for both Vertex AI (/v1/projects/.../models/{m}:generateContent) and consumer Gemini (/v1beta/models/{model}:generateContent).

  18. Cohere requires model field — returns 400 if model is missing from the request body.

Mount & Composition

mount() API

Mount additional mock services onto a running LLMock server. All services share one port, one health endpoint, and one request journal.

const llm = new LLMock({ port: 5555 });
llm.mount("/mcp", mcpMock); // MCP tools at /mcp
llm.mount("/a2a", a2aMock); // A2A agents at /a2a
llm.mount("/vector", vectorMock); // Vector DB at /vector
await llm.start();

Any object implementing the Mountable interface (a handleRequest method that returns boolean) can be mounted. Path prefixes are stripped before the service sees the request — /mcp/tools/list arrives as /tools/list.

createMockSuite()

Unified lifecycle for LLMock + mounted services:

import { createMockSuite } from "@copilotkit/aimock";

const suite = createMockSuite({
  port: 0,
  fixtures: "./fixtures",
  services: { "/mcp": mcpMock, "/a2a": a2aMock },
});

await suite.start();
// suite.llm — the LLMock instance
// suite.url — base URL

afterEach(() => suite.llm.resetMatchCounts()); // reset sequence counts, keep fixtures
afterAll(() => suite.stop());

aimock CLI config file

The aimock CLI reads a JSON config and serves all services on one port:

aimock --config aimock.json --port 4010

Config format:

{
  "llm": {
    "fixtures": "./fixtures",
    "latency": 0,
    "metrics": true
  },
  "services": {
    "/mcp": { "type": "mcp", "tools": "./mcp-tools.json" },
    "/a2a": { "type": "a2a", "agents": "./a2a-agents.json" }
  }
}

VectorMock

Mock vector database server for testing RAG pipelines. Supports Pinecone, Qdrant, and ChromaDB API formats.

import { VectorMock } from "@copilotkit/aimock";

const vector = new VectorMock();

// Create a collection and register query results
vector.addCollection("docs", { dimension: 1536 });
vector.onQuery("docs", [
  { id: "doc-1", score: 0.95, metadata: { title: "Getting Started" } },
  { id: "doc-2", score: 0.87, metadata: { title: "API Reference" } },
]);

// Upsert vectors
vector.upsert("docs", [
  { id: "v1", values: [0.1, 0.2, ...], metadata: { title: "Intro" } },
]);

// Dynamic query handler
vector.onQuery("docs", (query) => {
  return [{ id: "result", score: 1.0, metadata: { topK: query.topK } }];
});

// Standalone or mounted
const url = await vector.start();
// Or: llm.mount("/vector", vector);

VectorMock endpoints

ProviderEndpoints
PineconePOST /query, POST /vectors/upsert, POST /vectors/delete, GET /describe-index-stats
QdrantPOST /collections/{name}/points/search, PUT /collections/{name}/points, POST /collections/{name}/points/delete
ChromaDBPOST /api/v1/collections/{id}/query, POST /api/v1/collections/{id}/add, GET /api/v1/collections, DELETE /api/v1/collections/{id}

Service Mocks (Search / Rerank / Moderation)

Built-in mocks for common AI-adjacent services. Registered on the LLMock instance directly — no separate server needed.

Search (Tavily-compatible)

// POST /search — matches request `query` field
mock.onSearch("weather", [
  { title: "Weather Report", url: "https://example.com", content: "Sunny today" },
]);
mock.onSearch(/stock\s+price/i, [
  { title: "ACME Stock", url: "https://example.com", content: "$42", score: 0.95 },
]);

Rerank (Cohere-compatible)

// POST /v2/rerank — matches request `query` field
mock.onRerank("machine learning", [
  { index: 0, relevance_score: 0.99 },
  { index: 2, relevance_score: 0.85 },
]);

Moderation (OpenAI-compatible)

// POST /v1/moderations — matches request `input` field
mock.onModerate("violent", {
  flagged: true,
  categories: { violence: true, hate: false },
  category_scores: { violence: 0.95, hate: 0.01 },
});

// Catch-all — everything passes
mock.onModerate(/.*/, { flagged: false, categories: {} });

Pattern matching

All three services use the same matching logic:

  • String patterns — case-insensitive substring match
  • RegExp patterns — full regex test
  • First match wins — register specific patterns before catch-alls

Debugging Fixture Mismatches

When a fixture doesn't match:

  1. Inspect what the server received: mock.getLastRequest() → check body.messages array
  2. Check fixture order: mock.getFixtures() returns fixtures in registration order
  3. For userMessage: match is against the LAST role: "user" message only, substring match (not exact)
  4. Check the journal: mock.getRequests() shows all requests including which fixture matched (or null for 404)

E2E Test Setup Pattern

import { LLMock } from "@copilotkit/aimock";

// Setup — port: 0 picks a random available port
const mock = new LLMock({ port: 0 });
mock.loadFixtureDir("./fixtures");
await mock.start();
process.env.OPENAI_BASE_URL = `${mock.url}/v1`;

// Per-test cleanup — reset sequence match counts, keep the loaded fixtures
afterEach(() => mock.resetMatchCounts());

// Teardown
afterAll(async () => await mock.stop());

Static factory shorthand

const mock = await LLMock.create({ port: 0 }); // creates + starts in one call

API Quick Reference

MethodPurpose
addFixture(f)Append fixture (last priority)
addFixtures(f[])Append multiple
prependFixture(f)Insert at front (highest priority)
clearFixtures()Remove all fixtures
getFixtures()Read current fixture list
on(match, response, opts?)Shorthand for addFixture
onMessage(pattern, response, opts?)Match by user message
onEmbedding(pattern, response, opts?)Match by embedding input text
onJsonOutput(pattern, json, opts?)Match by user message with responseFormat
onToolCall(name, response, opts?)Match by tool name in tools[]
onToolResult(id, response, opts?)Match by tool_call_id
onTurn(turn, pattern, response, opts?)Match by turn index + user message
nextRequestError(status, body?)One-shot error, auto-removes
loadFixtureFile(path)Load JSON fixture file
loadFixtureDir(path)Load all JSON files in directory
start()Start server, returns URL
stop()Stop server
reset()Clear fixtures + journal + match counts
resetMatchCounts()Clear sequence match counts only
getRequests()All journal entries
getLastRequest()Most recent journal entry
clearRequests()Clear journal only
setChaos(opts)Set server-level chaos rates
clearChaos()Remove server-level chaos
onSearch(pattern, results)Match search requests by query
onRerank(pattern, results)Match rerank requests by query
onModerate(pattern, result)Match moderation requests by input
onImage(pattern, response)Match image generation by prompt
onSpeech(pattern, response)Match TTS by input text
onTranscription(response)Match audio transcription
onVideo(pattern, response)Match video generation by prompt
mount(path, handler)Mount a Mountable (VectorMock, etc.)
url / baseUrlServer URL (throws if not started)
portServer port number

Between tests that share a loaded fixture set, use resetMatchCounts() (not reset(), which also clears fixtures). For a MockSuite, call suite.llm.resetMatchCounts() — the suite itself has no resetMatchCounts().

Sequential responses use on() with sequenceIndex in the match — there is no dedicated convenience method.

Record-and-Replay (VCR Mode)

aimock supports a VCR-style record-and-replay workflow for ALL endpoints including multimedia (image, TTS, transcription, video): unmatched requests are proxied to real provider APIs, and the responses are saved as standard aimock fixture files for deterministic replay. Binary TTS responses are base64-encoded with format derived from Content-Type. Multimedia fixtures automatically include endpoint in their match criteria for correct routing on replay.

CLI usage

# Record mode: proxy unmatched requests to real OpenAI and Anthropic APIs
aimock --record \
  --provider-openai https://api.openai.com \
  --provider-anthropic https://api.anthropic.com \
  -f ./fixtures

# Strict mode: fail on unmatched requests (no proxying, no catch-all 404)
aimock --strict -f ./fixtures
  • --record enables proxy-on-miss. Requires at least one --provider-* flag.
  • --strict returns a 503 error when no fixture matches AND no proxy is configured (or the proxy attempt fails), instead of silently returning a 404. The proxy is still tried first when --record is set. Use this in CI to prevent unmatched requests from slipping through as silent 404s.
  • Provider flags: --provider-openai, --provider-anthropic, --provider-gemini, --provider-vertexai, --provider-bedrock, --provider-azure, --provider-ollama, --provider-cohere.

How it works

  1. Existing fixtures are served first — the router checks all loaded fixtures before considering the proxy.
  2. Misses are proxied — if no fixture matches and recording is enabled, the request is forwarded to the real provider API. Upstream URL path prefixes are preserved (e.g., https://gateway.company.com/llm/v1 correctly proxies to /llm/v1/chat/completions).
  3. All request headers are forwarded (auth headers NOT saved) — all client request headers are passed through to the upstream provider, except hop-by-hop headers and host/content-length/cookie/accept-encoding. Auth headers (Authorization, x-api-key, api-key) are forwarded but stripped from the recorded fixture.
  4. Responses are saved as standard fixtures — recorded files land in {fixturePath}/recorded/ and use the same JSON format as hand-written fixtures. Nothing special about them.
  5. Streaming responses are collapsed — SSE streams are collapsed into a single text or tool-call response for the fixture. The original streaming format is preserved in the live proxy response.
  6. Base64 embedding decoding — when the upstream returns base64-encoded embeddings (the default encoding_format in Python's openai SDK), the recorder decodes them into float arrays so fixtures contain readable numeric data instead of opaque base64 strings.
  7. Loud logging — every proxy hit logs at warn level so you can see exactly which requests are being forwarded.

Programmatic API

const mock = new LLMock({ port: 0 });
await mock.start();

// Enable recording at runtime
mock.enableRecording({
  providers: {
    openai: "https://api.openai.com",
    anthropic: "https://api.anthropic.com",
  },
  fixturePath: "./fixtures/recorded",
});

// ... run tests that hit real APIs for uncovered cases ...

// Disable recording (back to fixture-only mode)
mock.disableRecording();

Workflow

  1. Bootstrap: Run your test suite with --record and provider URLs. All requests that don't match existing fixtures are proxied and recorded.
  2. Review: Check the recorded fixtures in {fixturePath}/recorded/. Edit or reorganize as needed.
  3. Lock down: Run your test suite with --strict to ensure every request hits a fixture. No network calls escape.
  4. Maintain: When APIs change, delete stale fixtures and re-record.

MCP Recording and Fake Reports (1.45.0+)

Record a live MCP server

npx -p @copilotkit/aimock llmock -f ./fixtures \
  --mcp-record /mcp=http://localhost:3001/mcp
  • --mcp-record <mount>=<url> (repeatable) is on the llmock bin and the Docker image, not on aimock --config. It needs a local --fixtures path. --mcp-proxy-only <mount>=<url> forwards without writing and needs no --fixtures.
  • With aimock --config, set llm.enableMcpRecording: true and use llm.record.mcp: { "/mcp": "http://localhost:3001/mcp" }, or an object with upstream, fixturePath, proxyOnly, upstreamAuth, secretValues, strict, maxRecordBufferBytes. Without the opt-in, llm.record.mcp is ignored with a warning. llm.record.mcp alone does not turn on LLM recording.
  • In code: mcpMock.enableRecording({ upstream, ... }) and disableRecording().
  • Every MCP request needs a test id (X-Test-Id, percent-encoded, or ?testId=) or a context. Without one, the call is forwarded but not written.
  • The recording lands at <fixtures>/recorded/<slugified test id>/mcp.json. Replay it by starting aimock with the same --fixtures and no --mcp-record. A changed call then fails with MCP_FAKE_MISMATCH.
  • A call the file already answers is answered from the file, not forwarded. Delete the file to re-record.
  • Secrets: headers are never written; known secrets become [REDACTED] with a _warnings pointer. Add custom secret values with AIMOCK_RECORD_SECRET_VALUES (newline-separated, each at least 8 characters). Set an upstream credential with AIMOCK_MCP_UPSTREAM_AUTH="Name: value". OAuth is not supported through the recorder.
  • Recorded mcpFakes files may hold list, recorded and timing (block) and notifications and durationMs (call entry). Older aimock versions reject them. Replay sends recorded progress notifications (as SSE) at the recorded timing, scaled by --replay-speed. Recorded log notifications replay only after the client calls logging/setLevel.

Fail a test on a swallowed or unused fake

  • GET /__aimock/mcp/fakes/report?testId=<id> returns { ok, served, unconsumed, failures, unfaked, sharedUnconsumed, evicted } for one test id.
  • useAimock({ fixtures, fakesReport: "fail" }) (Vitest or Jest) checks each test's report in afterEach and throws AimockFakesReportError when a fake error was swallowed by the agent or a declared entry was never used. "warn" prints instead.
  • mock().fakesFor() returns { testId, mcpUrl, headers } for the current test. The default test id is <file> › <describe> › <test> (Vitest; Jest joins the names with spaces). Scope the mcpFakes block to that id. Pass an explicit id for concurrent tests and, in Jest, for duplicate test names.
  • assertFakesReport(report) does the same check anywhere. In pytest: aimock.fakes_for(), aimock.fakes_report() and aimock.assert_fakes_report(); the default id is the pytest node id.
  • Fakes work for any caller over MCP HTTP: agents, Mastra Workflow steps, LangGraph nodes and plain functions.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

このスキルの問題を報告する