Design agents for multi-turn tool use — sequential tool calls, result accumulation, error recovery, and complex task decomposition over multiple turns.
日本語の概要は準備中です。原文の説明を表示しています。
63 件 ・ 関連度順
概要と使いどころ
Design agents for multi-turn tool use — sequential tool calls, result accumulation, error recovery, and complex task decomposition over multiple turns.
日本語の概要は準備中です。原文の説明を表示しています。
OpenAI GPT Image generation and editing (gpt-image-2, 1.5, 1, mini). Text-to-image, mask-based inpainting, multi-reference composition, multi-turn conversational editing via Responses API, streaming with partial images. This skill should be used when generating or editing images via OpenAI's image models, when near-perfect text rendering in images is needed, when mask-based region-aware editing is required, when multi-turn conversational image editing is desired, or when streaming progressive image delivery is needed. Complements nano-banana-pro (Gemini) as a parallel image generation backend.
日本語の概要は準備中です。原文の説明を表示しています。
Design Agentforce conversations that span multiple turns without losing context: session variable scoping, conversation memory, clarifying-question patterns, topic-to-topic (now subagent) handoff, and the right abstractions for accumulating state across turns. NOT for deciding the topic boundaries themselves or out-of-scope behavior — use agentforce/agent-topic-design. NOT for single-turn agent actions and their input/output contracts — use agentforce/agent-actions.
日本語の概要は準備中です。原文の説明を表示しています。
Set up and run multi-turn RL training for interactive environments (terminal tasks, tool use, search/RAG, games) using the Tinker API. Use when the user wants multi-turn RL, agentic training, tool-use RL, or interactive environment training.
日本語の概要は準備中です。原文の説明を表示しています。
Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.
日本語の概要は準備中です。原文の説明を表示しています。
Generates images and text via reverse-engineered Gemini Web API. Supports text generation, image generation from prompts, reference images for vision input, and multi-turn conversations. Use when other skills need image generation backend, or when user requests "generate image with Gemini", "Gemini text generation", or needs vision-capable AI generation.
日本語の概要は準備中です。原文の説明を表示しています。
Guides the usage of Gemini Interactions API on Gemini Enterprise Agent Platform. Use when the user wants to use the stateful, server-managed Interactions API for multi-turn conversations, background execution, streaming, structured output, and function calling on the Agent Platform.
日本語の概要は準備中です。原文の説明を表示しています。
DeepEval evaluation workflow for AI agents and LLM applications. TRIGGER when the user wants to evaluate or improve an AI agent, tool-using workflow, multi-turn chatbot, RAG pipeline, or LLM app; add evals; generate datasets or goldens; use deepeval generate; use deepeval test run; send results to Confident AI; monitor production; run online evals; inspect traces; or iterate on prompts, tools, retrieval, or agent behavior from eval failures. AI agents are the primary use case. Covers Python SDK, pytest eval suites, CLI generation, traced evals, Confident AI reporting, and agent-driven improvement loops. DO NOT TRIGGER for unrelated generic pytest, non-AI test setup, or non-DeepEval observability work unless the user asks to compare or migrate to DeepEval; for instrumenting an app with DeepEval tracing, @observe, or framework integrations (use the `deepeval-tracing` skill); or for raw OpenTelemetry / OTLP export without the deepeval package (use the `deepeval-otel` skill).
日本語の概要は準備中です。原文の説明を表示しています。
Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.
日本語の概要は準備中です。原文の説明を表示しています。
Run the Omnigent load test and produce a results file explaining the latencies. Load when the user wants to load-test / stress-test / benchmark Omnigent under concurrency ("load test omnigent", "stress test the server", "how many hosts/sessions/turns can it handle", "load test real agent turns / conversations", "run a load test"). The test makes each simulated user a real omnigent host that creates host-bound sessions and drives real multi-turn conversations with a mocked LLM; it boots its own local stack (dev/loadtest/run.py). Gather inputs, run it, then read the generated summary.md and explain the latency distribution (avg/median/p95/p99, throughput, failures). NOT for single-request latency micro-benchmarks (that is dev/benchmarks/).
日本語の概要は準備中です。原文の説明を表示しています。
Facilitate workshop sessions in a one-step, multi-turn flow. Use when an interactive skill needs consistent pacing, options, and progress tracking.
日本語の概要は準備中です。原文の説明を表示しています。
LLM application red-teaming — prompt injection (direct + indirect), jailbreak, system-prompt leak, data exfiltration, guardrail bypass, multi-turn crescendo, cross-lingual + cipher + invisible-unicode token smuggling, excessive agency / tool abuse, insecure output handling. Canonical OWASP LLM Top 10 + ASI01-ASI10 mapping. Use when a target exposes a chat/completions/assistant/copilot endpoint, an AI feature that consumes user text or documents, or any /v1/chat, /api/chat, /mcp surface.
日本語の概要は準備中です。原文の説明を表示しています。
Use this skill when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, speech generation (TTS), voice design, voice replication, streaming responses, background research tasks, function calling, structured output, or migrating from the old generateContent API. Covers SDK usage and best practices for Gemini models and agents in Python and TypeScript.
日本語の概要は準備中です。原文の説明を表示しています。
Use this skill when OpenStoryline is already installed and the user wants to start the local MCP/Web services, create or continue a session, send editing instructions, perform multi-turn re-editing, and verify rendered video outputs, as well as Chinese requests like “启动 OpenStoryline”, “把 OpenStoryline 跑起来”, “用 OpenStoryline 剪视频”.
日本語の概要は準備中です。原文の説明を表示しています。
Conducts multi-turn iterative deep research on specific topics within a codebase with zero tolerance for shallow analysis. Use when the user wants an in-depth investigation, needs to understand how something works across multiple files, or asks for comprehensive analysis of a specific system or pattern.
日本語の概要は準備中です。原文の説明を表示しています。
Use this skill when a user wants to create, run, or analyze evaluation suites for Microsoft 365 Copilot declarative agents with the public @microsoft/m365-copilot-eval CLI. Trigger on intents such as "evaluate my agent", "test my agent", "run my evals", "create eval prompts", "add multi-turn tests", "tune evaluator thresholds", "why is my agent failing", or "set up eval environment variables".
日本語の概要は準備中です。原文の説明を表示しています。
Train and run models on Tinker (Thinking Machines) — LoRA and full fine-tuning, SFT with correct loss masking, RL/RFT with group-relative advantages and importance sampling, agentic RL over multi-turn tool-using agents, inference via the SamplingClient or the OpenAI-compatible endpoint, checkpoint management and export, every hyperparameter, and cost estimation from the live price table. Use this skill whenever the user mentions Tinker, tinker://, TINKER_API_KEY, console.tinker.ai, or the tinker CLI; wants to fine-tune, SFT, RFT, or RL a model on Tinker; asks about Tinker models, context windows, LoRA rank, sampling params, loss functions (cross_entropy, importance_sampling, ppo, cispo, dro), Datum construction, forward_backward or optim_step; wants to sample from, download, resume, merge, or export a Tinker checkpoint; or asks what a Tinker run will cost. Whenever Tinker work is initialized in a project, this skill also creates a TINKER_PRICING.md there so prices sit next to the training code. Reach for it even on vague asks like "train a model on this data" or "get inference working" when Tinker is the platform in play.
日本語の概要は準備中です。原文の説明を表示しています。
Generate structured research questions, testable hypotheses, and candidate empirical strategies from a topic, phenomenon, or dataset description. Use when user says "give me research ideas on X", "brainstorm questions about Y", "what could I study with this data?", "I'm looking for a paper idea on...", "generate hypotheses for...". One-shot generation, not multi-turn. For idea-refinement use `/interview-me`.
日本語の概要は準備中です。原文の説明を表示しています。
Interactive interview that formalizes a fuzzy research idea into a structured spec (RQ, hypotheses, identification, data needs, empirical strategy). Use when user says "interview me", "help me think through this idea", "I have a half-baked idea", "formalize this into a project", "walk me through framing a study". Multi-turn Q&A; saves spec to disk. NOT for lit review (`/lit-review`) or ideation from scratch (`/research-ideation`).
日本語の概要は準備中です。原文の説明を表示しています。
Delegate tasks to ANY CLI agent (claude, codex, aider, ...) running in a detached tmux session, with a race-safe done-signal protocol and multi-turn iteration. Use when delegating work to a non-Claude CLI agent, when the user says "tmux delegate", "run agent in tmux", "delegate to codex/aider", or when executor work should run in an observable background terminal instead of the Agent tool.
日本語の概要は準備中です。原文の説明を表示しています。
Edit, trim, cut, caption, subtitle, reframe, combine, add background music to, or export video, audio, and image files through Cassette. Use this skill whenever the user asks to change, preview, or render a media file in the project — even if they never say "Cassette" or name a tool — for example "trim the intro off demo.mp4", "add subtitles to this clip", "make me a 30-second cut with music", "why is there dead air at the start". Drives the local Oh My Cassette stdio MCP tools in Codex or Claude as one multi-turn conversation with the Cassette agent, with per-turn timeline previews, guided questions, and rendering only on explicit export.
日本語の概要は準備中です。原文の説明を表示しています。
Conversational search runtime: send messages, keep sessions consistent, and verify retrieval behavior and responses.
日本語の概要は準備中です。原文の説明を表示しています。
Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js). Use when the user asks to "write tests for my agent", "add a test for this tool", "test the handoff", "pin this bug", "why does my agent test fail", or after building or changing agent behavior that needs regression coverage. Covers the SDK's test session harness, assertions on messages, tool calls and handoffs, LLM judging of a reply against an intent, mocking tools, multi-turn tests, and judging whole conversations with the built-in judges. For interactive poking use debugging-livekit-agents. For grading whole conversations at scale use running-livekit-simulations.
日本語の概要は準備中です。原文の説明を表示しています。
Drives a multi-turn conversation with a LiveKit agent running locally to see what it does. Use when the user says "test my agent", "try my agent", "does this work", "why did it call that tool", "it says the wrong thing when I ask X", "test this change", or whenever you have edited an agent and need to check how it behaves. Wraps `lk agent debugger`: start the agent in text mode, send turns, read the tool calls, handoffs, errors and logs behind each reply, and restart after an edit. It runs without audio or a LiveKit room at one LLM call per turn, so it is the preferred way for a coding agent to live-test during development, and the default when the user says "test" without naming unit tests or simulations.
日本語の概要は準備中です。原文の説明を表示しています。