本文へ移動
cccskills
無料GitHub で公開

arkcli-understand

arkcli +understand:基于 Responses API 的 12 个多模态专项理解配方,支持临时 API Key/Base URL/Endpoint 执行与无副作用 dry-run。用户只给 Endpoint 时先用 resources resolve;转写、抽取、字幕、定位等明确产出走本 skill,开放式对话走 +chat,生成走 +gen。

インストール方法を見る

含まれるファイル(4)

  • SKILL.md9.6 KB
  • references/arkcli-understand.md15.7 KB
  • references/evals.md4.6 KB
  • references/sub-skills.md7.3 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

arkcli +understand

CRITICAL — 开始前 MUST 先用 Read 工具读取 ../arkcli-shared/SKILL.md,其中包含认证闸门、API Key 错误恢复、配置排查与命令选择顺序。 CRITICAL — 执行 +understand 之前,MUST 先用 Read 工具读取 references/arkcli-understand.md(命令/flag/返回值/错误)与 references/sub-skills.md(12 个 sub-skill 的用途与期望输出形态)。禁止凭印象拼命令。

核心概念

  • +understand 不是 12 套实现,而是 1 个引擎 + 一层语义:底层引擎就是 +chat 用的数据面 Responses API;每个 sub-skill 只是一条配方 {模态, fallback 模型, 内置 system prompt}。Platform Profile 省略 --model 时必须使用当前 Profile 的 Resources.Text.Default Endpoint;只有 Plan 类 Profile 才使用 recipe fallback 模型。
  • 命令形态:arkcli +understand <sub-skill> --input @file [prompt]。
    • args[0] 命中注册表(12 个之一)→ 当作显式 sub-skill,其余位置参数当 prompt 叠加在内置 prompt 之上。
    • args[0] 不命中 → 整段位置参数都当 prompt,服务端按首个 --input 的文件模态自动路由到该模态的默认 sub-skill。
  • 必须有 --input:sub-skill 是「对某个文件做理解」,没有 --input 会直接 missing_input 报错。纯 prompt 无法推导配方。
  • 返回值与 +chat 完全一致:arkcli 扁平 schema {id, model, content, reasoning_content, usage},不是 Responses 原生 output[].content[].text 嵌套。详见 references/arkcli-understand.md 的「返回值」段。
  • 多模态上传:image/video/doc 由 SDK file:// preprocessor 自动走 Files API;audio 是特例——内联为 base64 data URL,上限 25MB。详见 reference。
  • 用户临时提供 --api-key / --base-url / Endpoint 时,MUST 读取 ../arkcli-shared/references/execution-context.md。不要根据 Key 文本或 Endpoint 名称猜数据面/工作流。
  • --dry-run 只在本地解析配方、当前 Profile 已持久化的默认 Endpoint、显式模型和输入引用,输出统一 preview.v1;不读取在线 Endpoint/模型元数据、不上传文件、不调用 Responses API、不产生模型用量、不创建 response id 或持久化响应。在线才能补齐的值必须标为 unresolved,不能把预览当服务端 validation。

快速决策(understand vs chat vs gen)

用户意图走哪个
有明确产出形态的多模态理解任务(转写 / 翻译 / 字幕 / 框选定位 / GUI 操作 / 字段抽取 / 分章节总结 / 多说话人 / 会议纪要)+understand(命中某个 sub-skill)
开放式带图/视频对话、追问、推理、需要多轮接续(--store/--previous-response-id)、需要 tools(web_search/function)或 --text-format json_schema 严格 JSON../arkcli-chat/SKILL.md
生成图片 / 视频(不是理解已有素材)../arkcli-gen/SKILL.md

一句话判据:任务能映射到下面某个 sub-skill 名 → 用 +understand;否则开放对话用 +chat,生成用 +gen。

Agent 快速执行顺序

  1. 判断用户的理解任务命中哪个 sub-skill(见下表)。命中就显式带上 sub-skill 名,不要让它走自动路由去猜。
  2. 用户只给 ep-... → arkcli resources resolve <ep-id> --format json;只有用户任务属于明确理解产出,且 supported_workflows 包含 understand,才走本 skill。
  3. 用户给了临时 Key/Base URL/Endpoint → 按共享 execution-context 组合规则决定参数,不先切 profile。
  4. 没有完整 stateless 上下文时过认证闸门:不确定登录态先 arkcli auth status;鉴权失败按 ../arkcli-auth/references/auth-modes.md 的「API Key 模式的错误恢复」处理,不要原地重试。
  5. 备好输入:--input @<file>(可多次)。本地文件用 @ 前缀;远程用 https:// / tos:// URL。
  6. 默认不传 --model:Platform Profile 会使用 Resources.Text.Default 中已部署的 ep-xxx,与 +chat 保持一致;Plan 类 Profile 才使用 sub-skill 的 recipe fallback 模型。只有用户明确指定其他资源时才传完整版本化 ID 或 ep-xxx。
  7. 需要逐段输出加 --stream;想微调任务指令用 --system-prompt-append "...",或 --system-prompt-override "..."。
  8. 跑完回到用户原始目标,不要停在中间产物上。

sub-skill 速查表(12 个 / 4 模态)

下表是 Plan 类 Profile 使用的 recipe fallback 模型。Platform Profile 不使用这些裸模型名,而是使用当前 Profile 的 Resources.Text.Default Endpoint。确需覆盖时见上文执行顺序。

sub-skill模态Plan recipe fallback 模型一句话用途
image-captionimagedoubao-seed-1-6单图/多图描述、OCR(保留原文,不意译)
image-groundingimagedoubao-seed-1-6视觉定位:输出目标 bbox (x1,y1,x2,y2) + confidence
image-guiimagedoubao-seed-1-6GUI 截图 → 可执行操作序列(12 类操作 JSON 数组)
doc-extractfiledoubao-seed-1-6PDF/文档按 schema 抽取结构化 JSON(缺失填 null,保留 page)
video-summaryvideodoubao-seed-1-6视频总结:overall / segment / chapter + 关键时间点
video-qavideodoubao-seed-1-6视频问答:结合画面+音频+时间轴精准回答
vauvideodoubao-seed-2-0-lite音视频联合理解:分析视频内声音元素,出音频分析报告
asraudiodoubao-seed-2-0-lite语音转写:仅输出纯文本,无任何前后缀/格式
asr-alignaudiodoubao-seed-2-0-lite字幕打轴:默认 SRT 格式(序号+起止时间+文本)
asr-speakersaudiodoubao-seed-2-0-lite多说话人转写:标 [spk0] / [spk1] ...
astaudiodoubao-seed-2-0-lite语音翻译:把音频里的话译成文本
meeting-minutesaudiodoubao-seed-2-0-lite会议纪要:结构化 markdown(主题/参会人/要点/决议/待跟进)

逐条的期望输出形态与示例命令见 references/sub-skills.md。

省略 sub-skill 时的自动路由

args[0] 不命中注册表时,按首个 --input 的文件扩展名推断模态并兜底到该模态的默认 sub-skill:

推断模态兜底 sub-skill
image(.jpg/.png/.webp/...)image-caption
video(.mp4/.mov/...)video-qa
audio(.mp3/.wav/.m4a/...)asr
file(其他扩展名)doc-extract
  • 没有 --input 时无法自动路由(纯 prompt 推不出配方)→ 报 missing_sub_skill。
  • 目前只按文件模态兜底,不做 prompt 关键词级路由(例如说「框出」不会自动选 image-grounding)。Agent 想要非默认 sub-skill(如 grounding / gui / summary / 字幕 / 翻译)就显式写出 sub-skill 名,别依赖自动路由。

命令一览

命令说明
arkcli +understand image-caption --input @photo.jpg "描述这张图"图片描述
arkcli +understand image-caption --input @doc.png "OCR,原样保留文字"OCR
arkcli +understand image-grounding --input @scene.jpg "框出所有红色car"视觉定位 bbox
arkcli +understand doc-extract --input @invoice.pdf "抽取 发票号/金额/日期"文档字段抽取
arkcli +understand video-summary --input @clip.mp4 "按 chapter 总结"视频分章节总结
arkcli +understand video-qa --input @clip.mp4 "视频里出现了几辆车?"视频问答
arkcli +understand asr --input @speech.mp3语音转写(无需 prompt)
arkcli +understand asr-align --input @speech.mp3生成 SRT 字幕
arkcli +understand asr-speakers --input @meeting.wav多说话人转写
arkcli +understand ast --input @speech.m4a "译成英文"语音翻译
arkcli +understand meeting-minutes --input @meeting.m4a会议纪要
arkcli +understand "<prompt>" --input @photo.jpg省略 sub-skill → 按模态自动路由(此处 → image-caption)
arkcli +understand asr --input @speech.mp3 --stream流式逐段输出
arkcli +understand image-caption --input @x.jpg --system-prompt-append "用英文回答"在内置 prompt 后追加指令
arkcli +understand image-caption --input @x.jpg "描述图片" --dry-run --format json返回 recipe / 模型 / 请求摘要的 preview.v1;零网络、不上传、不推理

详细文档

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

arkcli agent:管理 ARK Managed Agents,包括 Agent / Skill / Env / Session / File / Memory Store / Vault / MCP OAuth。控制面优先走 ForTop/OpenTOP,Session 运行时和 Files 走数据面直联。

日本語の概要は準備中です。原文の説明を表示しています。

volcengine/ark-cli1422026年10月9日 更新

Inspect or invoke locally registered ArkCLI actions when product commands cannot cover a task. Use for registry errors or exact raw payloads. Not for public API catalogs or OpenAPI schemas.

日本語の概要は準備中です。原文の説明を表示しています。

volcengine/ark-cli1422026年10月9日 更新

arkcli 认证管理:交互式登录、Volc SSO 登录、查看状态、退出登录、生成 ARK API Key (apikey)、以及云开发机/CI 用 `arkcli init-volc` 从 VOLC_INIT_* 环境变量无交互引导 platform profile。0.1.16 起 SSO 登录走 Gate 1+2 自动绑定 Profile 切面 (type/region/project/owner_trn);AK/SK login 通道暂关。当用户需要初始化凭证、排查鉴权问题、切换认证方式、生成或重选 ARK API Key、或在已注入凭证的环境无交互引导时使用。反触发:用户问 TTS/ASR/语音模型能力、接入或调用时,不要引导 `auth apikey`,只转 models search 说明 arkcli 当前仅支持广场发现。

日本語の概要は準備中です。原文の説明を表示しています。

volcengine/ark-cli1422026年10月9日 更新

查询火山引擎 ARK 拆分账单明细(结算金额、Token 用量计费),支持按账期月、月范围、Endpoint、API Key、产品编码等维度过滤。当用户问账单、花了多少钱、对账、账期、按 EP / API Key 拆账、按产品拆账、月度账单、出账明细时使用。注意 billing 跟 usage stats 不同:stats 出推理量(近实时),billing 出结算金额(T+1 出账,财务口径)。

日本語の概要は準備中です。原文の説明を表示しています。

volcengine/ark-cli1422026年10月9日 更新

arkcli +chat:通过数据面 Responses API 快速对话/推理,支持多模态、流式、多轮、临时 API Key/Base URL/Endpoint 执行与无副作用 dry-run。当用户给出 Endpoint 但未说明工作流时,先用 resources resolve 识别候选;已经出现 Responses API capability/access 错误时,只读用 models get 核对精确模型的 api_support,不重试真实调用。有明确产出形态的多模态理解走 arkcli-understand。

日本語の概要は準備中です。原文の説明を表示しています。

volcengine/ark-cli1422026年10月9日 更新

arkcli +code-example:为指定基础模型生成多语言(Python / Go / Java / Node / curl)调用示例代码并写入本地文件。数据源是火山方舟 OpenTOP OpenGetSampleCode。当用户需要拿某个基础模型的 SDK / curl 调用示例、保存为本地接入模板时使用。反触发:TTS/ASR/语音模型没有 arkcli 示例代码路径,不能靠补版本解决,只能转 models search 说明当前不支持。

日本語の概要は準備中です。原文の説明を表示しています。

volcengine/ark-cli1422026年10月9日 更新

volcengine のスキルをすべて見る

このスキルの問題を報告する