本文へ移動
cccskills
無料GitHub で公開

byted-tos-image-process

Transforms and inspects image objects stored in Volcengine TOS. Use this skill only when the task explicitly involves a TOS bucket/object key, TOS image processing, TOS-to-TOS save-as output, or Volcengine TOS image process syntax such as image/info, image/resize, image/format, image/watermark, image/draw, image/blindwatermark, or image/understanding. Do not use this skill for ordinary uploaded screenshots, local images, UI screenshot analysis, generic OCR, face detection, or visual question answering unless the user clearly says the image is a TOS object or asks to process/save it through TOS.

インストール方法を見る

含まれるファイル(15)

  • SKILL.md8.0 KB
  • LICENSE11.1 KB
  • README.md8.1 KB
  • REFERENCE.md11.8 KB
  • requirements.txt4 B
  • scripts/image_blindwatermark.py10.2 KB
  • scripts/image_draw.py8.9 KB
  • scripts/image_format.py9.5 KB
  • scripts/image_info.py12.4 KB
  • scripts/image_process.py7.4 KB
  • scripts/image_resize.py9.1 KB
  • scripts/image_understanding.py7.8 KB
  • scripts/image_watermark.py15.6 KB
  • scripts/image_zoom.py11.4 KB
  • WORKFLOWS.md12.7 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Volcengine TOS Image Process

Inspect and transform images stored in Volcengine TOS — metadata, format conversion, resize, watermark, blind watermark, and AI-powered image understanding.

Setup (once per environment)

Install dependencies on first use:

cd {baseDir}
pip install -r {baseDir}/requirements.txt

Then run scripts with Python 3.7+:

python3 {baseDir}/scripts/<script>.py <args>

If you see a ModuleNotFoundError for tos, reinstall dependencies.

Environment Variables

This skill relies on the TOS identity declared in the metadata block. Common runtime variables are:

Environment VariableRequiredDescription
TOS_ACCESS_KEYYesTOS access key ID
TOS_SECRET_KEYYesTOS secret access key
TOS_ENDPOINTYesTOS endpoint URL
TOS_REGIONYesTOS region
TOS_BUCKETYesSource bucket that stores the image
TOS_OBJECT_KEYNoSource object key of the image. Can be overridden with --key
TOS_SECURITY_TOKENNoSTS session token when using temporary credentials
TOS_SAVEAS_BUCKETNoDefault target bucket for saving processed results
TOS_SAVEAS_OBJECT_PREFIXNoDefault key prefix for saving processed results

Quick start (common tasks)

# Read image metadata
python3 {baseDir}/scripts/image_info.py --key photo.jpg

# Convert to WebP
python3 {baseDir}/scripts/image_format.py --key photo.jpg --f webp --output converted.webp

# Resize to width 500
python3 {baseDir}/scripts/image_resize.py --key photo.jpg --width 500 --output resized.jpg

# Draw points and connecting lines
python3 {baseDir}/scripts/image_draw.py --key photo.jpg \
  --points 50x50-200x120-320x220 --line --color FF0000 --output draw.jpg

# Zoom by resize + crop
python3 {baseDir}/scripts/image_zoom.py --key photo.jpg \
  --resize-w 1200 --crop-w 500 --crop-h 400 --gravity center --output zoom.jpg

# Add visible text watermark
python3 {baseDir}/scripts/image_watermark.py --key photo.jpg \
  --text "My Brand" --font fangzhengshusong --color FF0000 --size 72 \
  --gravity center --output watermarked.jpg

# Embed blind watermark (requires ≥512×512 image and account permission)
python3 {baseDir}/scripts/image_blindwatermark.py --key photo.jpg \
  --kv text=HelloBlind --output blind.jpg

# Run a custom process string
python3 {baseDir}/scripts/image_process.py --key photo.jpg \
  --process "image/resize,w_300,h_300,m_fill" --output filled.jpg

# Preview the resolved request without calling TOS
python3 {baseDir}/scripts/image_resize.py --key photo.jpg \
  --width 500 --dry-run --json

# AI-powered understanding for a TOS image object (requires whitelist)
python3 {baseDir}/scripts/image_understanding.py --key photo.jpg \
  --prompt "Describe this image in detail"
python3 {baseDir}/scripts/image_understanding.py --key document.png \
  --prompt "识别图片中的所有文字内容"

Available scripts

ScriptPurpose
scripts/image_info.pyRead image metadata (format, dimensions, size). Falls back to local parsing when TOS returns raw bytes.
scripts/image_format.pyConvert format (jpg, png, webp) with optional quality setting.
scripts/image_resize.pyResize by width/height/mode.
scripts/image_draw.pyDraw points and optional connecting lines on an image with image/draw.
scripts/image_zoom.pyBuild agent-friendly zoom results by chaining image/resize and crop.
scripts/image_watermark.pyAdd visible text or image watermark with positioning, rotation, tiling, and opacity.
scripts/image_blindwatermark.pyEmbed blind watermark. Requires account-level permission and image ≥512×512 px.
scripts/image_process.pyPass any raw image/... process string.
scripts/image_understanding.pyAI-powered understanding for TOS image objects via VLM (doubao-seed-1.6-vision). Supports description, OCR, face detection, and visual Q&A only when the source image is a TOS object. Requires account whitelist.

All scripts support --key to override TOS_OBJECT_KEY, --output for local save, and --saveas-bucket/--saveas-object for TOS-to-TOS persistence. TOS_SAVEAS_BUCKET and TOS_SAVEAS_OBJECT_PREFIX are used as defaults when save-as CLI arguments are omitted. Most scripts support --json for machine-readable output, and all process-building scripts support --dry-run to preview the resolved request without calling TOS. Run any script with -h for full usage.

Out of scope

  • Editing images with local desktop tooling outside TOS.
  • Ordinary uploaded screenshots, local image files, mobile UI screenshots, generic OCR, face detection, and visual question answering that do not involve a TOS bucket/object key. Use the model's native vision or local file tools instead.
  • Video or document processing (use byted-tos-video-process or byted-tos-doc-process).
  • Non-TOS storage providers.

Rules

  • Authentication: Authentication is provided by the TOS identity declared in the metadata block above. Object selection can be overridden per script with --key.
  • Credential safety: Never print credential environment variable values such as TOS_ACCESS_KEY, TOS_SECRET_KEY, TOS_SECURITY_TOKEN, or model API keys. Validate behavior by running the scripts directly instead of echoing or dumping the environment.
  • Trigger boundary: Use this skill only for Volcengine TOS image objects or TOS image process workflows. If the user attaches or references a normal local image/screenshot and does not mention TOS, do not invoke these scripts; answer with native vision/local-file capabilities instead.
  • Dry run: Use --dry-run --json to validate parameters and inspect the generated process string. Dry-run does not call TOS and does not require AK/SK. If bucket/key are not provided, dry-run uses explicit placeholders where possible; real execution still requires valid TOS credentials, bucket/key, and network access.
  • Parameter source of truth: The exact process string syntax is defined by official Volcengine TOS documentation. When uncertain, check REFERENCE.md.
  • Watermark encoding: Text and font parameters in image/watermark require URL-safe Base64 encoding. The watermark script handles this automatically when you pass --text and --font.
  • Blind watermark constraints: The source image must be at least 512×512 pixels, and the account must have the blind watermark capability enabled. If the capability is missing, the script exits with [SKIP] (use --strict to fail hard).
  • Image understanding: Uses image/understanding with the doubao-seed-1.6-vision VLM model for TOS image objects only. The --prompt parameter is required. Requires account whitelist. Response time is typically 10-60 seconds.
  • Language: Reply in the user's preferred language.

Further reading

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Deploy, update, validate, and troubleshoot the AgentKit hybrid-cloud customer-service demo in this directory. Use when a user asks an AI coding agent to follow the README, deploy or update the demo, configure OpenAPI/Runtime/Knowledge/Memory/Sandbox/MCP/Skills/A2A, run customer-view validation, or record manual platform steps and failures.

日本語の概要は準備中です。原文の説明を表示しています。

bytedance/agentkit-samples4702026年10月9日 更新

Generate deterministic SVG algorithmic artwork. Invoke when the user asks for geometric, generative, or algorithmic visual art.

日本語の概要は準備中です。原文の説明を表示しています。

bytedance/agentkit-samples4702026年10月9日 更新

通过本地 Python CLI 和 OpenAPI 客户端管理、排查火山云手机资源。适用于查询实例和资源、截图、执行命令、查看任务、检查应用、主机和机房容量、标签、DNS、路由,以及操作已授权的测试云手机实例。

日本語の概要は準備中です。原文の説明を表示しています。

bytedance/agentkit-samples4702026年10月9日 更新

Launch and continue the Volcengine AI Research survey workflow for concept testing, audience design, questionnaire drafting, interview guide generation, execution confirmation, progress checks, and result queries. Use this skill when the user wants to create, revise, confirm, execute, or follow up on a real AI research survey task in ABCompass instead of doing generic brainstorming, copywriting, translation, summarization, or broad market discussion.

日本語の概要は準備中です。原文の説明を表示しています。

bytedance/agentkit-samples4702026年10月9日 更新

Create and check long-running video material evaluation tasks. Use this skill when the user wants to submit videos for evaluation, check an existing video evaluation task list, or fetch the result of a previously created video evaluation task. This skill is not a general-purpose video upload skill because upload is allowed only as an internal step of task creation. Authentication uses an API key passed as an Authorization bearer token.

日本語の概要は準備中です。原文の説明を表示しています。

bytedance/agentkit-samples4702026年10月9日 更新

火山引擎 AntiDDoSPro 高防域名只读巡检。用户要查域名健康、攻击、流量、CC/WAF/区域封禁、智能防护或 CCAI 状态时使用。

日本語の概要は準備中です。原文の説明を表示しています。

bytedance/agentkit-samples4702026年10月9日 更新

bytedance のスキルをすべて見る

このスキルの問題を報告する