本文へ移動
cccskills
無料GitHub で公開

image-to-text

Extract text from images using OCR. Use when the user shares a screenshot and you need to read the text content, copy UI labels, or extract copy from a design mockup.

インストール方法を見る

含まれるファイル(5)

  • SKILL.md3.1 KB
  • scripts/image-to-text.js1.2 KB
  • scripts/image-to-text.sh483 B
  • scripts/package-lock.json5.2 KB
  • scripts/package.json133 B

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Image to Text

Extract all readable text from an image using OCR (Tesseract). Returns the full text content along with word-level bounding boxes and confidence scores.

When to Use

  • Reading text content from a screenshot or design mockup
  • Extracting UI copy (labels, buttons, headings) so you don't have to retype it
  • Getting text positions and bounding boxes from a design image

How It Works

  1. The image is passed to Tesseract.js for optical character recognition
  2. Tesseract segments the image into lines and words
  3. Returns the full text plus word-level details (position, confidence)

Usage

bash <skill-path>/scripts/image-to-text.sh <image-path> [language]

Arguments:

  • image-path — Path to the image file (required)
  • language — OCR language code (optional, defaults to eng). Common: eng, fra, deu, spa, chi_sim, jpn

Examples:

# Extract text from a screenshot
bash <skill-path>/scripts/image-to-text.sh ./screenshot.png

# Extract French text
bash <skill-path>/scripts/image-to-text.sh ./mockup.png fra

Output

{
  "text": "Request work\nSuggestions\nPlumbing\nHVAC\nCleaning\nElectrical",
  "confidence": 87.4,
  "words": [
    {
      "text": "Request",
      "confidence": 94.2,
      "bbox": { "x0": 142, "y0": 180, "x1": 268, "y1": 204 }
    },
    {
      "text": "work",
      "confidence": 96.1,
      "bbox": { "x0": 274, "y0": 180, "x1": 332, "y1": 204 }
    }
  ],
  "lines": [
    {
      "text": "Request work",
      "confidence": 95.1,
      "bbox": { "x0": 142, "y0": 180, "x1": 332, "y1": 204 }
    }
  ]
}
FieldTypeDescription
textStringFull extracted text, newline-separated
confidenceNumberOverall confidence score (0-100)
wordsArrayEach word with text, confidence, and bounding box
linesArrayEach line with text, confidence, and bounding box

Present Results to User

After extracting text, present the content grouped by lines:

Extracted text (87.4% confidence):

  Request work
  Suggestions
  Plumbing
  HVAC
  Cleaning
  Electrical

Found 6 lines, 6 words.

Use the extracted text directly when implementing UI copy from a design.

Troubleshooting

Low confidence / garbled text — Tesseract works best with clean, high-contrast text. Screenshots of rendered UI work well. Photos of text at angles or with noise may produce poor results.

Wrong language — Pass the correct language code as the second argument. Tesseract needs the right language model to recognize characters.

First run is slow — Tesseract downloads language data (~4MB for English) on the first run. Subsequent runs are faster.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Multi-model agent orchestration using specialized agents for planning, coding, research, math/science, visual analysis, and adversarial review. Use when tasks are complex enough to benefit from different models' strengths, when you want adversarial review to catch blind spots, or when coordinating multi-step workflows across agent roles. Triggers on complex projects, multi-step tasks, architecture decisions, or when explicitly requested.

日本語の概要は準備中です。原文の説明を表示しています。

pascalorg/skills962026年9月11日 更新

Check color contrast ratios against WCAG AA and AAA accessibility standards. Use when the user wants to verify if their color combinations are accessible, check contrast between text and background colors, or audit a palette for accessibility.

日本語の概要は準備中です。原文の説明を表示しています。

pascalorg/skills962026年9月11日 更新

Audit a .glb or .gltf and make it small, correct, and fast to load in a browser, with a measured before/after report. Use when a model is "too big" or "slow to load", when the user asks to "optimize this GLB", "compress this glTF", or "export for web", or when they mention Draco, meshopt, gltfpack, KTX2, Basis, texture VRAM, draw calls, or a Blender / CAD / photogrammetry / Pascal export headed for Three.js, React Three Fiber (drei useGLTF), Babylon.js, or model-viewer. Covers inspection, spec validation, geometry and texture compression, scene-graph cleanup, and the correctness checks (metres, Y-up, node names, animations, PBR fidelity, alpha modes, vertex colors) that must not regress. Not for editing or modelling geometry, authoring or laying out a scene, generating 3D from text or images, converting from FBX/OBJ/CAD, or debugging a runtime frame rate that has nothing to do with asset size.

日本語の概要は準備中です。原文の説明を表示しています。

pascalorg/skills962026年9月11日 更新

Extract color palettes from images (screenshots, Figma exports, design mockups) to help implement matching UI. Use when the user shares a screenshot, design image, or asks to "match these colors", "extract colors from this image", "implement this design", or "get the color palette".

日本語の概要は準備中です。原文の説明を表示しています。

pascalorg/skills962026年9月11日 更新

Compare two images pixel-by-pixel and get a visual diff. Use when the user wants to compare their implementation against a design, spot differences between two screenshots, or verify visual regression.

日本語の概要は準備中です。原文の説明を表示しています。

pascalorg/skills962026年9月11日 更新

Web design reference for building production-grade interfaces. Covers layout, typography, color, spacing, shadows, animation, accessibility, responsive design, components, performance, and UX psychology. Use when building UI, reviewing design quality, choosing design tokens, or making any visual design decision.

日本語の概要は準備中です。原文の説明を表示しています。

pascalorg/skills962026年9月11日 更新

pascalorg のスキルをすべて見る

このスキルの問題を報告する