本文へ移動
cccskills
無料GitHub で公開

robust-pdf-read

Reliably extract text from PDFs using pdftotext when standard file reading fails.

インストール方法を見る

含まれるファイル(2)

  • SKILL.md2.4 KB
  • .skill_id29 B

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Robust PDF Text Extraction

Problem

Standard file reading tools (e.g., read_file) often fail to extract text from PDF documents. Instead of returning parsed text, they may return:

  • Raw binary data
  • Base64 encoded images
  • Garbled characters or null bytes

This occurs because PDFs are complex binary formats, not plain text files. Attempts to parse them using general-purpose Python libraries (like PyMuPDF) in sandboxed environments may also fail due to missing dependencies or environment restrictions.

Solution

Use the pdftotext command-line utility (part of poppler-utils) via run_shell. This tool is commonly pre-installed in Linux environments and reliably extracts text content from PDFs.

Procedure

1. Detect Extraction Failure

When attempting to read a PDF:

  • Check the content returned by read_file.
  • If the content contains null bytes (\x00), appears as base64, or is clearly binary/garbled, assume standard reading has failed.

2. Execute pdftotext

Run the following shell command using run_shell:

pdftotext -layout -nopgbrk <file_path> -
  • -layout: Maintains the physical layout of the text (optional but recommended).
  • -nopgbrk: Prevents inserting form feed characters between pages.
  • -: Outputs content to stdout instead of creating a new file.

3. Parse Output

Capture the stdout from the shell command. This string is the extracted text.

Example Usage

Scenario: You need to read document.pdf.

Step 1: Attempt standard read

content = read_file("document.pdf")
if "\x00" in content or not content.strip():
    # Fallback needed
    pass

Step 2: Fallback to shell

result = run_shell("pdftotext -layout -nopgbrk document.pdf -")
text = result.stdout

Prerequisites

  • The environment must have pdftotext installed (usually via poppler-utils).
  • If pdftotext is not found, attempt to install it (apt-get install poppler-utils) if permissions allow, or notify the user.

Benefits

  • Reliability: Bypasses Python library dependency issues in sandboxes.
  • Speed: Command-line tools are often faster than loading heavy Python libraries.
  • Compatibility: Works consistently across most Linux-based agent environments.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Incremental audio production with duration mismatch handling, adaptive stem extension, and pre-mix alignment verification

日本語の概要は準備中です。原文の説明を表示しています。

HKUDS/OpenSpace7,7552026年8月13日 更新

Incremental audio production with duration alignment handling, per-stem verification, and adaptive extension strategies

日本語の概要は準備中です。原文の説明を表示しています。

HKUDS/OpenSpace7,7552026年8月13日 更新

Create serverless API proxy endpoints that hide API keys and provide a unified backend for the dashboard frontend. Designed for Vercel deployment.

日本語の概要は準備中です。原文の説明を表示しています。

HKUDS/OpenSpace7,7552026年8月13日 更新

End-to-end audio production workflow with stems, effects, archiving, and verification

日本語の概要は準備中です。原文の説明を表示しています。

HKUDS/OpenSpace7,7552026年8月13日 更新

Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

日本語の概要は準備中です。原文の説明を表示しています。

HKUDS/OpenSpace7,7552026年8月13日 更新

Fallback pattern for executing Python code when execute_code_sandbox fails

日本語の概要は準備中です。原文の説明を表示しています。

HKUDS/OpenSpace7,7552026年8月13日 更新

HKUDS のスキルをすべて見る

このスキルの問題を報告する