本文へ移動
cccskills
無料GitHub で公開

pdf-reader

Read and extract content from PDF files — text, tables, metadata, and images. Use when asked to read a PDF, extract text from a PDF, summarize a PDF, analyze a PDF document, get tables from a PDF, or check PDF metadata. Also triggers on "open this PDF", "what does this PDF say", "parse PDF", "PDF to text", or when a .pdf file path or URL is provided.

インストール方法を見る

含まれるファイル(2)

  • SKILL.md6.2 KB
  • scripts/extract.py11.1 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

PDF Reader

Extract content from PDF files using pdftotext (Poppler) for text and pdfplumber (Python) for tables and structured extraction.

Quick Reference

TaskToolCommand
Full textpdftotextpdftotext file.pdf -
Text with layoutpdftotextpdftotext -layout file.pdf -
Specific pagespdftotextpdftotext -f 3 -l 5 file.pdf -
Tablespdfplumberpython3 scripts/extract.py tables file.pdf
Metadatapdfinfopdfinfo file.pdf
Page countpdfinfopdfinfo file.pdf | grep Pages
Imagespdfimagespdfimages -list file.pdf
Fontspdffontspdffonts file.pdf
OCR (scanned PDF)tesseractpython3 scripts/extract.py ocr file.pdf
Smart text + OCRextract.pypython3 scripts/extract.py text file.pdf
Quick surveyextract.pypython3 scripts/extract.py scan file.pdf

Workflow

Step 1: Get the PDF

If the user provides a URL, download it first:

curl -sL "URL" -o /tmp/document.pdf

Verify it's a valid PDF:

file /tmp/document.pdf  # should say "PDF document"
pdfinfo /tmp/document.pdf  # metadata + page count

Step 2: Choose Extraction Method

Plain text (most cases):

pdftotext file.pdf -

This pipes output to stdout. For large PDFs, use page ranges:

pdftotext -f 1 -l 10 file.pdf -    # pages 1-10

Layout-preserving text (columns, formatted docs):

pdftotext -layout file.pdf -

Use -layout when the PDF has multi-column layouts, tables rendered as text, or precise spacing that matters.

Tables (structured data):

python3 scripts/extract.py tables file.pdf

Or inline with pdfplumber:

import pdfplumber

pdf = pdfplumber.open("file.pdf")
for i, page in enumerate(pdf.pages):
    tables = page.extract_tables()
    for table in tables:
        print(f"\n--- Table on page {i+1} ---")
        for row in table:
            print(" | ".join(str(cell or "") for cell in row))
pdf.close()

Metadata only:

pdfinfo file.pdf

Returns: title, author, creator, producer, page count, page size, dates.

Step 3: Handle Large PDFs

For PDFs over ~50 pages, don't dump everything at once:

  1. Get page count: pdfinfo file.pdf | grep Pages
  2. Extract in chunks: pdftotext -f 1 -l 20 file.pdf -
  3. Process chunk, then continue: pdftotext -f 21 -l 40 file.pdf -

For targeted extraction (searching for specific content):

# Extract all text, grep for relevant sections
pdftotext file.pdf - | grep -n -i "keyword"

# Then extract the specific page range
pdftotext -f PAGE -l PAGE file.pdf -

Step 4: Handle Scanned PDFs (OCR)

If pdftotext returns empty or garbled output, the PDF is likely scanned.

Detection:

python3 scripts/extract.py scan file.pdf   # reports scanned pages
pdffonts file.pdf                           # empty = image-based

Smart extraction (auto-fallback):

text mode automatically detects scanned pages and OCRs them:

python3 scripts/extract.py text file.pdf

Pages with selectable text extract normally. Pages without selectable text fall back to OCR via Tesseract. No manual detection needed.

Force OCR on all pages:

python3 scripts/extract.py ocr file.pdf
python3 scripts/extract.py ocr file.pdf --pages 1-5
python3 scripts/extract.py ocr file.pdf --dpi 400        # higher quality
python3 scripts/extract.py ocr file.pdf --lang eng+nor   # multi-language

OCR options:

  • --dpi 300 — resolution for page-to-image conversion (default: 300, higher = slower but better)
  • --lang eng — Tesseract language pack (default: eng). Use + for multiple: eng+nor+deu
  • --pages 1-5 — limit to specific pages (recommended for large PDFs)

Available language packs:

tesseract --list-langs

Install additional languages via Homebrew:

brew install tesseract-lang    # all languages

Step 5: Handle Other Edge Cases

Mixed PDFs (some pages scanned, some not):

Just use text mode — it handles mixed PDFs automatically:

python3 scripts/extract.py text file.pdf

Selectable pages extract instantly, scanned pages get OCR'd. The output is tagged so you know which pages used OCR.

Password-protected PDFs:

pdftotext -upw "password" file.pdf -   # user password
pdftotext -opw "password" file.pdf -   # owner password

Encoding issues (garbled output):

pdftotext -enc UTF-8 file.pdf -

Extract images:

pdfimages -png file.pdf /tmp/images/img   # extracts as PNG
pdfimages -list file.pdf                  # list images without extracting

Decision Tree

Is it a URL? → curl -sL "URL" -o /tmp/doc.pdf
            ↓
Run: python3 scripts/extract.py scan file.pdf
            ↓
All pages have selectable text?
  YES → pdftotext file.pdf -           (fast, simple)
  NO  → python3 scripts/extract.py text file.pdf  (auto OCR fallback)
            ↓
Need tables?
  YES → python3 scripts/extract.py tables file.pdf

Tips

  • Start with scan on unknown PDFs — it reports pages, tables, scanned detection, and a preview
  • pdftotext is fastest for normal PDFs — try it first
  • Use -layout for multi-column documents (academic papers, reports)
  • pdfplumber is better for tables — it understands cell boundaries
  • text mode auto-detects scanned pages and OCRs only those — preferred over raw pdftotext for unknown PDFs
  • ocr mode is for forcing OCR on everything (useful when text extraction gives garbled output despite appearing selectable)
  • Higher --dpi gives better OCR accuracy but is slower (300 is a good default, 400+ for small text)
  • For PDFs from URLs, always download to /tmp/ first — don't pipe curl to tools
  • Large PDF text output may exceed context limits — use --pages to extract in ranges

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Audit, generate, update, and lint AGENTS.md files across all projects. Use when asked to check project context files, scaffold AGENTS.md for new projects, update stale ones, or run a cross-project audit.

日本語の概要は準備中です。原文の説明を表示しています。

espennilsen/pi1222026年9月22日 更新

blog-post

無料

Draft, edit, and publish blog posts for e9n.dev. Use when creating new posts, editing drafts, or refining existing content. Handles Eleventy frontmatter, Tailwind formatting, and Espen's authentic voice.

日本語の概要は準備中です。原文の説明を表示しています。

espennilsen/pi1222026年9月22日 更新

Generate a full operational status report for the Aivena bot. Checks all subsystems: extensions, webserver, Telegram, chat bridge, heartbeat, cron, database, memory, CRM, calendar, task management, jobs/telemetry, and storage. **Triggers — use this skill when:** - User asks for "status", "bot status", "system status", "operational status" - User asks "is everything running?", "how's Aivena doing?" - User says "health check", "diagnostics", "systems check" - User asks "what's the state of the bot?"

日本語の概要は準備中です。原文の説明を表示しています。

espennilsen/pi1222026年9月22日 更新

Parse git history and produce or update a CHANGELOG.md following the Keep a Changelog convention. Supports Conventional Commits, basic prefix conventions, and unstructured commit messages. Intelligently categorizes changes, detects breaking changes, links to PRs/issues, and handles both initial generation and incremental updates. **Triggers — use this skill when:** - User asks to "generate", "create", "update", or "write" a changelog - User mentions "CHANGELOG", "changelog", "release notes" - User says "document changes", "what changed since last release" - User wants to "prepare a release" and needs a changelog entry - User asks to "clean up" or "reformat" an existing changelog **Covers:** Any git-based project. Handles Conventional Commits (feat/fix/chore), Angular convention, basic prefixes (Add/Fix/Remove), and freeform commit messages. Outputs Keep a Changelog format with optional Common Changelog enhancements.

日本語の概要は準備中です。原文の説明を表示しています。

espennilsen/pi1222026年9月22日 更新

cmux

無料

Orchestrate cmux terminal panes — split terminals, run parallel processes, read output from other panes, and use the built-in browser. Use when working inside cmux and you need to run a dev server, watch tests, spawn sub-agents, or preview web pages.

日本語の概要は準備中です。原文の説明を表示しています。

espennilsen/pi1222026年9月22日 更新

Review code for quality, security, performance, and maintainability. Use when reviewing PRs, auditing a codebase, or refactoring. Covers TypeScript, Node.js, infrastructure-as-code, and full-stack web apps.

日本語の概要は準備中です。原文の説明を表示しています。

espennilsen/pi1222026年9月22日 更新

espennilsen のスキルをすべて見る

このスキルの問題を報告する