本文へ移動
cccskills
無料GitHub で公開

doc-to-markdown

Converts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing. Fixes pandoc grid tables, simple tables, image paths, CJK bold spacing, attribute noise, and code blocks; for PDFs also strips OCR garbage blocks, repeated headers/footers/watermarks, and absolute image paths from pymupdf4llm output. Benchmarked best-in-class (7.6/10) against Docling, MarkItDown, Pandoc raw, and Mammoth. Trigger on "convert document", "docx to markdown", "parse word", "doc to markdown", "解析word", "转换文档", "HTML to Markdown".

インストール方法を見る

含まれるファイル(22)

  • SKILL.md9.9 KB
  • assets/obsidian-links/fixed.md1.1 KB
  • assets/obsidian-links/legacy.md1.1 KB
  • assets/obsidian-links/source.html1.5 KB
  • assets/reader-pilot-evidence-template.json1.7 KB
  • references/benchmark-2026-03-22.md6.2 KB
  • references/conversion-examples.md6.6 KB
  • references/heavy-mode-guide.md3.8 KB
  • references/html-conversion.md9.5 KB
  • references/obsidian-link-examples.md4.9 KB
  • references/tool-comparison.md3.7 KB
  • scripts/batch_html.py10.5 KB
  • scripts/convert_path.py1.4 KB
  • scripts/convert.py46.5 KB
  • scripts/extract_pdf_images.py7.6 KB
  • scripts/html_to_markdown.py8.8 KB
  • scripts/merge_outputs.py13.2 KB
  • scripts/reader_pilot_gate.py22.7 KB
  • scripts/test_convert.py13.5 KB
  • scripts/test_html_to_markdown.py9.5 KB
  • scripts/test_reader_pilot_gate.py21.6 KB
  • scripts/validate_output.py16.2 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Doc to Markdown

Convert documents to high-quality markdown with intelligent multi-tool orchestration and automatic DOCX post-processing.

Architecture: Pandoc (best-in-class extraction) + 8 post-processing fixes (our value-add).

Quick Start

# DOCX → Markdown (one command, zero manual fixes)
uv run --with pymupdf4llm --with markitdown scripts/convert.py document.docx -o output.md --assets-dir ./media

# PDF → Markdown
uv run --with pymupdf4llm --with markitdown scripts/convert.py document.pdf -o output.md

# Saved HTML → Markdown (Pandoc required; no remote fetching)
uv run scripts/convert.py page.html -o page.md

# Run tests
uv run --with pytest pytest scripts/test_convert.py -v

Dual Mode

ModeSpeedQualityUse Case
Quick (default)FastGoodDrafts, simple documents
HeavySlowerBestFinal documents, complex layouts

Tool Selection

FormatQuick ModeHeavy Mode
PDFpymupdf4llmpymupdf4llm + markitdown
DOCXpandoc + post-processingpandoc + markitdown
PPTXmarkitdownmarkitdown + pandoc
XLSXmarkitdownmarkitdown
HTML/HTMpandoc + source-href retention checkunsupported; use the HTML quick path

Saved HTML And Website Manuals

For HTML/HTM, read references/html-conversion.md before converting. Use scripts/convert.py; the default converts the whole body and does not trim navigation. An explicit --html-selector selects exactly one tag, #id or .class. --html-heading-offset shifts parsed headings for assembly and rejects overflow beyond H6. Relative assets are not downloaded or copied.

Verify source href occurrences against the emitted Markdown AST, then rerun scripts/html_to_markdown.py after cleanup or merging. For manuals, map the book, chapters, lessons and internal headings before assembly; preserve code fences and reconcile rewritten anchors. Conversion success does not certify that figures are readable inside the recipient's actual Markdown reader.

DOCX Post-Processing (automatic)

When converting DOCX via pandoc, 8 cleanups are applied automatically:

ProblemFixTest coverage
Grid tables (+:---+)Single-column → blockquote, multi-column → pipe tableTestPostprocessPipeline
Simple tables ( ---- ----)Multi-column images → pipe table with captionsTestSimpleTable
Image path nesting (media/media/)Flatten to media/, absolute → relativetest_stats_tracking
Pandoc attributes ({width="..."})Removedtest_pandoc_attributes_removed
CJK bold spacing (**粗体**中文)Add space around ** for CJK bold spansTestCjkBoldSpacing (15 cases)
Indented dashed code blocks→ fenced ``` with language detectiontest_code_block_with_language
Escaped brackets (\[...\])→ [...]test_escaped_brackets_fixed
Double-bracket links ([[text]](url))→ [text](url)test_double_bracket_links_fixed

PDF Post-Processing (automatic, 2026-08-30 起)

When converting PDF via pymupdf4llm, 3 cleanups are applied automatically (skip with --no-postprocess):

ProblemFixTest coverage
Tesseract OCR garbage on image regions (<!-- Start of picture text -->...)Block removed; images themselves keptTestStripOcrPictureText
Repeated header/footer/watermark lines (same normalized line on ≥60% of pages, incl. diagonal watermarks)Detected via pymupdf cross-page scan, removed from markdown; bold-wrapped and merged-with-page-number variants also caughtTestRepeatingLines
Absolute image paths (![](/abs/tmp/assets/...))Rewritten relative to the output markdown file (portable output)TestImagePathsRelative

Heavy mode additionally prints a loud ⚠️ HEAVY MODE DEGRADED warning on stderr when one engine fails and the merge would otherwise silently degrade to single-engine output.

Known limits (learned from a 62-page Chinese research-report conversion, 2026-08-30):

  • pymupdf4llm may emit duplicated paragraphs (source text layer has only one copy) — not auto-fixed; spot-check.
  • Dotted TOC pages get detected as tables — rewrite the TOC manually if it matters.
  • Cross-page tables are NOT merged (each page's fragment keeps its own header row) — merge manually.
  • Table cells overlapped by diagonal watermarks can contain watermark character shards (dn, uFE, ...); the repeating-line stripper removes full lines only, not intra-cell shards. Watermark-heavy PDFs need cell-level rebuild (collect non-watermark spans per cell bbox).
  • Complex infographics (dense in-image text) come out as images only; transcribing in-image text needs a VLM pass, not this tool.

CJK Bold Spacing — why and how

DOCX uses run-level styling (no spaces between bold/normal runs in CJK text). Markdown renderers need whitespace around ** to recognize bold boundaries.

Rule: if a **content** span contains any CJK character, ensure both sides have a space — unless already spaced or at line boundary. This handles CJK punctuation, emoji adjacency, and mixed content.

Before: 打开**飞书**,就可以    → some renderers fail to bold
After:  打开 **飞书** ,就可以  → universally renders correctly

Heavy Mode Workflow

Heavy Mode runs multiple tools in parallel and selects the best segments:

  1. Parallel Execution: Run all applicable tools simultaneously
  2. Segment Analysis: Parse each output into segments (tables, headings, images, paragraphs)
  3. Quality Scoring: Score each segment based on completeness and structure
  4. Intelligent Merge: Select best version of each segment across tools

Merge Criteria

Segment TypeSelection Criteria
TablesMore rows/columns, proper header separator
ImagesAlt text present, local paths preferred
HeadingsProper hierarchy, appropriate length
ListsMore items, nested structure preserved
ParagraphsContent completeness

Image Extraction

# Extract images with metadata
uv run --with pymupdf scripts/extract_pdf_images.py document.pdf -o ./extracted-images

# Generate markdown references file
uv run --with pymupdf scripts/extract_pdf_images.py document.pdf --markdown refs.md

Output:

  • Images: extracted-images/img_page1_1.png, extracted-images/img_page2_1.jpg
  • Metadata: extracted-images/images_metadata.json (page, position, dimensions)

Quality Validation

# Validate conversion quality
uv run --with pymupdf scripts/validate_output.py document.pdf output.md

# Generate HTML report
uv run --with pymupdf scripts/validate_output.py document.pdf output.md --report report.html

Quality Metrics

MetricPassWarnFail
Text Retention>95%85-95%<85%
Table Retention100%90-99%<90%
Image Retention100%80-99%<80%

Merge Outputs Manually

# Merge multiple markdown files
python scripts/merge_outputs.py output1.md output2.md -o merged.md

# Show segment attribution
python scripts/merge_outputs.py output1.md output2.md -o merged.md --verbose

Path Conversion (Windows/WSL)

# Windows to WSL conversion
python scripts/convert_path.py "C:\Users\<windows-user>\Documents\file.pdf"
# Output: /mnt/c/Users/<windows-user>/Documents/file.pdf

Common Issues

"No conversion tools available"

# Install all tools
pip install pymupdf4llm
uv tool install "markitdown[pdf]"
brew install pandoc

FontBBox warnings during PDF conversion

  • Harmless font parsing warnings, output is still correct

Images missing from output

  • Use Heavy Mode for better image preservation
  • Or extract separately with scripts/extract_pdf_images.py

Tables broken in output

  • Use Heavy Mode - it selects the most complete table version
  • Or validate with scripts/validate_output.py

Bundled Scripts

ScriptPurpose
convert.pyMain orchestrator with Quick/Heavy mode + DOCX post-processing
html_to_markdown.pyPandoc HTML adapter and saved-output source-href retention verifier
test_convert.py31 tests covering all post-processing functions
merge_outputs.pyMerge multiple markdown outputs
validate_output.pyQuality validation with HTML report
extract_pdf_images.pyPDF image extraction with metadata
convert_path.pyWindows to WSL path converter

References

  • references/benchmark-2026-03-22.md - 5-tool benchmark (Docling/MarkItDown/Pandoc/Mammoth/ours)
  • references/heavy-mode-guide.md - Detailed Heavy Mode documentation
  • references/tool-comparison.md - Tool capabilities comparison
  • references/conversion-examples.md - Batch operation examples
  • references/html-conversion.md - Saved HTML scope, link retention, assets and manual heading assembly

Next Step: Clean Up Converted Content

After converting documents to markdown, suggest cleanup:

Conversion complete: [N] files converted to markdown.

Options:
A) Clean up docs — run /daymade-docs:docs-cleaner to consolidate redundant content (Recommended if multiple files)
B) Check facts — run /fact-checker to verify claims in the converted content
C) No thanks — the markdown conversion is sufficient

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Fixes web search on an agent whose model backend can't run it: a relay/reseller proxying Claude or Codex returns empty instead of failing. Use when web search returns nothing, a model insists a shipped product doesn't exist, someone wants to give an agent internet access, or the user is on a third-party base URL, relay, or 中转站. Diagnoses which built-in tools are dead, removes them, and installs a working replacement.

日本語の概要は準備中です。原文の説明を表示しています。

daymade/claude-code-skills1,4522026年10月11日 更新

抓取 A 股消息面情报:从财联社、华尔街见闻、金十、新浪 7x24、东财快讯、 证监会/央行/上交所/财政部政策公告、东方财富股吧等公开来源抓取与股票相关的 新闻、政策、情绪,输出结构化 JSON 或 Markdown。 当用户提到“A 股消息面”、“抓新闻”、“个股消息”、“政策监管”、“股吧情绪”、 “财联社”、“东财快讯”、“市场情绪”或需要把某只股票相关的公开情报聚合出来时 触发。也适用于“帮我看看 000001 最近有什么消息”这类口语化请求。

日本語の概要は準備中です。原文の説明を表示しています。

daymade/claude-code-skills1,4522026年10月11日 更新

Transcribes audio or video to speaker-labeled, timestamped text, locally with MLX on Apple Silicon or remotely. Use for 转录 / 录音转文字 / 说话人分离 / 字幕, and also for preparing audio for ASR without transcribing: 转格式, 降采样到 16kHz, merging recorder segments, or compressing and speeding up audio before 飞书妙记 — even when it looks like a one-line ffmpeg job.

日本語の概要は準備中です。原文の説明を表示しています。

daymade/claude-code-skills1,4522026年10月11日 更新

Routes audio: StepFun ASR/语音识别, StepFun TTS/配音, transcript/妙记→会议纪要, merge/review minutes. Reads one bundled specialist; generic ASR and correction stay direct.

日本語の概要は準備中です。原文の説明を表示しています。

daymade/claude-code-skills1,4522026年10月11日 更新

Diagnoses and repairs repository setup and guarded Git workflows for Claude Code or Codex — environment repair, startup sync, hook auditing, collaborator handoff. Use when a repo won't run, a teammate onboards, hook output duplicates, or commit/push/conflict needs guarding. Not for lost-commit recovery (use git-safety-net), GitHub ops (use github-ops), or history scrubbing (use github-sensitive-data-cleanup).

日本語の概要は準備中です。原文の説明を表示しています。

daymade/claude-code-skills1,4522026年10月11日 更新

Runs adversarial due-diligence on a benchmark the user envies — a founder, KOL, company, or product whose success looks inflated — splitting marketing bubble from real signal, then mapping the validated playbook onto the user's own resources. Use for 尽调/对标/拆解 a competitor, 抄/偷师 their playbook, or suspecting 水分/泡沫 in claims. Prefer over deep-research when debunking inflated claims, not a neutral briefing.

日本語の概要は準備中です。原文の説明を表示しています。

daymade/claude-code-skills1,4522026年10月11日 更新

daymade のスキルをすべて見る

このスキルの問題を報告する