本文へ移動
cccskills
無料GitHub で公開

databricks-unstructured-pdf-generation

Build RAG / unstructured-document evaluation datasets and demo documents (e.g. for Knowledge Assistant) on Databricks: generate synthetic PDFs locally, upload to Unity Catalog volumes, and pair each document with test questions for retrieval evaluation.

インストール方法を見る

含まれるファイル(5)

  • SKILL.md6.6 KB
  • agents/openai.yaml558 B
  • assets/databricks.png15.0 KB
  • assets/databricks.svg582 B
  • scripts/pdf_generator.py8.2 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Unstructured-Document for Demos and Eval Datasets on Databricks

Workflow for producing synthetic PDF documents + paired test questions as a Unity Catalog-resident dataset for Demos and RAG / unstructured-document retrieval evaluation on Databricks. The PDF-generation step uses standard local HTML → PDF tooling; the Databricks-specific value is the workflow shape — UC volume layout, paired question files, and integration with downstream Databricks retrieval / ai_extract / ai_parse_document evaluation.

Workflow

  1. Write HTML files to ./raw_data/html/ (write multiple files in parallel for speed) — domain-shaped to match the documents your retrieval pipeline will see in production.
  2. Convert HTML → PDF using <SKILL_ROOT>/scripts/pdf_generator.py (parallel conversion, wraps plutoprint).
  3. Upload PDFs to a Unity Catalog volume via databricks fs cp — same volume shape your production pipeline will read from.
  4. Generate ./raw_data/pdf/pdf_eval_questions.json pairing each document with retrieval-eval questions; this becomes the gold dataset for mlflow.genai.evaluate() or comparable retrieval-quality scorers.

If you only need ad-hoc PDFs (no Databricks workflow), any HTML → PDF tool (weasyprint, wkhtmltopdf, playwright pdf, plutoprint) works directly — this skill exists for the synthetic-dataset-on-UC end-to-end shape, not as a general PDF generator.

Path convention: <SKILL_ROOT> below = the directory containing this SKILL.md. Resolve to the absolute install path (e.g. ~/.claude/skills/databricks-unstructured-pdf-generation). ./raw_data/... paths are relative to your own project cwd.

Dependencies

uv pip install plutoprint

Step 1: Write HTML Files

mkdir -p ./raw_data/html

Write HTML documents to ./raw_data/html/filename.html. Use subdirectories to organize (structure is preserved).

Step 2: Convert to PDF

# Convert entire folder (parallel, 4 workers)
python <SKILL_ROOT>/scripts/pdf_generator.py convert --input ./raw_data/html --output ./raw_data/pdf

Skips files where PDF exists and is newer than HTML. Use --force to reconvert all.

Step 3: Upload to Volume

databricks fs requires the dbfs: scheme prefix even for UC Volume paths. -r copies the contents of the source directory into the target (the source directory name is not preserved), so name the target raw_data/pdf explicitly to keep the PDFs in their own folder on the volume. They land under raw_data/pdf/ — i.e. dbfs:/Volumes/my_catalog/my_schema/raw_data/pdf/report.pdf — so a Knowledge Assistant or ingest pipeline can point at that single folder.

databricks fs cp -r --overwrite ./raw_data/pdf dbfs:/Volumes/my_catalog/my_schema/raw_data/pdf

Step 4: Generate Test Questions

Create ./raw_data/pdf/pdf_eval_questions.json with questions for Knowledge Assistant (KA) or Multi-Agent Supervisor (MAS) evaluation. It's fine for this file to be uploaded to the volume alongside the PDFs — downstream agents can use it:

{
  "api_errors_guide.pdf": {
    "question": "What is the solution for error ERR-4521?",
    "expected_fact": "Call /api/v2/auth/refresh with refresh_token before the 3600s TTL expires"
  },
  "installation_manual.pdf": {
    "question": "What port does the service use by default?",
    "expected_fact": "Port 8443 for HTTPS, configurable via CONFIG_PORT environment variable"
  }
}

This JSON can be used to build KA test cases and validate retrieval accuracy.

Document Content Guidelines

When generating documents for Knowledge Assistant testing or demos:

  • Multi-page documents: Each PDF should be several pages with substantial content
  • Specific error codes and solutions: Include product-specific error codes, causes, and resolution steps
  • Technical details: API endpoints, configuration parameters, version numbers, specific commands
  • Simple CSS: Keep styling minimal for fast HTML creation and reliable PDF conversion
  • Queryable facts: Include details a KA must read the document to answer (not general knowledge)

Good document types:

  • Product user manuals with troubleshooting sections
  • API error reference guides (error codes, causes, solutions)
  • Installation/configuration guides with specific steps
  • Technical specifications with version-specific details

Example content: Instead of generic "Connection failed" errors, write:

  • "Error ERR-4521: OAuth token expired. Cause: Token TTL exceeded 3600s default. Solution: Call /api/v2/auth/refresh with your refresh_token before expiration. See Section 4.2 for token lifecycle management."

CLI Reference

python <SKILL_ROOT>/scripts/pdf_generator.py convert [OPTIONS]

  --input, -i     Input HTML file or folder (required)
  --output, -o    Output folder for PDFs (required)
  --force, -f     Force reconvert (ignore timestamps)
  --workers, -w   Parallel workers (default: 4)

Folder Structure

Subfolder structure is preserved:

./raw_data/html/                    ./raw_data/pdf/
├── report.html             →       ├── report.pdf
├── quarterly/                      ├── quarterly/
│   └── q1.html             →       │   └── q1.pdf
└── legal/                          └── legal/
    └── terms.html          →           └── terms.pdf

Bundled Script

This skill ships one helper script:

FileDescription
scripts/pdf_generator.pyHTML → PDF converter (wraps plutoprint); parallel folder conversion with timestamp-skip. Referenced by Step 2 and the CLI Reference.

The script ships at <SKILL_ROOT>/scripts/pdf_generator.py. If it is absent, recreate it from the CLI Reference above (a convert subcommand taking --input/--output/--force/--workers, wrapping plutoprint for HTML → PDF).

Troubleshooting

IssueSolution
"plutoprint not installed"uv pip install plutoprint
PDF looks wrongCheck HTML/CSS syntax
"Volume does not exist"databricks volumes create CATALOG SCHEMA VOLUME_NAME MANAGED (four separate positional args, not catalog.schema.volume)

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Create Agent Bricks: Knowledge Assistants (KA) for document Q&A and Supervisor Agents for multi-agent orchestration (MAS).

日本語の概要は準備中です。原文の説明を表示しています。

databricks/databricks-agent-skills3452026年10月10日 更新

Use Databricks built-in AI Functions (ai_classify, ai_extract, ai_summarize, ai_mask, ai_translate, ai_fix_grammar, ai_gen, ai_analyze_sentiment, ai_similarity, ai_parse_document, ai_prep_search, ai_query, ai_forecast) to add AI capabilities directly to SQL and PySpark pipelines without managing model endpoints. Also covers document parsing and building custom RAG pipelines (parse → prep_search → index → query).

日本語の概要は準備中です。原文の説明を表示しています。

databricks/databricks-agent-skills3452026年10月10日 更新

Databricks AI Runtime, the `databricks air` CLI commands for submitting and managing GPU training workloads on Databricks serverless compute. Use for: writing and submitting `databricks air` workload YAML, passing hyperparameters and secrets, checking run status, listing/cancelling runs, streaming a run's logs and watching its progress, custom Docker image setup, and environment configuration.

日本語の概要は準備中です。原文の説明を表示しています。

databricks/databricks-agent-skills3452026年10月10日 更新

Create Databricks AI/BI dashboards. Must use when creating, updating, or deploying Lakeview dashboards as Databricks Dashboard have a unique json structure. CRITICAL: You MUST test ALL SQL queries via CLI BEFORE deploying. Follow guidelines strictly.

日本語の概要は準備中です。原文の説明を表示しています。

databricks/databricks-agent-skills3452026年10月10日 更新

Design the UX of custom-code Databricks Apps (AppKit/React) data screens — KPI/overview pages, reports, charts, tables, and Genie/chat data assistants — mapped to concrete AppKit components. Use when BUILDING or reviewing the UI of an AppKit/React app that displays data or answers data questions: choosing genre, layout, charts, KPIs, semantic color, required states (loading/empty/error), IBCS notation, and AI-result trust (showing generated SQL/sources for Genie/chat). A plain "create a dashboard" request means a managed AI/BI (Lakeview) dashboard → use databricks-aibi-dashboards, NOT this skill. Also NOT for non-data frontend (forms, settings, auth, marketing) or scaffolding/build/deploy (→ databricks-apps). Complements databricks-apps; use it alongside whenever a custom app has a chart, table, KPI, report, or Genie/chat/AI surface.

日本語の概要は準備中です。原文の説明を表示しています。

databricks/databricks-agent-skills3452026年10月10日 更新

Build apps on Databricks Apps platform. Use when asked to create data apps, analytics tools, or custom interactive visualizations. A plain "create a dashboard" request means a managed AI/BI (Lakeview) dashboard → use databricks-aibi-dashboards, not this skill. Evaluates data access patterns (analytics vs Lakebase synced tables) before scaffolding. Invoke BEFORE starting implementation.

日本語の概要は準備中です。原文の説明を表示しています。

databricks/databricks-agent-skills3452026年10月10日 更新

databricks のスキルをすべて見る

このスキルの問題を報告する