Master 1Password: vaults, Watchtower, passkeys, SSH agent, CLI, and family/team administration. Use when getting full value from 1Password personally or administering it for others.
日本語の概要は準備中です。原文の説明を表示しています。
Building robust document ingestion — chunking, metadata, and multi-format pipelines for search and RAG.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Modern applications ingest documents for search, retrieval-augmented generation (RAG), and analytics. A document parser pipeline turns heterogeneous files (PDF, DOCX, HTML, spreadsheets, images) into clean, chunked, metadata-rich units ready for indexing. This skill covers the architecture and the details that determine retrieval quality.
Parse → clean → chunk → enrich. The pipeline stages: extract raw text with structure (headings, tables, lists); clean (fix hyphenation, normalize whitespace, drop boilerplate like repeated headers/footers); chunk into retrieval units; enrich each chunk with metadata (source, section, page, date). Skipping cleaning poisons everything downstream.
Chunking strategy drives retrieval quality. Fixed-size chunks with overlap are the baseline (e.g. ~500 tokens, 10–20% overlap). Better: respect document structure — chunk by section, keep tables intact, never split a table across chunks. A chunk should be self-contained enough to make sense out of context, which is why prepending section headings to each chunk helps enormously.
Metadata is a retrieval multiplier. Filters on metadata (date ranges, document type, source) often beat pure semantic search for precision. Capture rich metadata at ingest: title, section path, page numbers, author, date, URL, document version. Store it alongside embeddings, not just in them.
Tables need special handling. Serialize tables as Markdown or HTML within the chunk, keep them whole, and consider generating a textual summary of key figures ("Q3 revenue was $4.2M, up 12%") as an additional chunk — dense numeric tables embed poorly on their own.
Format routing. Detect file type and route to the right extractor (native PDF text vs OCR for scans, structure-aware DOCX parsing, readability extraction for HTML). One generic path for all formats produces mediocre results for every format.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Master 1Password: vaults, Watchtower, passkeys, SSH agent, CLI, and family/team administration. Use when getting full value from 1Password personally or administering it for others.
日本語の概要は準備中です。原文の説明を表示しています。
Create 3D visuals with modeling, texturing, lighting, rendering, and optimization for web and product.
日本語の概要は準備中です。原文の説明を表示しています。
Create 3D web experiences: scene setup, models, materials, lighting, animation, scroll-driven scenes, and performance budgets. Use when adding 3D to websites beyond basic demos.
日本語の概要は準備中です。原文の説明を表示しています。
Writing abstracts that get papers read — structured content, the 5-sentence core, and journal-specific constraints.
日本語の概要は準備中です。原文の説明を表示しています。
Learn effectively from courses and academies: choosing programs, studying actively, and converting courses into skills. Use when investing time/money in structured learning.
日本語の概要は準備中です。原文の説明を表示しています。
Audit designs for accessibility with WCAG checklists covering color, type, focus, motion, and content.
日本語の概要は準備中です。原文の説明を表示しています。