本文へ移動
cccskills
無料GitHub で公開

document-extraction-api

Extract structured data from documents using AI-powered field extraction.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md25.7 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Document Extraction API

Extract structured data from documents using AI-powered field extraction.

Cost: 1 credit per page

Prerequisites

You need an Iteration Layer API key. Get one at platform.iterationlayer.com during the 7-day trial.

For full integration guidance (SDKs, auth, MCP, error handling), see the Iteration Layer Integration Guide.

API Reference

Extract structured data from any document with a single API call. Send one or more files and a schema defining the fields you need, and receive typed, validated results with confidence scores and source citations.

Key Features

  • Multi-Format Support — Extract from 40+ file formats: PDF, Office documents (DOCX, PPTX, ODT, ODS, XLSX), public website URLS, EPUB, LaTeX, email (EML), Jupyter notebooks, images, and text/markup formats.
  • Rich Field Types — Primitive types (text, number, date, boolean, email, enum) plus purpose-built validated types: IBAN (validated against the standard format), ADDRESS (returns a structured object with street, city, region, postal code, and country), CURRENCY_CODE (normalized to ISO 4217), CURRENCY_AMOUNT (numeric monetary value), and COUNTRY (normalized to ISO 3166-1 alpha-2). These validated types mean you get clean, usable data — not raw strings you have to parse yourself.
  • Structured Arrays — Extract repeating data like invoice line items with nested schemas.
  • Calculated Fields — Define arithmetic operations (sum, subtract, multiply, divide) computed from other extracted fields.
  • Confidence Scores — Every extracted value includes a confidence score between 0 and 1.
  • Source Citations — Verbatim quotes from the document that support each extracted value.
  • Website Image Context — Public website URL inputs include useful referenced images in the extraction context while skipping logos, tracking pixels, decorative graphics, and semantically unhelpful images.
  • Schema Validation — Field schemas are validated before extraction, catching errors like circular dependencies or type mismatches early.

Overview

The Document Extraction API analyzes documents and extracts structured data based on a schema you define. You send one or more files (base64 or URL) and a schema with field definitions, and receive a JSON response with typed values, confidence scores, and citations. Public website URLs are supported when you want website content included in a broader multi-file extraction.

Endpoint: POST /document-extraction/v1/extract

Limits:

  • Max files per request: 20
  • Max file size: 50 MB per file

Supported File Formats

  • Documents: PDF, DOCX, PPTX, ODT, EPUB, RTF
  • Spreadsheets: XLSX, XLS, ODS, CSV, TSV
  • Email: EML, MSG (headers, body, and attachment extraction — attachments are ingested and included in extraction context)
  • Notebooks: Jupyter (.ipynb)
  • Academic & Publishing: LaTeX (.tex, .latex), BibTeX (.bib), Typst (.typst, .typ)
  • Markup & Text: HTML, Markdown, JSON, XML, YAML, TOML, RST, Org, Djot, MDX, TXT
  • Images: PNG, JPEG, GIF, WebP, AVIF, HEIF, BMP, TIFF, JP2, PNM/PBM/PGM/PPM, SVG

How It Works

Every extraction runs the same pipeline:

  1. Validate — the schema is checked before any files are touched. Three things are validated: CALCULATED source field references (each must exist in the schema and be a numeric type), circular dependencies between CALCULATED fields, and default values matching their field type. The first failure stops validation with a descriptive error. No LLM call is made if the schema is invalid.
  2. Ingest — files are converted to a normalized text representation.
  3. Extract — each field is extracted according to its type and configuration. All non-CALCULATED fields are handled together, with all source files available during extraction. Every extracted value is tagged with the file it came from.
  4. Calculate — CALCULATED fields are computed from the extraction results. Operations are pure arithmetic: sum, subtract, multiply, divide. The confidence of a CALCULATED field is the minimum confidence of its source fields. Division by zero returns 0. The result is deterministic — if the source values are correct, the computed value is exact.
  5. Consolidate — default values are applied to fields that were not extracted, and required field constraints are checked.

Intelligent extraction — the API automatically selects the best extraction approach based on the complexity of your schema and the nature of the documents. You don't configure this.

For a technical deep dive into the ingestion pipeline that runs before schema extraction, see How Our Document Ingestion Pipeline Turns Files into LLM-Ready Markdown.

Request Format

<!-- tabs -->
curl -X POST \
  https://api.iterationlayer.com/document-extraction/v1/extract \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "files": [
      {
        "type": "base64",
        "name": "invoice.pdf",
        "base64": "<base64-encoded-file>"
      }
    ],
    "schema": {
      "fields": [
        {
          "name": "invoice_number",
          "type": "TEXT",
          "description": "The invoice number"
        },
        {
          "name": "total_amount",
          "type": "CURRENCY_AMOUNT",
          "description": "Total invoice amount"
        }
      ]
    }
  }'
import { IterationLayer } from "iterationlayer";
const client = new IterationLayer({
  apiKey: "YOUR_API_KEY",
});

const result = await client.extractDocument({
  files: [
    {
      type: "base64",
      name: "invoice.pdf",
      base64: new Uint8Array([/* file bytes */]),
    },
  ],
  schema: {
    fields: [
      {
        name: "invoice_number",
        type: "TEXT",
        description: "The invoice number",
      },
      {
        name: "total_amount",
        type: "CURRENCY_AMOUNT",
        description: "Total invoice amount",
      },
    ],
  },
});
from iterationlayer import IterationLayer
client = IterationLayer(api_key="YOUR_API_KEY")

result = client.extract_document(
    files=[
        {
            "type": "base64",
            "name": "invoice.pdf",
            "base64": b"...",
        }
    ],
    schema={
        "fields": [
            {
                "name": "invoice_number",
                "type": "TEXT",
                "description": "The invoice number",
            },
            {
                "name": "total_amount",
                "type": "CURRENCY_AMOUNT",
                "description": "Total invoice amount",
            },
        ]
    },
)
import il "github.com/iterationlayer/sdk-go"
client := il.NewClient("YOUR_API_KEY")

result, err := client.ExtractDocument(il.ExtractDocumentRequest{
	Files: []il.FileInput{
		il.FileInput{Type: "base64", Name: "invoice.pdf", Base64: []byte{ /* file bytes */ }},
	},
	Schema: il.ExtractionSchema{
		Fields: []any{
			il.TextFieldConfig{Type: "TEXT", Name: "invoice_number", Description: "The invoice number"},
			il.CurrencyAmountFieldConfig{Type: "CURRENCY_AMOUNT", Name: "total_amount", Description: "Total invoice amount"},
		},
	},
})
<!-- response -->
{
  "success": true,
  "data": {
    "invoice_number": {
      "type": "TEXT",
      "value": "INV-2024-001",
      "confidence": 0.98,
      "citations": ["Invoice No: INV-2024-001"],
      "source": "invoice.pdf"
    },
    "total_amount": {
      "type": "CURRENCY_AMOUNT",
      "value": 1250.00,
      "confidence": 0.95,
      "citations": ["Total: €1,250.00"],
      "source": "invoice.pdf"
    }
  }
}
<!-- /tabs -->

Top-Level Fields

FieldTypeRequiredDescription
filesarrayYesList of file inputs to extract from (see File Input below)
schemaobjectYesExtraction schema defining the fields to extract
webhook_urlstringNoHTTPS URL to receive results asynchronously. If provided, returns 201 immediately. See Webhooks.

Async Mode

Add a webhook_url parameter to process the request in the background. The API returns 201 Accepted immediately and delivers the result to your webhook URL when processing completes. See Webhooks for payload format and retry behavior.

File Input

Provide each file as either base64 or a URL:

FieldTypeRequiredDescription
typestringYes"base64" or "url"
namestringRequired for base64, optional for urlFilename with extension (e.g., invoice.pdf). URL inputs without a name may be resolved as website pages.
base64stringIf type is base64Base64-encoded file content
urlstringIf type is urlPublicly accessible HTTPS URL to fetch the file from. HTTP is not accepted.
fetch_optionsobjectNoWebsite retrieval options for URL inputs. Use only when the URL should be treated as a website page with explicit retrieval controls.

For websites the API always checks the site's robots.txt before fetching the page. If the URL is disallowed for the effective User-Agent, the request fails with 400 Bad Request and the page is not fetched. Sitemap entries in robots.txt are recorded as metadata only; they are not traversed.

Website URL fetches are rate-limited per destination host. The service automatically respects robots.txt crawl-delay hints, slows down after upstream 429 Too Many Requests responses, and uses upstream Retry-After or rate-limit reset headers when provided. This protects public sites and may delay later requests to the same host; it is not configurable through fetch_options.

If a website URL fetch looks like a WAF or security challenge, the API automatically retries the same URL in an isolated Chromium browser session. Detection covers common Cloudflare, Akamai, AWS WAF/CloudFront, Imperva/Incapsula, DataDome, PerimeterX, Sucuri, F5/BIG-IP, hCaptcha, reCAPTCHA, and generic verification pages. Successful browser fallback is recorded under metadata.security; failed fallback returns 400 Bad Request with a clear security-challenge message.

Fetch Options

fetch_options lets you control retrieval behavior for a URL input treated as a website page. Robots compliance and per-host rate limiting are mandatory and are not configurable.

FieldTypeRequiredDescription
localestringNoBCP 47 locale tag sent as Accept-Language (e.g., "en-US", "de-DE", "fr").
user_agentstringNoCustom User-Agent header string (1–500 characters).
authobjectNoAuthentication for website URL inputs. Supports bearer tokens, HTTP Basic auth, and a custom auth header. Secret values are not returned in metadata.
headersobjectNoAdditional request headers for website URL inputs. Header names and values are validated; unsafe headers such as Cookie, Set-Cookie, Host, Content-Length, hop-by-hop headers, and browser-controlled Sec-* headers are rejected.
timeout_msintegerNoWebsite fetch timeout in milliseconds. Must be between 1000 and 60000.
should_render_javascriptbooleanNoUse Chromium browser rendering before extraction. Default: false.

Custom Headers

Custom headers are sent on direct fetches and, when JavaScript rendering is enabled or security-challenge fallback is needed, as Chromium request headers where supported by the browser. Each Chromium fetch uses an isolated browser context that is disposed after the request; cookies and session state are not persisted across URLs or requests.

Fetch options control retrieval behavior, not billing.

Auth Examples

Use one auth shape per request.

{ "auth": { "type": "bearer", "token": "..." } }
{ "auth": { "type": "basic", "username": "...", "password": "..." } }
{ "auth": { "type": "custom_header", "name": "x-api-key", "value": "..." } }

URL Example

Document Extraction can also ingest a public website URL as one input in a broader multi-file extraction. Use Website Extraction instead when the request should contain exactly one public website page.

{
  "files": [
    {
      "type": "url",
      "url": "https://example.com/pricing"
    }
  ],
  "schema": {
    "fields": [
      {
        "name": "plan_name",
        "type": "TEXT",
        "description": "The name of the pricing plan"
      }
    ]
  }
}

Schema Definition

The schema object contains a fields array. Each field has these base properties:

FieldTypeRequiredDescription
namestringYesUnique field identifier (used as key in the response)
typestringYesOne of the supported field types (see Field Types below)
descriptionstringYesNatural language description of what to extract
is_requiredbooleanNoIf true, returns an error when the field cannot be extracted and has no default. Default: false

Each field type supports additional type-specific properties described below.

Field Types

TEXT

Short single-line text value.

PropertyTypeRequiredDescription
max_lengthintegerNoMaximum character length (> 0)
default_valuestringNoDefault value if not found
{
  "name": "company_name",
  "type": "TEXT",
  "description": "Name of the company"
}

TEXTAREA

Multi-line text value.

PropertyTypeRequiredDescription
max_lengthintegerNoMaximum character length (> 0)
default_valuestringNoDefault value if not found
{
  "name": "notes",
  "type": "TEXTAREA",
  "description": "Additional notes or comments"
}

INTEGER

Whole number value.

PropertyTypeRequiredDescription
minintegerNoMinimum value (inclusive)
maxintegerNoMaximum value (inclusive)
unitstringNoUnit label (e.g., "kg", "items")
default_valueintegerNoDefault value if not found
{
  "name": "quantity",
  "type": "INTEGER",
  "description": "Number of items ordered",
  "min": 1
}

DECIMAL

Floating-point number value.

PropertyTypeRequiredDescription
minfloatNoMinimum value (inclusive)
maxfloatNoMaximum value (inclusive)
decimal_pointsintegerNoNumber of decimal places to round to (>= 0)
unitstringNoUnit label
default_valuefloatNoDefault value if not found
{
  "name": "weight",
  "type": "DECIMAL",
  "description": "Package weight",
  "unit": "kg",
  "decimal_points": 2
}

DATE

Calendar date, extracted as an ISO 8601 string (YYYY-MM-DD).

PropertyTypeRequiredDescription
allow_future_datesbooleanNoWhether to allow dates in the future
allow_past_datesbooleanNoWhether to allow dates in the past
{
  "name": "invoice_date",
  "type": "DATE",
  "description": "Date the invoice was issued"
}

DATETIME

Date and time, extracted as an ISO 8601 datetime string.

PropertyTypeRequiredDescription
allow_future_datesbooleanNoWhether to allow dates in the future
allow_past_datesbooleanNoWhether to allow dates in the past
{
  "name": "timestamp",
  "type": "DATETIME",
  "description": "Transaction timestamp"
}

TIME

Time value (e.g., "14:30:00"). No additional parameters.

{
  "name": "delivery_time",
  "type": "TIME",
  "description": "Scheduled delivery time"
}

ENUM

One or more values from a predefined list. Extracted as a string array.

PropertyTypeRequiredDescription
valuesstring[]YesAllowed options
min_selectedintegerNoMinimum number of selected values (>= 0)
max_selectedintegerNoMaximum number of selected values (> 0)
default_valuestring[]NoDefault selected values
{
  "name": "payment_method",
  "type": "ENUM",
  "description": "How the invoice was paid",
  "values": ["bank_transfer", "credit_card", "cash", "paypal"],
  "max_selected": 1
}

BOOLEAN

True or false value.

PropertyTypeRequiredDescription
default_valuebooleanNoDefault value if not found
{
  "name": "is_paid",
  "type": "BOOLEAN",
  "description": "Whether the invoice has been paid"
}

EMAIL

Email address string.

PropertyTypeRequiredDescription
default_valuestringNoDefault value if not found
{
  "name": "contact_email",
  "type": "EMAIL",
  "description": "Contact email address"
}

IBAN

International Bank Account Number. Validated against the pattern ^[A-Z]{2}\d{2}[A-Z0-9]{11,30}$.

PropertyTypeRequiredDescription
default_valuestringNoDefault value if not found
{
  "name": "bank_account",
  "type": "IBAN",
  "description": "Recipient IBAN"
}

COUNTRY

ISO 3166-1 alpha-2 country code (e.g., "DE", "US").

PropertyTypeRequiredDescription
default_valuestringNoMust be a valid ISO 3166-1 alpha-2 code
{
  "name": "origin_country",
  "type": "COUNTRY",
  "description": "Country of origin"
}

CURRENCY_CODE

ISO 4217 currency code (e.g., "EUR", "USD").

PropertyTypeRequiredDescription
default_valuestringNoMust be a valid ISO 4217 code
{
  "name": "currency",
  "type": "CURRENCY_CODE",
  "description": "Invoice currency"
}

CURRENCY_AMOUNT

Numeric monetary amount.

PropertyTypeRequiredDescription
minfloatNoMinimum value (inclusive)
maxfloatNoMaximum value (inclusive)
decimal_pointsintegerNoNumber of decimal places to round to (>= 0)
default_valuefloatNoDefault value if not found
{
  "name": "total_amount",
  "type": "CURRENCY_AMOUNT",
  "description": "Total invoice amount",
  "decimal_points": 2
}

ADDRESS

Structured address object. Extracted as an object with street, city, region, postal_code, and country fields.

PropertyTypeRequiredDescription
allowed_country_codesstring[]NoRestrict to specific ISO 3166-1 alpha-2 country codes

Response value shape:

{
  "street": "123 Main St",
  "city": "Berlin",
  "region": "Berlin",
  "postal_code": "10115",
  "country": "DE"
}
{
  "name": "billing_address",
  "type": "ADDRESS",
  "description": "Billing address",
  "allowed_country_codes": ["DE", "AT", "CH"]
}

ARRAY

A list of structured objects, each conforming to a nested schema. Use this for repeating data like line items.

PropertyTypeRequiredDescription
fieldsarrayYesArray of field configurations for each item

The fields array uses the same field configuration format as top-level fields.

{
  "name": "line_items",
  "type": "ARRAY",
  "description": "Invoice line items",
  "fields": [
    {
      "name": "description",
      "type": "TEXT",
      "description": "Item description"
    },
    {
      "name": "quantity",
      "type": "INTEGER",
      "description": "Quantity ordered",
      "min": 1
    },
    {
      "name": "unit_price",
      "type": "CURRENCY_AMOUNT",
      "description": "Price per unit",
      "decimal_points": 2
    },
    {
      "name": "total",
      "type": "CURRENCY_AMOUNT",
      "description": "Line item total",
      "decimal_points": 2
    }
  ]
}

CALCULATED

A derived numeric value computed from other extracted fields. Not extracted from the document — calculated locally after all source fields are resolved.

PropertyTypeRequiredDescription
operationstringYesOne of: "sum", "subtract", "multiply", "divide"
source_field_namesstring[]YesNames of fields to apply the operation to, in order
unitstringNoUnit label

Source fields must be numeric types: INTEGER, DECIMAL, CURRENCY_AMOUNT, or another CALCULATED. Circular dependencies are detected and rejected at validation time.

The confidence score of a CALCULATED field is the minimum confidence of its source fields. Division by zero returns 0.

{
  "name": "tax_amount",
  "type": "CALCULATED",
  "description": "Tax amount (total minus net)",
  "operation": "subtract",
  "source_field_names": ["total_amount", "net_amount"]
}

Response Format

Success Response

{
  "success": true,
  "data": {
    "invoice_number": {
      "type": "TEXT",
      "value": "INV-2024-001",
      "confidence": 0.98,
      "citations": ["Invoice No: INV-2024-001"],
      "source": "invoice.pdf"
    },
    "total_amount": {
      "type": "CURRENCY_AMOUNT",
      "value": 1250.00,
      "confidence": 0.95,
      "citations": ["Total: €1,250.00"],
      "source": "invoice.pdf"
    }
  }
}

Each field in data contains:

FieldTypeDescription
typestringThe field type from the schema
valuevariesExtracted value (type depends on field type — see below)
confidencefloatConfidence score between 0.0 and 1.0
citationsstring[]Verbatim quotes from the source document
sourcestringFilename the value was extracted from

Value types by field type:

Field TypeValue Type
TEXT, TEXTAREA, EMAIL, IBAN, COUNTRY, CURRENCY_CODE, DATE, DATETIME, TIMEstring
INTEGER, DECIMAL, CURRENCY_AMOUNT, CALCULATEDnumber
BOOLEANboolean
ENUMstring[]
ADDRESSobject
ARRAYobject[]

Fields that could not be extracted and have no default_value are omitted from data (unless is_required is true, which causes an error). Fields resolved via default_value have a confidence of 1.0.

Best Practices

Use confidence scores as workflow inputs when extracted values can update records, trigger business decisions, or appear in generated outputs.

For low-risk fields, your application can auto-accept values above a field-specific threshold. For high-impact fields such as payment amounts, IBANs, tax IDs, contract dates, consent fields, or customer-facing claims, route low-confidence values to human review before the workflow continues. Store extracted values separately from approved values so downstream systems can tell whether a value is still a candidate or has been reviewed.

Downstream systems should read from approved values, not raw extracted values. An approved value may be auto-accepted or human-corrected. Keep confidence scores as routing signals, citations as review context, and audit events as metadata-only workflow facts rather than application logs that retain full document content.

For auditability, record the extraction schema version, provider route where your policy requires it, workflow version, source record ID, confidence summary, review decision, approved-value snapshot, and downstream delivery status. These fields let operators explain which workflow accepted a value without turning operational logging into a second content store.

See Workflow Design Best Practices for reliability primitives such as confidence gates, review branches, approved values, audit records, retention boundaries, and clear data-flow behavior.

Recipes

For complete, runnable examples see the Recipes page.

  • Extract Invoice Data -- Extract line items, totals, and vendor details from an invoice into structured JSON.
  • Extract Resume Data -- Extract contact info, work history, and skills from a resume into structured data.
  • Extract Medical Record -- Extract patient details, diagnoses, and medications from a medical record into structured JSON.
  • Extract Receipt Data -- Extract merchant, amount, date, and line items from a receipt image or PDF.
  • Extract Multi-Invoice Data -- Extract structured data from multiple invoice files in a single API call using an array schema.

Error Responses

All errors return a JSON body with { "success": false, "error": "<message>" }.

StatusDescription
400Invalid request (missing files/schema, invalid base64, URL fetch failure, file size exceeded, invalid field config)
401Missing or invalid API key
402Insufficient credits or budget cap exceeded
422Processing error (circular dependency in CALCULATED fields, required field not extractable, LLM parsing failure)
429Rate limit exceeded

Links

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Resize, sharpen, and compress an image to fit email platform size limits in a single pipeline.

日本語の概要は準備中です。原文の説明を表示しています。

iterationlayer/skills42026年6月9日 更新

Compress an image to fit within a specific file size in bytes using quality-first compression.

日本語の概要は準備中です。原文の説明を表示しています。

iterationlayer/skills42026年6月9日 更新

Convert a contract PDF to clean markdown for clause extraction or LLM analysis.

日本語の概要は準備中です。原文の説明を表示しています。

iterationlayer/skills42026年6月9日 更新

Convert external documents — specs, contracts, reports — to markdown for knowledge base ingestion.

日本語の概要は準備中です。原文の説明を表示しています。

iterationlayer/skills42026年6月9日 更新

Convert a document to clean markdown suitable for chunking and embedding in a RAG pipeline.

日本語の概要は準備中です。原文の説明を表示しています。

iterationlayer/skills42026年6月9日 更新

Convert an image between PNG, JPEG, and WebP formats with quality control for web optimization.

日本語の概要は準備中です。原文の説明を表示しています。

iterationlayer/skills42026年6月9日 更新

iterationlayer のスキルをすべて見る

このスキルの問題を報告する