本文へ移動
cccskills
無料GitHub で公開

scrapling

Scrape pages locally with anti-bot bypass, TLS impersonation, and adaptive element tracking — no API keys, no cloud. Handles Cloudflare protection, CSS/XPath element extraction, and survives site redesigns. Complements firecrawl (cloud) with 100% local execution. Triggers on Cloudflare bypass, anti-bot scraping, stealth fetch, local scraping, Scrapling.

インストール方法を見る

含まれるファイル(7)

  • SKILL.md7.9 KB
  • README.md3.9 KB
  • references/adaptive-scraping.md3.2 KB
  • references/cli-reference.md3.6 KB
  • references/python-api-reference.md6.6 KB
  • scripts/scrapling_fetch.py5.5 KB
  • scripts/scrapling_install.sh838 B

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Scrapling -- Local Stealth Web Scraping

100% local Python library (BSD-3, D4Vinci/Scrapling). No API keys, no cloud dependencies. Built-in Cloudflare solver, TLS impersonation, and adaptive element tracking.

When to Use Scrapling vs Firecrawl

NeedUseWhy
Clean markdown from a URLfirecrawl scrape --only-main-contentOptimized for LLM markdown conversion
Bypass Cloudflare/anti-botscrapling stealth fetchBuilt-in Turnstile solver, Patchright stealth
Extract specific elementsscrapling with CSS/XPath selectorsElement-level precision, adaptive tracking
No API key availablescrapling100% local, zero credentials
Batch cloud scrapingfirecrawl crawl / batch-scrapeCloud infrastructure, parallel processing
Site redesign resiliencescrapling adaptive modeSQLite-backed similarity matching
Full-site concurrent crawlscrapling Spider frameworkScrapy-like with pause/resume
Web search + scrapefirecrawl search --scrapeCombined search + extraction

Installation

Run once to set up Scrapling with all features:

~/.claude/skills/scrapling/scripts/scrapling_install.sh

Installs scrapling[all] via uv, downloads Chromium + system dependencies, and verifies all fetchers load.

Quick Start -- Stdout Wrapper

The wrapper uses Scrapling's Python API directly (faster than CLI, avoids curl_cffi cert issues) and outputs to stdout for piping into filter_web_results.py:

# Basic HTTP fetch (fastest, TLS impersonation)
python3 ~/.claude/skills/scrapling/scripts/scrapling_fetch.py https://example.com

# Stealth mode (Patchright, anti-bot bypass)
python3 ~/.claude/skills/scrapling/scripts/scrapling_fetch.py https://protected.site --stealth

# Stealth + Cloudflare solver
python3 ~/.claude/skills/scrapling/scripts/scrapling_fetch.py https://cf-protected.site --stealth --solve-cloudflare

# Dynamic (Playwright Chromium, JS rendering)
python3 ~/.claude/skills/scrapling/scripts/scrapling_fetch.py https://js-heavy.site --dynamic

# With CSS selector for targeted extraction
python3 ~/.claude/skills/scrapling/scripts/scrapling_fetch.py https://example.com --css ".product-list"

# Pipe through firecrawl's filter for token efficiency
python3 ~/.claude/skills/scrapling/scripts/scrapling_fetch.py https://example.com --stealth | \
  python3 ~/.claude/skills/firecrawl/scripts/filter_web_results.py --sections "Pricing" --max-chars 5000

Full path: python3 ~/.claude/skills/scrapling/scripts/scrapling_fetch.py Flags: --stealth, --dynamic, --css SELECTOR, --solve-cloudflare, --impersonate BROWSER, --format {text,html}, --no-headless, --timeout SECONDS

For --network-idle, --real-chrome, or POST requests, use the CLI direct path below.

Quick Start -- CLI Direct

For file-based output (Scrapling's native CLI):

# HTTP fetch -> markdown
scrapling extract get 'https://example.com' content.md

# HTTP with CSS selector and browser impersonation
scrapling extract get 'https://example.com' content.md --css-selector '.main-content' --impersonate chrome

# Dynamic fetch (Playwright, JS rendering)
scrapling extract fetch 'https://example.com' content.md

# Stealth with Cloudflare bypass
scrapling extract stealthy-fetch 'https://protected.site' content.md --solve-cloudflare

# Stealth with CSS selector, visible browser for debugging
scrapling extract stealthy-fetch 'https://site.com' content.md --css-selector '.data' --no-headless

Three Fetcher Tiers

CLI verbEngineStealthJSSpeed
getcurl_cffi (HTTP)TLS impersonationNoFast
fetchPlaywright/ChromiumMediumYesMedium
stealthy-fetchPatchright/ChromeMaximumYesSlower

Python API

For element-level extraction, automation, or when the CLI is insufficient:

from scrapling.fetchers import Fetcher, StealthyFetcher, DynamicFetcher

# Simple HTTP fetch with CSS extraction
page = Fetcher.get('https://example.com', impersonate='chrome')
titles = page.css('.item h2::text').getall()
links = page.css('a::attr(href)').getall()

# Stealth with Cloudflare bypass
page = StealthyFetcher.fetch('https://protected.site',
    headless=True, solve_cloudflare=True,
    hide_canvas=True, block_webrtc=True)
data = page.css('.content').get_all_text()

# Page automation (login, click, fill)
def login(page):
    page.fill('#username', 'user')
    page.fill('#password', 'pass')
    page.click('#submit')

page = StealthyFetcher.fetch('https://app.example.com', page_action=login)

Parsing (no fetching)

from scrapling.parser import Selector

page = Selector("<html>...</html>")
page.css('.item::text').getall()       # CSS with pseudo-elements
page.xpath('//div[@class="item"]')     # XPath
page.find_all('div', class_='item')    # BeautifulSoup-style
page.find_by_text('Add to Cart')       # Text search
page.find_by_regex(r'Price: \$\d+')    # Regex search

Adaptive Scraping

Scrapling's signature feature. Elements are fingerprinted to SQLite and relocated by similarity scoring after site redesigns.

from scrapling.fetchers import Fetcher

Fetcher.adaptive = True
page = Fetcher.get('https://example.com')

# First run: save element fingerprint
products = page.css('.product-list', auto_save=True)

# Later, after site redesign breaks the selector:
products = page.css('.product-list', adaptive=True)  # Still finds it

Read: references/adaptive-scraping.md

Spider Framework

For full-site crawling with concurrency, pause/resume, and session routing:

from scrapling.spiders import Spider, Response, Request
from scrapling.fetchers import FetcherSession, StealthySession

class ResearchSpider(Spider):
    name = "research"
    start_urls = ["https://example.com/"]
    concurrent_requests = 10

    def configure_sessions(self, manager):
        manager.add("fast", FetcherSession(impersonate="chrome"))
        manager.add("stealth", StealthySession(headless=True), lazy=True)

    async def parse(self, response: Response):
        for link in response.css('a::attr(href)').getall():
            if "protected" in link:
                yield Request(link, sid="stealth")
            else:
                yield Request(link, sid="fast")
        for item in response.css('.article'):
            yield {"title": item.css('h2::text').get(), "url": response.url}

result = ResearchSpider().start(crawldir="./data")  # Pause/resume enabled
result.items.to_json("output.json")

Troubleshooting

  • Browser not found: Run scrapling install to download browser dependencies
  • Import error on fetchers: Install the full package: uv pip install "scrapling[all]"
  • Cloudflare still blocking: Combine flags: --solve-cloudflare --block-webrtc --hide-canvas
  • SSL cert error with HTTP fetcher: curl_cffi's bundled CA certs can be stale on macOS/pyenv. The wrapper handles this automatically (verify=False). For Python API, pass verify=False to Fetcher.get().
  • Slow stealthy fetch: Expected--browser automation is inherently slower than HTTP requests. Use get (HTTP) when stealth is not needed.

Reference Documentation

FileContents
references/cli-reference.mdFull CLI extract command reference (get, post, fetch, stealthy-fetch)
references/python-api-reference.mdPython API (Fetcher, DynamicFetcher, StealthyFetcher, Sessions, Spiders)
references/adaptive-scraping.mdAdaptive element tracking deep-dive (save, match, similarity scoring)

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Search academic papers, build literature reviews, and synthesize research findings — combines Exa MCP (research_paper category, arxiv filtering) with arxiv-mcp-server for paper discovery, download, and deep analysis. Triggers on academic paper, literature review, research synthesis, arxiv, find papers, scholarly search.

日本語の概要は準備中です。原文の説明を表示しています。

tdimino/claude-code-minoan412026年9月28日 更新

Automates browser interactions for web testing, form filling, screenshots, and data extraction. Use when the user needs to navigate websites, interact with web pages, fill forms, take screenshots, test web applications, or extract information from web pages.

日本語の概要は準備中です。原文の説明を表示しています。

tdimino/claude-code-minoan412026年9月28日 更新

This skill should be used when creating, auditing, or maintaining AGENTS.md files and Codex CLI configuration for any project—including initializing AGENTS.md for cross-agent compatibility (Codex, Cursor, Copilot, Devin, Jules, Amp, Gemini CLI), generating config.toml or .rules files, scaffolding .agents/skills/, converting CLAUDE.md to AGENTS.md, or auditing existing agent configs for bloat and staleness. Complementary to codex-orchestrator (which executes subagents; this skill creates the config files they consume).

日本語の概要は準備中です。原文の説明を表示しています。

tdimino/claude-code-minoan412026年9月28日 更新

Academic research skill for Biblical Hebrew, Semitic linguistics, cuneiform studies, and comparative Ancient Near Eastern research. Provides Sefaria API for Hebrew Bible, CDLI/ORACC for cuneiform databases, and web discovery via Omnisearch, Exa, Firecrawl, and Obscura for finding scholarly sources across JSTOR, Perseus, Persée, Google Scholar, and academia.edu. Triggers on Hebrew quotes, cuneiform, Sefaria, ANE research, Minoan, search for scholarship, find papers, literature review, scholarly search, academic search, Genesis/Tehom, Ugaritic, Talmudic sources, extract from PDF, OCR academic.

日本語の概要は準備中です。原文の説明を表示しています。

tdimino/claude-code-minoan412026年9月28日 更新

Build comprehensive ARCHITECTURE.md files following matklad's canonical guidelines — bird's-eye views, ASCII/Mermaid diagrams, codemaps, invariants, and layer boundaries. Triggers on document the architecture, create ARCHITECTURE.md, map this codebase, architectural overview.

日本語の概要は準備中です。原文の説明を表示しています。

tdimino/claude-code-minoan412026年9月28日 更新

Generate physically-based atmospheric scattering shaders — sky domes, planetary atmospheres, LUT-optimized pipelines, depth-aware post-processing. Four modes — sky-dome, atmosphere-post, planet, lut. Triggers on atmospheric scattering, sky shader, sunset rendering, planet atmosphere, Rayleigh scattering, Mie scattering, volumetric sky, sky dome, atmosphere post-processing, aerial perspective, transmittance LUT, sky rendering, realistic sky, planetary rendering.

日本語の概要は準備中です。原文の説明を表示しています。

tdimino/claude-code-minoan412026年9月28日 更新

tdimino のスキルをすべて見る

このスキルの問題を報告する