本文へ移動
cccskills
無料GitHub で公開

vs-crawler

Crawl websites (news, blogs, papers, GitHub, product docs, RSS feeds) into a fixed-schema JSONL file, then create a dataset and a searchable application in Viking AI Search. Supports one-time crawl and scheduled recurring crawl with automatic incremental sync.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md15.5 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Viking Content Crawler

When to Use

Use this skill when the user wants to crawl content from websites and import it into Viking AI Search to build a searchable knowledge base. This covers news sites, blogs, academic papers, GitHub repositories, product documentation, RSS feeds, and similar web content sources.

The agent writes crawler code tailored to the target sites, outputs data in a fixed JSONL schema, and then hands off to the vs-item-onboarding skill for dataset creation and import.

Do not use this skill when:

  • The user already has a local file ready to import (use vs-item-onboarding directly).
  • The user wants to import from a database (use vs-item-onboarding directly with MySQL).

Version Check

Before starting this skill workflow, run vs version check --json. Continue only when status is up-to-date. If status is update-available, stop and tell the user to update the cloned vs repository, then run git pull --ff-only, bash ./scripts/install.sh, and bash ./scripts/install-skills.sh all --target auto --force (PowerShell: scripts/install.ps1 and scripts/install-skills.ps1). If the status is unknown, stop and report that the CLI version could not be verified.

Fixed Schema

All crawled records MUST conform to this schema. Every record is a flat JSON object written as one line in a JSONL file.

FieldTypeRequiredDescription
idstringyesUnique identifier. Use a source-native stable ID (e.g., arXiv ID, GitHub owner/repo, post slug) when available; otherwise derive a deterministic ID from title + author + published_at. Must be deterministic so re-crawling the same item produces the same ID.
titlestringyesContent title (headline, post title, paper title, repo name, doc page title).
summarystringyesShort abstract or description (100-500 characters recommended).
contentstringyesFull text body with HTML stripped to plain text. For GitHub repos, concatenate README content. For PDF/DOC documents, extract the text content directly into this field.
categorystringyesOne of: news, blog, paper, github, docs, other.
sourcestringyesHuman-readable source name, e.g. "Hacker News", "arXiv", "Viking Docs".
authorstringnoAuthor name(s); multiple authors separated by commas.
published_atstringnoISO 8601 datetime, e.g. "2026-07-16T10:30:00Z". Use crawl time if unavailable.
tagsarray<string>noTags, keywords, or topics.
languagestringnoISO 639-1 code: "en", "zh", etc.
source_urlstringnoCanonical URL of the source page (the URL the record was crawled from). Must be a fully-qualified URL with scheme and host.
metadataobjectnoStructured key-value data. Must be flat (one level deep, no nested objects). Values must be scalar (string, number, boolean) — no arrays or objects inside. Only the standard keys listed below are allowed; do not add custom keys. All sources must use the same metadata schema.

Example Record

{
  "id": "viking-blog-introducing-viking-ai-search",
  "title": "Introducing Viking AI Search",
  "summary": "Viking AI Search is a new generation of hybrid search engine combining BM25 and vector search...",
  "content": "Full article text with HTML removed and paragraphs separated by newlines...",
  "category": "blog",
  "source": "Viking Blog",
  "author": "Jane Doe",
  "published_at": "2026-07-15T08:00:00Z",
  "tags": ["search", "vector database", "hybrid search"],
  "language": "en",
  "source_url": "https://viking.example.com/blog/introducing-viking-ai-search",
  "metadata": {
    "read_time": "8 min",
    "word_count": 2340,
    "views": 12580
  }
}

Standard Metadata Fields

Only these keys are allowed in metadata. Do not add custom keys — every crawled record, regardless of source or category, must use exactly these keys when the data is available, and omit keys whose data is unavailable. This guarantees schema consistency across all crawl sources so downstream consumers (schema inference, search relevance tuning) see a uniform shape.

KeyTypeCategoryDescription
read_timestringcontentEstimated reading time, e.g. "8 min".
word_countnumbercontentWord count of the article / document body.
viewsnumberengagementView count or page view count.
likesnumberengagementLike / upvote / thumbs-up count.
commentsnumberengagementComment count.
sharesnumberengagementShare count.
starsnumberrepo / paperGitHub stars (for github category) or citation-equivalent metric.
forksnumberrepoGitHub fork count (for github category).
citationsnumberpaperCitation count (for paper category).
venuestringpaperPublication venue, e.g. "NeurIPS 2025", "arXiv".
doistringpaperDigital Object Identifier, e.g. "10.1234/abcde".

Values must be flat scalars (string / number / boolean). No nested objects, no arrays. If a data point does not map to any standard key, omit it rather than inventing a new key.

Preconditions

  • vs CLI >= 0.2.0 is installed and authenticated (vs auth status and vs doctor succeed).
  • The crawl target is reachable from the execution environment.
  • A suitable runtime is available (Python 3.8+ with requests and beautifulsoup4 recommended).

Commands

This skill delegates dataset creation and import to vs-item-onboarding. The crawler workflow itself uses:

StageActionPurpose
CrawlRun agent-written crawler scriptFetch content and write JSONL
OnboardInvoke vs-item-onboarding skillCreate dataset, infer schema, import data, optionally start sync
ScheduleSet up cron/launchd wrapperFor scheduled mode: periodically re-crawl and append new lines

Workflow

Run in strict order.

  1. Confirm crawl mode — resolve whether the user wants one-time crawl or scheduled recurring crawl. Only skip the question when the request contains an explicit, unambiguous signal (apply detection to whatever language the user is writing in):

    • Explicit one-time: phrases carrying "once", "one-time", "just this time", or equivalent single-crawl semantics.
    • Explicit scheduled: phrases carrying "daily", "scheduled", "keep updated", "auto-crawl", "sync", "incremental", or equivalent recurring semantics.
    • If the request is neutral — e.g. "crawl X", bare "crawl", mentions target sites but says nothing about scheduling/once — you MUST ask the user to choose. The bare crawl verb is NOT a one-time signal; it is ambiguous. Never silently default to one-time.
  2. Identify crawl targets and write the crawler. Based on the user's target sites, write a crawler script. The crawler MUST:

    • Output records conforming to the Fixed Schema as JSONL (one record per line).
    • Write output to a stable path: /tmp/viking/crawler/<job-name>/items.jsonl.
    • For scheduled mode: support incremental crawling — track the last crawl cursor (most recent published_at or last seen item IDs) in /tmp/viking/crawler/<job-name>/state.json so subsequent runs only fetch new content.
    • Deduplicate by id within each run and against previous state.
    • Strip HTML to plain text; never include raw HTML in content.
    • When encountering PDF, DOC, or other document links, download the document and extract its text content directly into the content field. Use available libraries (e.g. PyPDF2/pypdf for PDF, python-docx for DOCX, beautifulsoup4 for HTML) to extract readable text. Do not store document links in records; put the extracted full text in content.
    • Be polite: set a descriptive User-Agent, respect robots.txt, add 1-3 second delays between requests, retry transient errors with backoff.
    • Prefer structured sources (RSS/Atom feeds > sitemap.xml > official APIs > HTML scraping).
    • Log per-item errors and continue; do not abort on single-page failures.
    • Strictly follow the Fixed Schema defined above — the same field names, types, and metadata key set, regardless of the source. Do not add source-specific top-level fields or metadata keys. Print a summary to stdout: crawled count, new count, output path.
  3. Run the crawler to produce the initial JSONL file at /tmp/viking/crawler/<job-name>/items.jsonl.

  4. Hand off to vs-item-onboarding. Invoke the vs-item-onboarding skill with the following context:

    • Source type: JSONL file
    • File path: /tmp/viking/crawler/<job-name>/items.jsonl
    • Import mode: one-time import if the user chose one-time crawl; one-time import + ongoing incremental sync if the user chose scheduled crawl.
    • App creation: required — the user wants both a dataset AND an application so the crawled content is immediately searchable. Tell vs-item-onboarding to run through app creation and dataset attachment (steps 12–13) rather than stopping after dataset creation.
    • Schema confirmation: auto-confirm — the crawler outputs a fixed, well-defined schema (see Fixed Schema above). When vs-item-onboarding reaches the Schema Confirmation step (step 7), automatically reply yes to proceed without surfacing the confirmation prompt to the user. Only surface it if the backend returns warnings that indicate actual schema problems (e.g. missing PK BizAttr).
    • Readiness: do NOT block waiting for Ready. After vs-item-onboarding prints its hand-off block with console links, the workflow is complete. Do not run vs app wait-ready, do not poll for readiness, do not add any extra waiting steps. The user will check the console themselves.
    • Let vs-item-onboarding handle all subsequent steps (schema inference, confirmation, dataset creation, data write, app creation, dataset attach, optional sync start, console hand-off).
    • Do NOT re-implement the onboarding steps yourself — defer entirely to vs-item-onboarding.
  5. (Scheduled mode only) Set up recurring crawl + sync. After vs-item-onboarding completes successfully and the dataset is created:

    • The JSONL connector sync (set up by vs-item-onboarding during step 4) already watches the JSONL file for new lines and imports them automatically. You do NOT need to separately configure vs connector init/run for the file — vs-item-onboarding handles this when it chooses the sync path.
    • Create a wrapper script that:
      1. Runs the crawler in incremental mode (using state.json to skip already-crawled content), appending new records to /tmp/viking/crawler/<job-name>/items.jsonl.
      2. Exits cleanly if no new records are found.
    • Schedule the wrapper script using the platform-appropriate mechanism:
      • cron (macOS/Linux): add a crontab entry. Recommended interval: 30 minutes to a few hours depending on how frequently the source updates.
      • launchd (macOS): create a LaunchAgent plist with StartInterval.
    • Surface the schedule info, log file path, and how to stop/inspect the job in the hand-off.

Customer Environment Principle

  • In customer environments, assume repository source code is unavailable.
  • Execute tasks using only the installed skills, the packaged vs CLI surface (--help, command output, observed runtime behavior), and explicit user-provided information.
  • If the installed CLI behavior conflicts with a skill, trust the installed CLI behavior first.

Constraints

  1. Never write raw HTML into content. Always strip to plain text.
  2. Never hardcode credentials in crawler code. Use environment variables for API keys.
  3. Always generate a stable id. Use a source-native stable ID (e.g., arXiv ID, GitHub owner/repo, post slug) when available; otherwise derive a deterministic ID from title + author + published_at. Must be deterministic so re-crawling the same item produces the same ID.
  4. All datetime values MUST be ISO 8601 (e.g., "2026-07-16T10:30:00Z").
  5. All output MUST be valid JSONL: one JSON object per line, UTF-8 encoded.
  6. The category field MUST use the predefined values (news, blog, paper, github, docs, other).
  7. Dataset creation and import MUST go through vs-item-onboarding. Do not call vs dataset create, vs data write, etc. directly from this skill.
  8. For scheduled mode, incremental sync is handled by the JSONL file connector (configured by vs-item-onboarding). The scheduled job only needs to run the crawler to append new lines to the JSONL file; the connector daemon picks up new lines automatically.
  9. Respect rate limits and robots.txt. Add polite delays between requests.
  10. Extract text from PDF/DOC documents. When encountering PDF, DOCX, or other document links, download the file and extract its text content directly into the content field using appropriate libraries (e.g., pypdf for PDF, python-docx for DOCX). Do not store document links in output records.
  11. Auto-confirm Schema Confirmation during onboarding. The crawler produces records against the Fixed Schema defined above, which is stable and well-defined. When handing off to vs-item-onboarding, instruct it to automatically reply yes at the Schema Confirmation step without surfacing the prompt to the user. Only pause and surface schema details if the backend inference returns genuine errors (e.g. missing primary-key BizAttr) that require user intervention.
  12. Never block waiting for dataset/app readiness. After vs-item-onboarding completes its hand-off (printing console links + readiness reminder), end your turn. Do NOT run vs app wait-ready, vs dataset wait-ready, or any polling loop to wait for the Ready state. Readiness is an asynchronous backend process; tell the user to check the console links themselves.
  13. metadata keys are fixed — only standard keys allowed. All records from all sources must use only the standard metadata keys listed in the Standard Metadata Fields table. Never invent custom keys. If a data point does not fit any standard key, omit it. This guarantees uniform schema across all crawl sources.
  14. Before executing any concrete vs ... command, first consult vs-product-qa to verify the current command surface and required flags.

Recovery Hints

  • Crawler returns zero records → verify target site/feed accessibility, check for rate limiting (HTTP 429), review error logs.
  • Duplicate records appear → verify id generation is deterministic (same item always produces the same ID).
  • Content extraction produces garbled text → ensure HTTP response encoding is correctly detected.
  • PDF text extraction fails or is garbled → try a different PDF library (e.g., switch from pypdf to pdfplumber) or fall back to extracting abstract/metadata only.
  • Sync is not picking up new lines → verify the JSONL connector daemon is running via vs connector status --job <job>.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Volcengine official documentation lookup helper. Supports both document search and full-content fetch across Volcengine products, developer tools, support content, best practices, pricing, deployment, troubleshooting, API, SDK, and policy pages.

日本語の概要は準備中です。原文の説明を表示しています。

volcengine/SearchCLI1,1932026年9月19日 更新

Provide system alias mapping for Search CLI. Invoke this skill when user mentions "Search CLI", "search_cli", or tries to execute search_cli commands.

日本語の概要は準備中です。原文の説明を表示しています。

volcengine/SearchCLI1,1932026年9月19日 更新

vs-chat

無料

Conversational search runtime: send messages, keep sessions consistent, and verify retrieval behavior and responses.

日本語の概要は準備中です。原文の説明を表示しています。

volcengine/SearchCLI1,1932026年9月19日 更新

onboarding workflow for creating datasets and applications in Viking AI Search. Supports one-time import from local files (JSON, JSONL, CSV) and MySQL databases, plus scheduled incremental sync for append-only JSONL files and MySQL. All sources are first exported to a bootstrap JSONL file; backend-driven schema inference handles detection, and optional background sync keeps the dataset up to date as new data arrives.

日本語の概要は準備中です。原文の説明を表示しています。

volcengine/SearchCLI1,1932026年9月19日 更新

Answer Viking AI Search product questions, CLI usage questions, API/auth questions, configuration questions, and troubleshooting questions by grounding every claim in either the installed `vs` CLI's own output or official Volcengine documentation. Never fabricate.

日本語の概要は準備中です。原文の説明を表示しています。

volcengine/SearchCLI1,1932026年9月19日 更新

Create Viking web projects, start and verify a local preview, or deploy a generated project to Volcengine IGA Pages when explicitly requested. Includes agent-guided feature, eligible application, dataset, scene, and authentication choices. Use only after confirming the installed CLI exposes `vs project`; otherwise stop without taking action.

日本語の概要は準備中です。原文の説明を表示しています。

volcengine/SearchCLI1,1932026年9月19日 更新

volcengine のスキルをすべて見る

このスキルの問題を報告する