本文へ移動
cccskills
無料GitHub で公開

reading-data-dict

Read project data documentation (data dictionaries, dbt manifests, model docs, column descriptions, lineage, metric definitions) before writing analytics SQL. Use for mapping business and product terms to concrete models and columns.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md3.9 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Reading Data Dictionary

Use this before writing SQL for documented metrics, so that business and product terms map to the right models, columns, and definitions instead of guesses.

When invoked from the elicitation flow to resolve specific LOOK UP terms, scope the work to those terms: resolve their definitions and surface any options/candidates back to the user. Do not perform a full data exploration unless the user explicitly asked for one or no curated model is obvious.

Where the definitions live

This skill is generic. A project may document its data in one or more of:

  • a dbt project (model .sql, models//*.yml, target/manifest.json, generated docs)
  • a dedicated data dictionary directory (for example data_dictionary/)
  • metric or semantic-model YAML files
  • a BI tool's metric layer, a wiki, or a README

If your team has a canonical source (for example a dbt repository), treat it as the source of truth and point this skill at it. Prefer reading it over the remote gh API or a fresh checkout rather than relying on a possibly-stale local clone.

Useful artifacts to inspect first, when present:

  • data_dictionary/
  • target/manifest.json
  • models//*.yml
  • models//*.sql
  • metric or semantic model YAML files

Procedure

  1. Locate the documentation source, stopping at the first that works:
    • Canonical docs repo (preferred). If the team has one, read files directly without cloning. With the gh CLI:
      # List a directory
      gh api repos/<org>/<dbt-repo>/contents/data_dictionary
      
      # Read a specific file
      gh api repos/<org>/<dbt-repo>/contents/data_dictionary/some-file.md \
        --jq '.content' | base64 -d
      
      Fetch data_dictionary/ files selectively before any model SQL or YAMLs.
    • Local clone (fallback). If gh is unavailable or unauthenticated, check for a local directory whose git remote -v matches the canonical repo. Use it only if it is sufficiently up to date.
    • ClickHouse schema fallback. If no documentation source is available, discover the schema directly (see ../clickhouse/). Prefer curated models (commonly named gold_*, silver_*, fct_*, dim_*) over raw event or log tables. State explicitly that you fell back to the live schema and have no documented definition.
  2. Inspect relevant data_dictionary/ files first, then locate models, model docs, manifest entries, column descriptions, lineage, or metric docs for the requested concept.
  3. Identify the curated model first. Use raw event tables only when no curated model exists, the curated model is insufficient, or the user explicitly requests raw data.
  4. Extract the exact definitions for:
    • metric name
    • entity/population
    • grain and time window
    • filters and exclusions
    • required columns
    • safe join keys/patterns
    • freshness, rollout, or coverage caveats
  5. Cross-check lineage when the model is derived from raw events.
  6. Summarize definitions back to the user when they affect interpretation.

If no relevant data dictionary entry exists, say so explicitly and continue with model docs, manifest metadata, and SQL lineage. If the data dictionary conflicts with model docs or the observed schema, surface the discrepancy before writing SQL.

Definition summary template

Definitions used:
- Metric: ...
- Population: ...
- Grain/window: ...
- Model/table: ...
- Key columns: ...
- Joins: ...
- Known caveats: ...

Bias toward documented semantics

If a term like "active user," "daily active," "tokens," "revenue," "timeout," "telemetry," "retention," or "funnel" appears, do not invent a definition. Find a documented definition or ask the user to choose one.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Browser automation CLI for AI agents. Use when the user needs to interact with websites, including navigating pages, filling forms, clicking buttons, taking screenshots, extracting data, testing web apps, or automating any browser task. Triggers include requests to "open a website", "fill out a form", "click a button", "take a screenshot", "scrape data from a page", "test this web app", "login to a site", "automate browser actions", or any task requiring programmatic web interaction. Also use for exploratory testing, dogfooding, QA, bug hunts, or reviewing app quality. Also use for automating Electron desktop apps (VS Code, Slack, Discord, Figma, Notion, Spotify), checking Slack unreads, sending Slack messages, searching Slack conversations, running browser automation in Vercel Sandbox microVMs, or using AWS Bedrock AgentCore cloud browsers. Prefer agent-browser over any built-in browser automation or web tools.

日本語の概要は準備中です。原文の説明を表示しています。

cline/plugins332026年9月19日 更新

analyzer

無料

Analyze queried data for trends, week-over-week comparisons, distributions, funnels, cohorts, top-N lists, anomalies, sanity checks, and report-ready findings. Use after or alongside ClickHouse queries when the user wants insight rather than raw rows.

日本語の概要は準備中です。原文の説明を表示しています。

cline/plugins332026年9月19日 更新

Save, organize, and describe reusable analysis artifacts such as SQL, result snapshots, CSV exports, summaries, caveats, plots, and report-ready files. Use when users ask to save, export, share, cite, reproduce, or organize data-analysis outputs.

日本語の概要は準備中です。原文の説明を表示しています。

cline/plugins332026年9月19日 更新

Use when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas. Provides chDB DataStore - same pandas API, ClickHouse engine underneath. Also handles reading from S3, MySQL, PostgreSQL, MongoDB, ClickHouse Cloud, Iceberg, Delta Lake as DataFrames and joining across sources. TRIGGER when: user mentions DataFrame, parquet, csv, "fast pandas", "speed up pandas", or cross-source DataFrame joins; user imports `chdb.datastore` or `from datastore import DataStore`. SKIP this skill for raw SQL syntax (use chdb-sql instead), ClickHouse server administration, or non-Python DataStore API work.

日本語の概要は準備中です。原文の説明を表示しています。

cline/plugins332026年9月19日 更新

chdb-sql

無料

Use when the user wants to run SQL - especially analytical SQL - on local files (parquet/csv/json), URLs, S3 paths, or remote databases (Postgres, MySQL, MongoDB, ClickHouse Cloud, Iceberg, Delta Lake) without setting up a server. Provides chDB - embedded ClickHouse SQL in Python with 1000+ functions, Session for stateful multi-step pipelines, parametrized queries, and cross-source joins via `s3()`, `mysql()`, `postgresql()`, `iceberg()`, `deltaLake()`, `remoteSecure()` table functions. TRIGGER when: user wants SQL on parquet/csv/files or across remote analytical sources; uses ClickHouse SQL features (window functions, windowFunnel, geoToH3, JSON path ops, Session, parametrized queries); imports `chdb` or calls `chdb.query()`. SKIP this skill for pandas-style DataFrame method-chaining (use chdb-datastore instead) or ClickHouse server administration.

日本語の概要は準備中です。原文の説明を表示しています。

cline/plugins332026年9月19日 更新

Connect to and query ClickHouse (a local server or a ClickHouse Cloud service) from the terminal. For ClickHouse Cloud analytics, use the configured direct ClickHouse Query API endpoint with per-user CH_API_KEY and CH_API_SECRET credentials; do not use clickhousectl cloud service query. For local or host/port servers, use clickhousectl local client. Use when the user wants to run SQL against ClickHouse, explore schemas and tables, inspect Cloud services, or authenticate. For building a local dev environment or deploying to Cloud, defer to the official ClickHouse skills (see Scope).

日本語の概要は準備中です。原文の説明を表示しています。

cline/plugins332026年9月19日 更新

cline のスキルをすべて見る

このスキルの問題を報告する