本文へ移動
cccskills
無料GitHub で公開

data-analysis

End-to-end R data analysis pipeline — exploration → cleaning → regression → publication-ready tables and figures. Use when user says "analyze this dataset", "run a regression on X", "explore this CSV", "full analysis workflow", "get me summary stats and a regression", or points at a `.csv`/`.rds`/`.dta` and asks for empirical results. Produces numbered R scripts in `scripts/R/` and outputs to `output/`.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md8.9 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Data Analysis Workflow

Run an end-to-end data analysis in R: load, explore, analyze, and produce publication-ready output.

Input: $ARGUMENTS — a dataset path (e.g., data/county_panel.csv) or a description of the analysis goal (e.g., "regress wages on education with state fixed effects using CPS data").


Constraints

  • Follow R code conventions in .claude/rules/r-code-conventions.md
  • Save all scripts to scripts/R/ with descriptive names
  • Save all outputs (figures, tables, RDS) to output/
  • Use saveRDS() for every computed object — Quarto slides may need them
  • Use project theme for all figures (check for custom theme in .claude/rules/)
  • Run r-reviewer on the generated script before presenting results

Workflow Phases

Phase 0: Pre-Flight Report

Before writing any analysis code, produce a Pre-Flight Report showing you read the inputs. This prevents the common failure mode where the agent hallucinates variable names or skips project conventions.

Output block (in your response to the user, before Phase 1):

## Pre-Flight Report

**Dataset:** [path]
- Variables found: [list from head()/names()]
- Rows: [count]
- Key types: [e.g., "outcome=numeric, treatment=binary, state=factor"]
- Missing-data summary: [% missing per key var]

**Project conventions read:**
- `.claude/rules/r-code-conventions.md` — [one-line summary of most relevant rule]
- `.claude/rules/content-invariants.md` — [INV-9, INV-10, INV-11, INV-12 applicable]

**Task interpretation:** [one sentence restating what the user asked for]

**Plan:** [3-5 bullet outline of the R script structure]

If any input cannot be read (missing file, unreadable format), stop and ask the user before proceeding.

Phase 1: Setup and Data Loading

  1. Create R script with proper header (title, author, purpose, inputs, outputs)
  2. Load required packages at top (library(), never require())
  3. Set seed once at top in YYYYMMDD format (per r-code-conventions.md), e.g. set.seed(20260415) (INV-9)
  4. Load and inspect the dataset

Phase 2: Exploratory Data Analysis

Generate diagnostic outputs:

  • Summary statistics: summary(), missingness rates, variable types
  • Distributions: Histograms for key continuous variables
  • Relationships: Scatter plots, correlation matrices
  • Time patterns: If panel data, plot trends over time
  • Group comparisons: If treatment/control, compare pre-treatment means

Save all diagnostic figures to output/diagnostics/.

Phase 3: Main Analysis

Based on the research question:

  • Regression analysis: Use fixest for panel data, lm/glm for cross-section
  • Standard errors: Cluster at the appropriate level (document why)
  • Multiple specifications: Start simple, progressively add controls
  • Effect sizes: Report standardized effects alongside raw coefficients

Specification ledger — every run, kept or not

Every specification estimated in this phase, including the ones dropped and the ones that failed, gets one row appended to quality_reports/spec-ledger.md as it is run. It is the record a referee's "what else did you try?" deserves: it keeps the search visible rather than preventing it, and it records specifications without advising which to run.

  • Append only. Never edit or delete a row; a correction is a new row. A commit that changes a committed row fails the repo-hygiene gate.
  • Status is kept, dropped or failed; Why says why for anything not kept.
  • Estimate is optional. On restricted data, leave it empty until the number has cleared disclosure review (confidential-data.md).
  • Escape a | inside a formula as \|, or the table breaks. A specification containing a backtick (a Stata local macro such as `x') goes in a double-backtick span: `` reg y `controls' ``.
LEDGER=quality_reports/spec-ledger.md
[ -f "$LEDGER" ] || printf '%s\n' "# Specification ledger" "" \
  "Every specification estimated, kept or not, one row each, appended as it is run. Never edit a past row; a correction is a new row." "" \
  "| Date | Commit | Script:line | Outcome | Specification | Sample | Status | Why | Estimate |" \
  "|---|---|---|---|---|---|---|---|---|" > "$LEDGER"
if REV=$(git rev-parse --short HEAD 2>/dev/null); then       # -dirty = anything uncommitted outside quality_reports/, untracked files included
  [ -n "$(git status --porcelain -- . ':(exclude)quality_reports' 2>/dev/null)" ] && REV="$REV-dirty"
else
  REV="no-commit"
fi
printf '| %s | %s ' "$(date +%F)" "$REV" >> "$LEDGER"          # one printf + one heredoc per specification
cat >> "$LEDGER" <<'EOF'
| scripts/R/03_analyze.R:42 | log_wage | `feols(log_wage ~ treat + age \| id + year, cluster = ~id)` | panel 2010-2019 | kept | main specification | 0.082 (0.021) |
EOF

Phase 4: Publication-Ready Output

Tables:

  • Use modelsummary for regression tables (preferred) or stargazer
  • Include all standard elements: coefficients, SEs, significance stars, N, R-squared
  • Export as .tex for LaTeX inclusion and .html for quick viewing

Figures:

  • Use ggplot2 with project theme
  • Set bg = "transparent" for Beamer compatibility
  • Include proper axis labels (sentence case, units)
  • Export with explicit dimensions: ggsave(width = X, height = Y)
  • Save as both .pdf and .png

Phase 5: Save and Review

  1. saveRDS() for all key objects (regression results, summary tables, processed data)
  2. Create output/ subdirectories as needed with dir.create(..., recursive = TRUE)
  3. Run the r-reviewer agent on the generated script:
Delegate to the r-reviewer agent:
"Review the script at scripts/R/[script_name].R"
  1. Address any Critical or High issues from the review.

Script Structure

Follow this template:

# ============================================================
# [Descriptive Title]
# Author: [from project context]
# Purpose: [What this script does]
# Inputs: [Data files]
# Outputs: [Figures, tables, RDS files]
# ============================================================

# 0. Setup ----
library(tidyverse)
library(fixest)
library(modelsummary)

set.seed(20260415)  # YYYYMMDD per r-code-conventions.md (INV-9)

dir.create("output/analysis", recursive = TRUE, showWarnings = FALSE)

# 1. Data Loading ----
# [Load and clean data]

# 2. Exploratory Analysis ----
# [Summary stats, diagnostic plots]

# 3. Main Analysis ----
# [Regressions, estimation]

# 4. Tables and Figures ----
# [Publication-ready output]

# 5. Export ----
# [saveRDS for all objects, ggsave for all figures]

Important

  • Reproduce, don't guess. If the user specifies a regression, run exactly that.
  • Show your work. Print summary statistics before jumping to regression.
  • Check for issues. Look for multicollinearity, outliers, perfect prediction.
  • Use relative paths. All paths relative to repository root.
  • No hardcoded values. Use variables for sample restrictions, date ranges, etc.

Long-running fits: use the Monitor tool (Apr 2026)

For regressions, simulations, or bootstrap loops that take more than a couple of minutes, launch via Bash with run_in_background: true and then use Anthropic's Monitor tool to stream R stdout into the conversation in real time. Pattern:

  1. Background-launch with Bash run_in_background: true, sending all output to a log: mkdir -p output && Rscript scripts/R/03_analyze.R > output/03_analyze.log 2>&1. The background job notifies you by itself when the process exits.
  2. Start Monitor with a command that follows that log and filters for milestones and failures, e.g. tail -f output/03_analyze.log | grep --line-buffered -E "Coefficients table written|Error|Execution halted". Monitor has no job-id parameter: the stdout of its own command is the event stream. tail -f never exits, so set timeout_ms above the expected runtime (or persistent: true) and stop the monitor with TaskStop once the job finishes.
  3. Continue or course-correct based on what the stream reveals.

This avoids the polling-loop anti-pattern (sleep 30; check; sleep 30; check) and avoids burning cache on idle waits. Especially useful when paired with the Cost-Conscious Parallelism section of the guide.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Turn an incoming set of findings — from an AI reviewer, a referee report, a code review, a linter, or a second model — into verified fixes, without letting a confident misread damage correct work. Every finding is a CANDIDATE until checked against the actual source. Use whenever you receive review comments, audit findings, or a critique you did not write yourself, especially when the reviewer is a model or when the volume is too large to check by feel.

日本語の概要は準備中です。原文の説明を表示しています。

pedrohcgs/claude-code-my-workflow1,6592026年9月28日 更新

Enforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs. Report PASS/FAIL per claim against tolerance thresholds. Use before submission and before releasing a replication package.

日本語の概要は準備中です。原文の説明を表示しています。

pedrohcgs/claude-code-my-workflow1,6592026年9月28日 更新

Before and after changing anything shared — a function's return value, a signature, a schema, a label set, a config default, a constant, a file format — find every consumer and actually run them. Catches the change that looks purely additive but silently breaks a contract in a file you never opened. Use when editing shared code, adding a field/column/return element, renaming, changing units or defaults, or touching a pipeline that produces reported numbers.

日本語の概要は準備中です。原文の説明を表示しています。

pedrohcgs/claude-code-my-workflow1,6592026年9月28日 更新

Snapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt / environment.yml / uv.lock, Stata version + ado package list), records seeds and RNG kind, optionally writes a pinning Dockerfile, and produces a paste-ready "Computational requirements" block. Use when user says "capture the environment", "snapshot my dependencies", "pin the versions", "make a renv.lock / requirements.txt", "make this byte-reproducible", or before releasing a replication package to openICPSR / the AEA Data Editor.

日本語の概要は準備中です。原文の説明を表示しています。

pedrohcgs/claude-code-my-workflow1,6592026年9月28日 更新

challenge

無料

Stress-test a finding against the choices you did not make. Enumerates the discrete forks a competent analyst could have taken (measure definition, sample filter, control set, clustering level, weighting, functional form), runs the specification grid, and reports the distribution rather than a point estimate — then attacks the identifying assumption with named, computable sensitivity statistics. Use when the user says "is this robust", "challenge this result", "specification curve", "multiverse", "how sensitive is this", "what if I'd used a different measure", "stress-test my estimate", or before a result becomes a headline claim. NOT a reviewer of prose or code — it challenges the CLAIM.

日本語の概要は準備中です。原文の説明を表示しています。

pedrohcgs/claude-code-my-workflow1,6592026年9月28日 更新

Save a structured state snapshot before stopping or handing off. Captures the active plan, recent decisions, file pointers (with line numbers), open questions, and the next 1–3 actions into a checkpoint file under `quality_reports/checkpoints/`. Optionally proposes `[LEARN]` entries to add to MEMORY.md. Use when user says "checkpoint", "save state", "snapshot before I stop", "where am I", "wrap up the session for handoff", or before a long break / model switch / collaborator handoff. Companion to (NOT replacement for) the narrative session-log workflow.

日本語の概要は準備中です。原文の説明を表示しています。

pedrohcgs/claude-code-my-workflow1,6592026年9月28日 更新

pedrohcgs のスキルをすべて見る

このスキルの問題を報告する