本文へ移動
cccskills
無料GitHub で公開

dataset

公开数据集发现入口。覆盖 Kaggle / UCI / HuggingFace / 天池。题目要求"自行 查找数据"或"补充外部数据"时启用;附件已含数据时**不调用**。

インストール方法を見る

含まれるファイル(1)

  • SKILL.md2.1 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

dataset — 公开数据集发现

何时使用

  • 题目附件没有数据集,但题面要求"查找类似公开数据"。
  • 需要历史基准数据集(如 MNIST / Iris / 波士顿房价)做模型对比。
  • 不要在已有附件数据时启用本子 skill。

入口

数据源入口配置
Kagglekaggle datasets list -s <kw> / kaggle datasets download -d <user/name>~/.kaggle/kaggle.json(注册免费下载)
UCI ML直接 HTTPS pd.read_csv(url)无
HuggingFacefrom datasets import load_datasetEZMM_HF_TOKEN(私有数据集)
天池浏览器 + 手动下载无

命令模板

# Kaggle 搜索 + 下载
kaggle datasets list -s "vegetable retail price"
kaggle datasets download -d <user/dataset-name> -p workdir/.../attachments/external/kaggle --unzip

# HuggingFace
python -c "from datasets import load_dataset; ds = load_dataset('squad', split='train[:1%]')"

# UCI(直接 URL)
python -c "import pandas as pd; df = pd.read_csv('https://archive.ics.uci.edu/...'); df.to_csv('workdir/.../attachments/external/uci/iris.csv', index=False)"

落盘规范

外部下载的数据放在:

workdir/{task_id}/attachments/external/<source>/<dataset>/

并在同目录写 SOURCES.md:

- 数据集名: <name>
- 来源: <kaggle url>
- License: <CC0 / CC-BY / Apache-2.0 / 其他>
- 下载时间: <ISO timestamp>
- 用途说明: <一句话>

失败诊断

情况处理
Kaggle token 未配置提示用户 ~/.kaggle/kaggle.json 配置
License 不允许商用数据可用于学术建模报告;论文中标明出处 + license
文件 > 1GBcoder 阶段用 chunksize 处理(参考 prompts/coder.md)
网络受限写诊断;建议用户手动下载放入 attachments/

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Pre-pipeline aggregator that scans AI agent cache directories (.claude, .cursor, .antigravity, .openclaw) or any user-specified directory for experimentation logs, extracts insights and numeric results, and formats them as PaperOrchestra-ready inputs (idea.md + experimental_log.md). TRIGGER when the user says "aggregate my agent logs for paper writing", "extract experiments from my coding agent history", "prepare PaperOrchestra inputs from my cache", "turn my agent logs into a paper", mentions a folder or directory they want to use as the basis for a paper, or wants to run PaperOrchestra but only has scattered agent experiment histories rather than structured inputs. Run this BEFORE paper-orchestra. Also called automatically by paper-orchestra when workspace/inputs/idea.md or workspace/inputs/experimental_log.md are missing.

日本語の概要は準備中です。原文の説明を表示しています。

woodfishhhh/EZ_math_model422026年7月28日 更新

Use when EZ_math_model model selection is unclear after the modeling decision tree, the problem spans multiple domains, or modeler needs several candidate approaches before writing modeling_plan.md.

日本語の概要は準備中です。原文の説明を表示しています。

woodfishhhh/EZ_math_model422026年7月28日 更新

Step 5 of the PaperOrchestra pipeline (arXiv:2604.05018). Iteratively refine drafts/paper.tex by simulating peer review and applying targeted revisions, with strict accept/revert halt rules. Maintains a worklog and snapshots each iteration so revert is real, not symbolic. TRIGGER when the orchestrator delegates Step 5 or when the user asks to "refine the draft", "iterate on the paper", or "run peer review on this paper".

日本語の概要は準備中です。原文の説明を表示しています。

woodfishhhh/EZ_math_model422026年7月28日 更新

Use when EZ_math_model runs in multi or hybrid mode and has two or more independent subtasks that can be assigned to subagents without shared files, shared state, or dataflow dependencies.

日本語の概要は準備中です。原文の説明を表示しています。

woodfishhhh/EZ_math_model422026年7月28日 更新

docx

無料

Use when EZ_math_model needs to convert paper.md to paper.docx, read a DOCX problem statement, preserve equations and tables, or package a modeling report as a Word document.

日本語の概要は準備中です。原文の説明を表示しています。

woodfishhhh/EZ_math_model422026年7月28日 更新

Use when EZ_math_model multi or hybrid mode needs several domain-specialist subagents to gather outside knowledge, industry context, model theory, or sensitivity-analysis references in parallel.

日本語の概要は準備中です。原文の説明を表示しています。

woodfishhhh/EZ_math_model422026年7月28日 更新

woodfishhhh のスキルをすべて見る

このスキルの問題を報告する