本文へ移動
cccskills
無料GitHub で公開

adjudication-sheets

Build human adjudication / hand-labeling sheets from LLM-pipeline data without evidence truncation. Use when: (1) preparing a CSV/Excel sheet for a human to rule on cases an LLM classifier or rater panel judged, (2) a labeler reports "there is no information to label from" or cells look empty in Excel, (3) excerpt columns cluster at one exact length (e.g. all 1,500 chars — a hard truncation cap). Covers: full rating-basis recovery, Excel 32,767-char cell cap, multi-line CSV mangling, ruling dropdowns, companion text files.

インストール方法を見る

含まれるファイル(2)

  • SKILL.md4.5 KB
  • README.md3.9 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Human Adjudication Sheets from LLM-Pipeline Data

Problem

Adjudication sheets built from pipeline intermediates tend to carry truncated "excerpt" columns (previews made for machine diffing, not human judging). A human asked to rule "does this text show X?" on a fragment produces invalid rulings: the decisive sentence may sit past the cutoff. Separately, multi-line text in CSV cells renders as broken rows/empty cells in Excel, so the labeler reports "there is no data here" even when the column is populated.

Context / Trigger Conditions

  • A labeler says the sheet has "no information" — check for embedded newlines in CSV fields first.
  • Excerpt lengths cluster at an exact value (all 1,204 / 1,500 / 800 chars) or end with "[…]" — that is a hard cap, not natural length.
  • The ruling question is "does the record/text show X?" — the ruling basis must be byte-identical to what the classifier/raters saw, or the ruling grades a different object.

Solution

  1. Trace the true rating basis from the scoring code, not from intermediates. Find the line where text is passed to the model (e.g. prompt = ... + text) and reproduce that construction exactly. Verify: stored excerpt == full_text[:cap] (0 mismatches), or cache-key match against the LLM request cache.
  2. Deliver .xlsx, not .csv, for any sheet with multi-line text: wrap_text + frozen header + generous width on evidence columns; data-validation dropdown on the ruling column with the exact allowed codes (e.g. E_pos/E_neg/unclear, or 0/1/2/unclear).
  3. Handle the Excel cell cap (32,767 chars) explicitly: if full text exceeds it, inline the first ~30k ending with a loud marker ("[CONTINUES — see companion file]") AND write the complete text to a per-case .txt (header block with case metadata + full text); add a text_file path column. Write companion files for ALL cases regardless — they are the comfortable reading surface.
  4. Keep the original CSV untouched as the machine-readable artifact; the xlsx is the human ruling surface; harvest rulings from the xlsx. Distribution: text_file paths are project-relative and break the moment the sheet is emailed or copied out of the tree (the labeler asks "where are the attendant/companion files?"). If the labeler is not working inside the project folder, ship ONE zip containing the xlsx and the companion-texts folder side by side, with the instruction: extract together; the file for row X is <texts_folder>\<case_id>.txt next to the spreadsheet. Require the FILLED xlsx back and harvest from the returned copy, not the original.
  5. Tell the labeler the epistemics: rule from the provided record only; silence = negative; outside research (Google) goes in ruling_notes, never the ruling — otherwise the labels leak information the classifier could never see.

Verification

  • No excerpt column has >30% of rows at one exact length; no "[…]" markers remain.
  • pd.read_excel round-trip shows full lengths (compare min/median/max vs the old excerpts).
  • Dropdown rejects free-text entries; a saved test ruling survives reopen.

Example

Specialist Directors US, 2026-07-08: director sheet excerpts capped at ~1,200 chars (full dossiers up to 4,090); filing sheet capped at 1,500 chars while raters had scored the whole ~80k-char Item 1A — the PI caught both mid-sitting. Fix: src/director_v1/make_sitting_xlsx.py, tools/make_adjudication_v2_xlsx.py (full-text xlsx + 76 companion txts + dropdowns).

Notes

  • Uniform-length clustering is the fastest tell; check it BEFORE handing any sheet to a human.
  • Harvest gotcha: when comparing harvested rulings to prior labels, coerce BOTH sides to numeric first — Excel/pandas round-trips floats as "1.0" vs the dropdown's "1", so a string compare falsely flags every row as a disagreement.
  • If the pipeline caches LLM requests content-addressed, verify the rebuilt basis against the cache rather than trusting a deterministic builder to have been stable.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

N-round adversarial review pipeline for empirical research output — the chain from data to LaTeX tables to a manuscript that cites them. A Claude drafter proposes minimal diffs, a deterministic mechanical battery gates every diff from a clean state with a regression gate, a Codex reviewer files check-backed critiques, and a blind judge panel decides residual disputes. Manual-invoke ONLY: trigger when the user explicitly runs /adversarial-empirical-review or names 'adversarial-empirical-review' / 'adversarial empirical review'. Do NOT auto-trigger on generic 'review my results', 'check my tables', or manuscript-editing requests. For prose-style refinement use style-emulation instead; this skill AUDITS WHETHER THE TABLES ARE CORRECT — that each number in the tables is what the analysis code computes, reproduces from the data, and is internally consistent. It is an empirical + code review: the manuscript is read only to resolve table numbering, and prose is not examined.

日本語の概要は準備中です。原文の説明を表示しています。

kennethkhoocy/applied-micro-skills222026年9月5日 更新

Before designing, training, or auditing ANY model that replicates human-annotated labels, audit the annotation protocol's INPUT — the exact document/evidence the human labelers consulted — and give the model that same input. Use when: (1) designing a classifier/LLM extractor whose target is a hand-coded label set, (2) a label-replication model shows low recall concentrated in a label subset and the diagnosis on offer is "the label's information is not in the features", (3) reviewers propose construct splits (e.g. "designation vs record-evident"), adjudication sittings, or per-domain stop rules to explain residual disagreement with gold, (4) validating an extraction pipeline against labels transcribed from a source document. Symptom of the underlying failure: elaborate theory accumulates to explain why gold is "partially unpredictable" when the model was simply never shown the document the annotators read.

日本語の概要は準備中です。原文の説明を表示しています。

kennethkhoocy/applied-micro-skills222026年9月5日 更新

Raise real concurrency in asyncio LLM batch scorers built on the OpenAI SDK (AsyncOpenAI, including OpenAI-compatible providers like DeepSeek). Use when: (1) raising an asyncio.Semaphore above ~100 produces no throughput gain, (2) a batch pipeline saturates near 100 in-flight requests despite a larger semaphore, (3) planning a high-concurrency campaign against a provider with no hard rate limit (DeepSeek v4-flash tolerates 2000+ in flight). Root cause: AsyncOpenAI's default httpx pool caps max_connections at 100, silently bottlenecking any larger semaphore — you must pass a custom http_client with httpx.Limits sized to the semaphore.

日本語の概要は準備中です。原文の説明を表示しています。

kennethkhoocy/applied-micro-skills222026年9月5日 更新

Place pre-screened literature citations into a LaTeX or Word manuscript, or restyle the citations already in one. Three modes: (1) inline placement — inline \cite{}/\citet{}/\citep{} with a compiled references.bib, for author-date journals (APA, MLA, Harvard, Chicago author-date, IEEE, Vancouver); (2) footnote placement — full formatted \footnote{} or OOXML footnotes for legal and notes styles (Bluebook, OSCOLA, Chicago, APA, McGill) with Id./supra short forms; (3) restyle — convert existing footnote citations from one style to another. This skill is manual-invoke ONLY — trigger ONLY when the user explicitly runs /cite-placement or explicitly names the "cite-placement" skill. Do NOT auto-trigger on general citation, footnote, or reference requests.

日本語の概要は準備中です。原文の説明を表示しています。

kennethkhoocy/applied-micro-skills222026年9月5日 更新

Download the actual PDF binary from bot-gated sites (taxpolicycenter.org, urban.org, SSRN-hosted mirrors, think-tank/publisher sites) via the Wayback Machine id_ URL form. Use when: (1) curl/WebFetch of a .pdf URL returns HTML instead of a PDF even with a browser User-Agent, (2) pypdf fails with "invalid pdf header: b'<!DOC'" or "EOF marker not found" on a freshly downloaded file, (3) Firecrawl can parse the PDF to markdown but you need the original file on disk (e.g., filing a reference copy).

日本語の概要は準備中です。原文の説明を表示しています。

kennethkhoocy/applied-micro-skills222026年9月5日 更新

Complete methodology for computing publication-quality cumulative abnormal returns with proper event-study test statistics, matching the robustness of Kaspereit's eventstudy2 for Stata. Covers dateline construction, event-date mapping, estimation and event windows, thin-trading adjustment, OLS with Theil prediction error correction, abnormal return computation, CAR/CAAR/AAR accumulation, boundary contamination guards, and common tests such as Patell, BMP, Kolari-Pynnonen, generalized sign, Wilcoxon, and GRANK-T. Use when the user mentions abnormal returns, event windows, market-model regressions, CARs, CAAR, AAR, eventstudy2, thin trading, trade-to-trade returns, or event-study test statistics.

日本語の概要は準備中です。原文の説明を表示しています。

kennethkhoocy/applied-micro-skills222026年9月5日 更新

kennethkhoocy のスキルをすべて見る

このスキルの問題を報告する