本文へ移動
cccskills
無料GitHub で公開

asyncopenai-concurrency-httpx-pool

Raise real concurrency in asyncio LLM batch scorers built on the OpenAI SDK (AsyncOpenAI, including OpenAI-compatible providers like DeepSeek). Use when: (1) raising an asyncio.Semaphore above ~100 produces no throughput gain, (2) a batch pipeline saturates near 100 in-flight requests despite a larger semaphore, (3) planning a high-concurrency campaign against a provider with no hard rate limit (DeepSeek v4-flash tolerates 2000+ in flight). Root cause: AsyncOpenAI's default httpx pool caps max_connections at 100, silently bottlenecking any larger semaphore — you must pass a custom http_client with httpx.Limits sized to the semaphore.

インストール方法を見る

含まれるファイル(2)

  • SKILL.md3.9 KB
  • README.md3.5 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

AsyncOpenAI Concurrency: the Hidden httpx Pool Cap

Problem

Async batch scorers typically gate concurrency with asyncio.Semaphore(N). Raising N above ~100 silently does nothing: the OpenAI SDK's default httpx transport caps the connection pool at max_connections=100, so excess tasks queue inside httpx instead of reaching the provider. The semaphore looks like the throttle but is not the binding constraint — there is no error, just a throughput ceiling.

Context / Trigger Conditions

  • asyncio.Semaphore(N) with N > 100 around client.chat.completions.create shows the same throughput as N = 100
  • Client constructed as AsyncOpenAI(api_key=..., base_url=...) with no http_client argument (the default transport)
  • Provider is known to allow high concurrency (DeepSeek v4-flash: ~2500)
  • Symptom check: requests-in-flight measured at the server never exceeds ~100

Solution

Size the httpx pool to the semaphore when constructing the client:

import httpx
from openai import AsyncOpenAI

CONCURRENCY = 2000
client = AsyncOpenAI(
    api_key=..., base_url="https://api.deepseek.com",
    http_client=httpx.AsyncClient(limits=httpx.Limits(
        max_connections=CONCURRENCY,
        max_keepalive_connections=CONCURRENCY)))
sem = asyncio.Semaphore(CONCURRENCY)

Both edits are required; either alone caps the other. Keep the per-request retry loop — at high concurrency transient failures are more likely, and the retry envelope is what turns them into non-events.

Verification

Throughput scales with N. Verified 2026-07-17 on DeepSeek v4-flash (deepseek-chat, JSON-mode unit scoring, ~1.5k-token prompts): at semaphore 50 a cold 13.8k-request chunk took ~70 min; at semaphore 2000 + matched pool, a 14.4k-request chunk took ~11 min (~40 req/s sustained, ~250/s burst on a 1.7k-request tail chunk), 0 failed requests, 0 schema-invalid responses. Effective speedup ~7x rather than 40x — server-side queuing absorbs the rest — but with zero reliability cost.

Example

Specialist Directors US T1 campaign: final 8 chunks (100,243 calls) completed in ~55 minutes for $8.08 after the fix, versus a projected ~9 hours at the old setting. The edit is two lines in the scorer; the semaphore constant alone would have been a silent no-op.

Notes

  • DeepSeek publishes no hard rate limit and handled 2000 in-flight cleanly; the practical ceiling reported is ~2500. Other providers enforce RPM/TPM caps — check before sizing.
  • Windows: default asyncio proactor loop handled 2000 sockets without tuning; no ulimit-style adjustment needed.
  • Companion ops lesson from the same campaign: when a run's plan changes scale (two-night legs -> one-shot), re-audit the launch glue's hardcoded limits — a wrapper MAX-HOURS=10 safety cap sized for the old plan hard-killed a healthy runner at 97/108 chunks. Caps and budgets in supervisor scripts must be revisited whenever expected duration changes.
  • Concurrency is a pure throughput knob: per-request outputs are unchanged (temperature 0, independent requests), so raising it mid-campaign does not create a scoring seam — unlike model/prompt changes, which do (see [llm-campaign-drift-gate]).

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Build human adjudication / hand-labeling sheets from LLM-pipeline data without evidence truncation. Use when: (1) preparing a CSV/Excel sheet for a human to rule on cases an LLM classifier or rater panel judged, (2) a labeler reports "there is no information to label from" or cells look empty in Excel, (3) excerpt columns cluster at one exact length (e.g. all 1,500 chars — a hard truncation cap). Covers: full rating-basis recovery, Excel 32,767-char cell cap, multi-line CSV mangling, ruling dropdowns, companion text files.

日本語の概要は準備中です。原文の説明を表示しています。

kennethkhoocy/applied-micro-skills222026年9月5日 更新

N-round adversarial review pipeline for empirical research output — the chain from data to LaTeX tables to a manuscript that cites them. A Claude drafter proposes minimal diffs, a deterministic mechanical battery gates every diff from a clean state with a regression gate, a Codex reviewer files check-backed critiques, and a blind judge panel decides residual disputes. Manual-invoke ONLY: trigger when the user explicitly runs /adversarial-empirical-review or names 'adversarial-empirical-review' / 'adversarial empirical review'. Do NOT auto-trigger on generic 'review my results', 'check my tables', or manuscript-editing requests. For prose-style refinement use style-emulation instead; this skill AUDITS WHETHER THE TABLES ARE CORRECT — that each number in the tables is what the analysis code computes, reproduces from the data, and is internally consistent. It is an empirical + code review: the manuscript is read only to resolve table numbering, and prose is not examined.

日本語の概要は準備中です。原文の説明を表示しています。

kennethkhoocy/applied-micro-skills222026年9月5日 更新

Before designing, training, or auditing ANY model that replicates human-annotated labels, audit the annotation protocol's INPUT — the exact document/evidence the human labelers consulted — and give the model that same input. Use when: (1) designing a classifier/LLM extractor whose target is a hand-coded label set, (2) a label-replication model shows low recall concentrated in a label subset and the diagnosis on offer is "the label's information is not in the features", (3) reviewers propose construct splits (e.g. "designation vs record-evident"), adjudication sittings, or per-domain stop rules to explain residual disagreement with gold, (4) validating an extraction pipeline against labels transcribed from a source document. Symptom of the underlying failure: elaborate theory accumulates to explain why gold is "partially unpredictable" when the model was simply never shown the document the annotators read.

日本語の概要は準備中です。原文の説明を表示しています。

kennethkhoocy/applied-micro-skills222026年9月5日 更新

Place pre-screened literature citations into a LaTeX or Word manuscript, or restyle the citations already in one. Three modes: (1) inline placement — inline \cite{}/\citet{}/\citep{} with a compiled references.bib, for author-date journals (APA, MLA, Harvard, Chicago author-date, IEEE, Vancouver); (2) footnote placement — full formatted \footnote{} or OOXML footnotes for legal and notes styles (Bluebook, OSCOLA, Chicago, APA, McGill) with Id./supra short forms; (3) restyle — convert existing footnote citations from one style to another. This skill is manual-invoke ONLY — trigger ONLY when the user explicitly runs /cite-placement or explicitly names the "cite-placement" skill. Do NOT auto-trigger on general citation, footnote, or reference requests.

日本語の概要は準備中です。原文の説明を表示しています。

kennethkhoocy/applied-micro-skills222026年9月5日 更新

Download the actual PDF binary from bot-gated sites (taxpolicycenter.org, urban.org, SSRN-hosted mirrors, think-tank/publisher sites) via the Wayback Machine id_ URL form. Use when: (1) curl/WebFetch of a .pdf URL returns HTML instead of a PDF even with a browser User-Agent, (2) pypdf fails with "invalid pdf header: b'<!DOC'" or "EOF marker not found" on a freshly downloaded file, (3) Firecrawl can parse the PDF to markdown but you need the original file on disk (e.g., filing a reference copy).

日本語の概要は準備中です。原文の説明を表示しています。

kennethkhoocy/applied-micro-skills222026年9月5日 更新

Complete methodology for computing publication-quality cumulative abnormal returns with proper event-study test statistics, matching the robustness of Kaspereit's eventstudy2 for Stata. Covers dateline construction, event-date mapping, estimation and event windows, thin-trading adjustment, OLS with Theil prediction error correction, abnormal return computation, CAR/CAAR/AAR accumulation, boundary contamination guards, and common tests such as Patell, BMP, Kolari-Pynnonen, generalized sign, Wilcoxon, and GRANK-T. Use when the user mentions abnormal returns, event windows, market-model regressions, CARs, CAAR, AAR, eventstudy2, thin trading, trade-to-trade returns, or event-study test statistics.

日本語の概要は準備中です。原文の説明を表示しています。

kennethkhoocy/applied-micro-skills222026年9月5日 更新

kennethkhoocy のスキルをすべて見る

このスキルの問題を報告する