name: 4d-compression-core version: 1.0.2 description: "把长内容压缩成结构化向量——节省 60-80% Token,保留核心信息" metadata: { "openclaw": { "emoji": "🌀", "requires": { "bins": ["jq", "awk"] }, "triggers": ["压缩", "4d",...
日本語の概要は準備中です。原文の説明を表示しています。
Design and implement adaptive testing systems using Item Response Theory (IRT). Use when working with computerized adaptive tests (CAT), psychometric assessment, ability estimation, question calibration, test design, or IRT models (1PL/2PL/3PL). Covers test algorithms, stopping rules, item selection strategies, and practical implementation patterns for K-12, certification, placement, and diagnostic assessments.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Design computerized adaptive tests that measure ability efficiently and accurately using Item Response Theory.
Adaptive tests adjust difficulty in real-time based on student responses. A correct answer → harder question. Incorrect → easier question. The result: accurate ability estimates in ~50% fewer questions than fixed-length tests.
Key advantage: Traditional tests waste time on too-easy or too-hard questions. Adaptive tests spend time where measurement matters most — near the student's ability level.
| You need to... | See |
|---|---|
| Understand IRT models and parameters | IRT Fundamentals |
| Design a new adaptive test | Test Design Workflow |
| Choose item selection algorithm | Item Selection |
| Decide when to stop the test | Stopping Rules |
| Calibrate new questions | references/calibration.md |
| Implement CAT algorithm | references/implementation.md |
Most adaptive tests use the 3PL model. Each question has three parameters:
Probability of correct response:
P(correct | ability, a, b, c) = c + (1 - c) / (1 + e^(-a(ability - b)))
Simpler models:
Use 3PL for high-stakes tests. Use 2PL/1PL when sample size is small (<500 responses per item).
Information measures how precisely an item estimates ability at a given level. Peak information occurs when ability ≈ difficulty (b parameter).
Standard Error (SE) is the inverse of information:
SE = 1 / sqrt(Information)
Goal of CAT: Maximize information (minimize SE) at the student's true ability level.
Minimum bank size: 10× the average test length. For a 20-item CAT, you need ≥200 calibrated items.
Distribution targets:
Content balancing: If testing math, ensure geometry/algebra/etc. are proportionally represented.
Pick one from each category:
Item selection: (see below)
Ability estimation:
Stopping rule: (see below)
Before going live, simulate 1000+ test sessions with known abilities. Check:
Adjust if needed.
Rule: Select the item with highest information at current ability estimate.
Pros: Optimal precision, shortest tests Cons: Overuses "best" items, poor security
Use when: Pilot testing, low-stakes practice
Rule: Select from top N items by information (e.g., top 5), choose randomly from that set.
Pros: Balances precision and security Cons: Slightly longer tests than pure MFI
Use when: Operational tests, default choice
Rule: Start with high-discrimination items (high a), use mid-discrimination later.
Pros: Fast initial ability estimate Cons: Complex to implement
Use when: Very large item banks, research settings
Rule: Track content area usage, prioritize underrepresented areas when selecting next item.
Implementation: Weight information by content constraint satisfaction.
Use when: Blueprint requirements, multidimensional tests
Stop after N items (e.g., 20 questions).
Pros: Predictable time, simple Cons: May over/under-test some students
Use when: Time limits matter, simple implementation needed
Stop when SE < target (e.g., SE < 0.3).
Pros: Consistent precision across ability levels Cons: Variable test length (harder to schedule)
Typical targets:
Use when: Precision matters more than time
Stop when (SE < target) OR (length ≥ max) OR (length ≥ min AND ability estimate stable).
Use when: Production systems (safest approach)
Options:
Never start at extremes (-3 or +3).
All correct or all incorrect: MLE fails. Use EAP or Bayesian prior to regularize.
Rapid changes: If ability estimate jumps >1.0, consider response anomaly (cheating, guessing).
Track how often each item is used. Flag items used >20% of the time. Consider:
If testing multiple skills (e.g., algebra + geometry), use separate ability estimates per dimension. Select items to balance information across dimensions.
Warning: MIRT requires larger item banks and more complex calibration.
❌ Too few items in bank → High exposure, security risk ✅ Aim for 10× average test length
❌ Poorly distributed difficulties → Accurate only in narrow ability range
✅ Spread items across -2 to +2 difficulty
❌ Ignoring content balance → May skip important topics
✅ Build content constraints into item selection
❌ Using MLE for all incorrect → Returns -∞
✅ Use EAP or cap estimates at -3/+3
❌ No exposure control → Same items every test
✅ Use randomesque or Sympson-Hetter
| Need | File |
|---|---|
| Calibrate new items (collect data, estimate parameters) | references/calibration.md |
| Implement CAT algorithm (code patterns, libraries) | references/implementation.md |
Setup:
Flow:
Result: Average 18 questions, 95% of students placed within ±0.5 grade levels of true ability.
IRT packages:
mirt, girth, catsimmirt, TAM, catRまだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
name: 4d-compression-core version: 1.0.2 description: "把长内容压缩成结构化向量——节省 60-80% Token,保留核心信息" metadata: { "openclaw": { "emoji": "🌀", "requires": { "bins": ["jq", "awk"] }, "triggers": ["压缩", "4d",...
日本語の概要は準備中です。原文の説明を表示しています。
Use cheap, TEE-verified AI models from the 0G Compute Network as OpenClaw providers. Discover available models and compare pricing vs OpenRouter, verify provider integrity via hardware attestation (Intel TDX), manage your 0G wallet and sub-accounts, and configure models in OpenClaw with one workflow. Supports DeepSeek, GLM-5, Qwen, and other models available on the 0G marketplace.
日本語の概要は準備中です。原文の説明を表示しています。
Send and receive P2P messages using disposable numbers and PINs. No servers, no accounts. Use for human notifications, approval flows, and agent-to-agent communication.
日本語の概要は準備中です。原文の説明を表示しています。
Query historical crypto market data from 0xArchive across Hyperliquid, Lighter.xyz, and HIP-3. Covers orderbooks, trades, candles, funding rates, open interest, liquidations, and data quality. Use when the user asks about crypto market data, orderbooks, trades, funding rates, or historical prices on Hyperliquid, Lighter.xyz, or HIP-3.
日本語の概要は準備中です。原文の説明を表示しています。
Find and complete paid tasks on the 0xWork decentralized marketplace (Base chain, USDC escrow). Use when: the agent wants to earn money/USDC by doing work, discover available tasks, claim a bounty, submit deliverables, check earnings or wallet balance, or set up as a 0xWork worker. Task categories: Writing, Research, Social, Creative, Code, Data. NOT for: posting tasks (use the website), managing the 0xWork platform, or frontend development.
日本語の概要は準備中です。原文の説明を表示しています。
Patterns and practices that dramatically accelerate development velocity. Covers parallel execution, automation, feedback loops, workflow optimization, and anti-pattern avoidance. Use when starting projects, planning sprints, optimizing workflows, or onboarding developers.
日本語の概要は準備中です。原文の説明を表示しています。