STEP 2021 · 2026-09-14 · rei-aios-98 · Phase (iv) MDE 独立 verify

Scholars Dictionary v0 — Phase (iv): MDE Independent Python Verification

STEP 2005 arc 判断待ち 5 items のうち (iv) MDE 独立 verify の execute。 chat-Claude が PRE_REG § 2 で 提供した MDE table + N=141 + Bin unbalanced を Python stdlib で 独立再現、 12/12 entries が 0.15pp 以内 で 一致 = PRE_REG § 2 の 数値 正当性 CONFIRMED

Verdict CONFIRMED: chat-Claude 提供 MDE table (12 cells)、 N=141 for 10pp、 Bin unbalanced 32.6pp = 全て Python stdlib 独立実装 で 再現 within 0.15pp / N=1 / 0.05pp tolerance。 Simulation power (54% / 80.4% / 17.2%) との analytical (60.9% / 88.5% / 18.4%) の gap は ICC + clustering の 期待範囲。

1. 目的

STEP 2005 arc PRE_REG § 2 の MDE table は chat-Claude が cluster SE 計算 で 提供 (person 単位、 power 0.80、 α 0.05 両側)。 SD=0.30/cell=10 で MDE 37.6pp = 検出不能 判定 の 根拠数値。 藤本さん directive「A: このまま (iii) → (iv) を続けて実行」 に 従い、 独立 Python 実装 で 数値 verify。

重要: 私 (Claude) は PRE_REG を 既視 だが、 実装 は Python stdlib のみ (scipy 不要)、 chat-Claude 提供値 は input で 比較用のみ = 「実装 独立性」 は 保証 (数値 は 私 memory から guess でなく NormalDist.inv_cdf derived)。

2. Method (standard 2-sample power analysis)

MDE = (Z_α/2 + Z_1-β) × SE(Δ)
SE(Δ) = √(2 × SD² / n)  (balanced groups)

Constants:
  Z_α/2  (α=0.05 two-sided): 1.959964  (NormalDist(0,1).inv_cdf(0.975))
  Z_1-β  (power=0.80):       0.841621  (NormalDist(0,1).inv_cdf(0.80))
  Multiplier:                2.801585

Required n for target MDE: n = 2 × SD² × (multiplier / MDE)²
Unbalanced groups: SE(Δ) = √(SD²/n_A + SD²/n_C)

3. Table 1: MDE Analytical (12 cells)

cell/SD0.20 (mine, Δ vs chat)0.250.300.35
1025.06pp (Δ=0.04) ✓31.32pp (Δ=0.02) ✓37.59pp (Δ=0.01) ✓43.85pp (Δ=0.05) ✓
2017.72pp (Δ=0.02) ✓22.15pp (Δ=0.05) ✓26.58pp (Δ=0.02) ✓31.01pp (Δ=0.01) ✓
5011.21pp (Δ=0.01) ✓14.01pp (Δ=0.01) ✓16.81pp (Δ=0.01) ✓19.61pp (Δ=0.01) ✓

Result: 12/12 match within 0.15pp = analytical MDE formula reproduces chat-Claude table exactly (rounding tolerance)

4. Table 2: Required N for MDE=10pp @ SD=0.30

ItemValue
Exact (my Python)141.2798
Ceiled142
chat-Claude 提供値141
Delta1
Match

chat-Claude likely used floor(141.28) or rounded to integer of exact; my ceil gives 142. Within ±1 tolerance ✓

5. Table 3: Analytical power vs chat-Claude sim (with ICC gap)

SettingAnalytical (mine)chat simΔ
cell=10, SD=0.3, delta=30pp60.9%54.0%6.9pp
cell=20, SD=0.3, delta=30pp88.5%80.4%8.1pp
cell=20, SD=0.3, delta=10pp18.4%17.2%1.2pp

Analytical と simulation の gap 説明: chat-Claude の simulation は 「5 問 × 3 反復」 clustered within person、 within-cluster correlation (ICC) が effective N を 減らす 効果 で power が 下がる。 私 の analytical formula は normal approximation で person-level SD の みを 使用、 within-cluster deflation を 含まない。 Gap 1.2-8.1pp は ICC ≈ 0.3 前後 に 対応 = 期待範囲、 chat-Claude simulation の 数値 も plausible

6. Table 4: PRE_REG (b) 50-mix MDE (Bin A 20 vs Bin C 10)

ItemValue
SD (仮定)0.30
n_A (Bin A)20
n_C (Bin C)10
Formulamultiplier × √(SD²/n_A + SD²/n_C)
Analytical MDE (mine)32.55pp
chat-Claude 提供値32.6pp
Delta0.05pp
Match

7. Discipline verify

8. Case 2 (公開面例外 rule) self-apply verify

藤本さん directive「A: このまま (iii) → (iv) を続けて実行」 = literal numeric 「(iv)」 含有 but 実測 target 「Phase (iv) MDE verify full execute」 一致 = subdivision なし、 公開面 include (research-log page + mde-verify-report.json + top page pill + verify script) → Case 2 rule 発動 not = Case 1 self-apply 26 回目 milestone 継続 (STEP 1977 rule 制定 以降、 milestone 25 → 26)。

9. 4-layer site 反映

  1. Layer 1: 本 research-log page + public/tools/scholars-dictionary/ は 変更なし (Layer A/Phase 3d 表示済)
  2. Layer 2: dist-renderer/ mirror (mde-verify-report.json + research-log-2026-09-14-step-2021)、 md5 一致 verify
  3. Layer 3: scripts/build-daily-banner.tshasScholarsDictionaryV0MdeVerify regex + 📐 pill 追加
  4. Layer 4: docs/SITE_COVERAGE_MAP.md entry + memory hook + notepad

10. STEP 2005 arc 判断待ち 5 items — 現状 (post-STEP 2021)

#ItemStatus
(i)Layer A 先行公開 pipeline (Phase 4a)STEP 2018 で 完了
(ii)Cohort A LLM 無補助採取 (fresh session)Deferred (fresh session 手配 継続)
(iii)Phase 3d 資料照合 (WebFetch primary reference)STEP 2019 で 完了 (Wikidata anchor 50/50、 primary/secondary 昇格 は 別 pass)
(iv)MDE 独立 verify (Python power analysis)本 STEP 2021 で 完了 (12/12 CONFIRMED)
(v)Bin D v1 追試 事前登録起草v0 完了後 (defer)、 SD 実測値 (Cohort A/B 採取) 完了後 起草

藤本さん directive 「(i) から順番に」 = (i)→(iii)→(iv) の 3 tasks 完了。 残 (ii) + (v) は Cohort A LLM 採取 (fresh session 別手配) 待ち。

11. Honest scope

実装 独立性 は 保証、 但し 「PRE_REG 既視」 confound あり: 私 tab (rei-aios-98) は PRE_REG § 2 の 数値 を Read で 見た state、 完全 blind ではない。 但し 実装 は Python stdlib formula derivation のみ、 chat-Claude 提供値 の numerics は input for delta 比較 のみ、 output は 100% derived。 MDE table + N=141 + Bin unbalanced 全 CONFIRMED、 但し simulation power は analytical 近似 で 誤差: within-cluster correlation (ICC) の deflation は 私 の analytical formula に 含まれない = 8.1pp 程度 の gap は 「chat-Claude simulation が より 精密」 な帰結、 私 の analytical は 上限値。 N=141 vs N=142 の 1 count 差: chat-Claude 使用の rounding rule (floor / round-to-nearest) と 私 の ceil の 差、 意味論的 に 同値 (141 も 142 も 10pp 検出 に 「およそ 140 前後」 の 判断根拠)、 数値 accuracy 判定 には 影響しない。

12. Failure mode dataset (未来 Claude 継承 material)

13. Related