# STEP 1955 — orphan-gauge-v05-recalibration

**Timestamp**: 2026-09-11T02:00 (JST)
**Tab worktree**: main (rei-aios-9b)
**Commit**: (this commit)

## 一行 summary

v0.5 real RUIN 1,555 の stratified sample 50 review → **A=0, B=50, C=0, D=0** = actionable rate **0%** (95% CI 0-7%)、Rei per-STEP methodology で 実質全 TOLERATE、device 大転換 finding。

## 主要 finding / evidence

### Sample composition (v0.5 deterministic RUIN 1,555 から stratified n=50)

- **23 files** `data/lean4-mathlib/CollatzRei/**/*.lean` (46%): per-STEP exhibits (Step846, Step950, Step987 等) + topic proofs (Fermat, Sophie Germain, Poincare, Baker, Siegel 等) + MathlibPrep upstream targets + ComparativeLogicAtlas arc
- **3 files** `src/axiom-os/**/*.ts`: bio-computation-spiral-engine + education-db + fujimoto-dimension-algebra
- **3 files** `scripts/**/*.{ts,py}`: publish-paper-to-miraheze + step789-reverse-math-oracle + check-invention-novelty
- **2 files** each `data/*.json` (STEP 803 + STEP 855 data)、`src/renderer/components/*` (SiteAudit + AncientMysteryTab)、`data/open-problems/**` (leech-e8 + kourovka)、`src/aios/**` (dashboard + knowledge-api)、`data/theory-to-circuit/generated-full/` (generated code)
- **1 file** each from 26 other clusters

### Verdict distribution

| Verdict | Count | Rate | 意味 |
|---|---|---|---|
| **A** ACTIONABLE_CLEAN | **0** | **0%** | 実 cleanup 候補 |
| **B** TOLERATE_ARCHIVE | **50** | **100%** | historical/audit/mirror/snapshot/exhibit |
| **C** FALSE_POSITIVE | 0 | 0% | device 誤 flag |
| **D** UNKNOWN | 0 | 0% | 判定不能 |

**Precision-as-suspicious-flag** (A+B): 100% (device 完全精確)
**Actionable rate**: **0% (0/50)**、95% CI 0-7% (rule of 3: 3/50 = 6%)
**FP rate**: 0%

### Verdict resolution process

- **38 samples**: filename heuristics で B (STEP-numbered exhibits + Rei topic proofs + MathlibPrep + historical data + generated code + tests)
- **12 samples** initially D UNKNOWN → **batch grep resolution**:
  - 11 files hits ≥ 2 (referenced somewhere) → B
  - 1 file `BakerSeparationAttempt.lean` hits=0 → 追加 file-content-check → **substantive pilot** (STEP 1275 sub-2、Eric Merle Lean 4 conditional dependencies、load-bearing Rei arc) → B
- **Pattern (pp) 検証**: STEP 1949 で SnstRealComparison を A→B に corrigendum した pattern と同じ、BakerSeparationAttempt も名前 "Attempt" だけで A 判定 したら wrong だった、file-content-check で substantive pilot 判明 = discipline (pp) 「destructive 提案時 は 対象 file head-Read 必須」 が 再確認された

### ★★★ Device 大転換 finding: Rei actionable rate ~0%

**Historical trajectory**:
- **STEP 1947 (v0.1 RUIN 1,136 sample)**: A=2/50 (4%) — 内訳: SnstRealComparison (STEP 1949 で B に corrigendum) + karpathy watch (STEP 1950 で archive 実行) = **retroactive real A rate = 1/50 = 2%**
- **STEP 1955 (v0.5 RUIN 1,555 sample)**: **A=0/50 (0%)**

**Interpretation**:
- Low-hanging cleanup fruit (karpathy dormant watch) は STEP 1950 で 既 picked
- 残 RUIN は 実質全て **intentional preservation** (Rei append-only 学問体系 + per-STEP proof exhibit style)
- Device precision (100%) と device practical utility は **独立次元**、 precision は完璧、cleanup value は 実測 0%

**User-agreement rate** = 100% (all 50 verdict = user wants to keep) → device output と user desire が **完全乖離**

### Device utility recontextualization

Orphan-gauge の 真の value proposition を STEP 1955 evidence で 再定義:

| Role | Value | Evidence |
|---|---|---|
| **Cleanup guidance tool** | ~0% actionable | STEP 1955 n=50 sample 全 B |
| **Measurement / audit tool** | 100% precise flag | STEP 1954 determinism verify + 2 runs identical |
| **Methodology characterization** | Rei per-STEP style を quantitative visualize | STEP 1949 60.9% isolated finding |
| **Discipline enforcement (retroactive corrigendum trigger)** | 2 event (STEP 1949 SnstRealComparison + STEP 1955 BakerSeparationAttempt) | Both were "A candidate" until file-content-check → B |

## Honest scope

- **主張しないこと**:
  - n=50 は 全 tree RUIN 1,555 の 3.2% sample、cluster-weighted global A rate CI は wide (0-7% by rule of 3)
  - "Actionable rate 0%" は **今日時点** の判定、時間経過 + 藤本さん judgment 変化で個別 file の判定は変わりうる
  - 装置 useless では ない — measurement value + discipline value + methodology characterization value は 残る
  - v0.6+ で "cleanup guidance" 以外の value proposition (e.g., 「装置有用性 monitoring dashboard」) を 開発する path も 開ける
- **前提**:
  - v0.5 stability fix (STEP 1954) 完了で measurement noise 排除、sample n=50 は re-runnable の deterministic evidence
  - Rei per-STEP methodology 認識 (STEP 1949 meta finding) が判定精度を support
  - 私 (Claude Code rei-aios-9b) 単独 verdict + auto-signal batch grep、独立 second reviewer なし

## Failure mode (機械学習用 dataset)

- **(eee) low-hanging fruit picked 後 の 装置 utility 過大主張予防**: STEP 1950 で karpathy 1 件 cleanup 実行 = low-hanging fruit picked、次 sampling で actionable rate が 0 に落ちる の は 予期可能。 Prevention: 装置 utility を time series で 追跡、 「1 件 cleanup 後 の 装置有用性 = 低下」 を honest 提示、 「装置 = 継続 useful」 誤主張 avoid。
- **(fff) precision 100% × actionable 0% の 装置 dilemma 誤解**: precision と actionable rate は 独立次元、precision 完璧でも actionable が 0 なら 「cleanup guidance tool として 実 utility 0」 という 正直な 状態。 Prevention: 装置 value proposition を multi-role (cleanup + measurement + discipline + characterization) で 記述、 単一 metric 「useful vs useless」 の 二値判定 禁止。
- **(ggg) 「A verdict は 必ず A」 の bias 予防**: STEP 1955 の 12 D UNKNOWN のうち 1 file (BakerSeparationAttempt) は grep hits=0 で「これは A だ」と 判定しそうになったが file-content-check で B 判明 = STEP 1949 SnstRealComparison pattern の 再発 予防。 Prevention: A verdict の 最終確定 前 に file-content-check (head-Read 30 行) を **無条件必須** (STEP 1949 pp) と 同じ discipline)、 auto-signal だけで A 判定禁止。

## 詳細参照

- Sample: `data/orphan-gauge/step-1955/v05-sample-50.tsv` (50 rows + config)
- Verdicts: `data/orphan-gauge/step-1955/v05-sample-50-verdicts.tsv` (row × path × verdict × reason)
- Source data: STEP 1954 `data/orphan-gauge/step-1954/v05-report-summary.json` (RUIN 1,555 full list)
- STEP 1947 previous calibration (v0.1): `data/orphan-gauge/step-1947/sample-50-verdicts.tsv`
- Related STEP: 1926 (v0.1) → 1947 (v0.1 calibration) → 1948 (v0.2) → 1949 (v0.3-lite + SnstRealComparison corrigendum) → 1950 (cleanup 1 件) → 1951 (site update) → 1952 (v0.4 + judgments) → 1953 (top page pill) → 1954 (v0.5 stability fix) → **1955 (this — v0.5 recalibration)** = **10-STEP arc**

## 追加 section — v0.6 candidates rethink

STEP 1955 finding で v0.5 candidates roster の 優先順位 が 大幅変更:

**Retired/Deprioritized**:
- ~~Cleanup guidance value 向上 pattern~~ → actionable rate 0% で ROI 実質 0、 defer 永続化
- ~~Lean import-graph parser~~ → per-STEP style で isolated normal、 pilot detection は cleanup source として不成立 (STEP 1949 で既 finding)

**New v0.6 candidates (device role shift)**:
- **v0.6-alpha "audit dashboard"**: 装置を "cleanup" から "measurement/monitoring" に shift、time series で LIVE/ORPHAN/RUIN/DEAD 推移 tracking、Rei プロジェクト の 健康度 metric として 使う
- **v0.6-beta "explicit archive marker"**: `intentionally-preserved: true` YAML frontmatter or `.arch-marker` sidecar file で 個別 file を 「装置無視」 marking、 誤 flag 排除
- **v0.6-gamma "device retirement"**: RUIN concept を Rei-context で "warning-worthy" から "informational" に demote、 UX 上 の 表示強度 弱める

**Recommendation**: v0.6-alpha (audit dashboard) が cost/value 最高 (実測 evidence 継続 = 装置 有用性 の time-series 保持)。v0.6-beta は false 高 cost + limited return。v0.6-gamma は 装置 全 shutdown 判断 = 藤本さん judgment 領分。

**私 推奨**: v0.6 defer 継続、次 arc は 装置外 topic に振り替え、装置 role は 現状 保持 (measurement / discipline enforcement)。
