# STEP 1947 — orphan-gauge-ruin-50-sample-calibration

**Timestamp**: 2026-09-11T00:15 (JST)
**Tab worktree**: main (rei-aios-9b、shared tree に data/orphan-gauge/step-1947/ 3 artifact tsv 追加)
**Commit**: (this commit)

## 一行 summary

STEP 1926 orphan-gauge 装置の RUIN 1,136 から stratified sample 50 review、precision 100% / 実 actionable 4% / 主 finding = 96% が TOLERATE_ARCHIVE。

## 主要 finding / evidence

- **Sample**: 50 files, sqrt-weighted stratified over 39 clusters (39 clusters, top-3 = external-oss 654 / lean4-mathlib 297 / invention 109 = 93.3% of RUIN)、deterministic seed `sha256("20260911" + path)`
- **Verdicts**:
  - **A (ACTIONABLE_CLEAN)**: **2/50 (4%)** — row 11 `SnstRealComparison.lean` (isolated pilot、Grep で self-ref のみ) + row 19 `karpathy-llm-wiki-watch.json` (dormant radar watch、src/scripts に refs 0)
  - **B (TOLERATE_ARCHIVE)**: **48/50 (96%)** — 内訳: invention approval trail 4 / lean4-transfer staging 2 / external-oss 3rd-party mirror 4 / lean4-mathlib 4 (mathlib-prep + bridge proofs) / paper artifact 3 / publish-log 1 / historical snapshot (`data/*/latest.json` 完了 domain) 18 / dated dataset 12
  - **C (FALSE_POSITIVE)**: **0/50 (0%)** — device が明確に誤 flag した 事例なし
  - **D (UNKNOWN)**: **0/50** — 初回 2 UNKNOWN (rows 11, 19) は追加 Grep で A に解消
- **Device metrics**:
  - **Precision-as-suspicious-flag**: (A + B) / 50 = **100%** — 装置は staleness × referencedness の 交差 flag として 完全精確
  - **Actionable rate**: A / 50 = **4%** — flag された file の うち 実 cleanup 候補は 4%
  - **False positive rate**: C / 50 = **0%**
- **Cluster-weighted global rough estimate** (wide CI、n=50 small sample):
  - lean4-mathlib: 297 files × A rate 20% (1/5) → ~59 actionable
  - research-radar: 1 file × 100% → 1 actionable
  - 他 838 files × A rate 0/44 → 0 (upper CI ~7% → 59)
  - **Rough global A estimate**: **60-120 files (5-11% of 1,136)** — CI 幅大、上限 evidence 弱

## Honest scope

- **主張しないこと**:
  - この review は **n=50 sqrt-weighted stratified sample**、大 cluster (external-oss 654) は 4 sample のみ で 母集団 precision の 信頼区間 は 広い
  - 私 (Claude) 単独 review = **独立 second reviewer なし**、藤本さん or chat-Claude 較正 未実施
  - A verdict 2 件 は 実 cleanup 実行前 に 藤本さん verify 必須 (destructive action 前提)
  - B verdict 48 件 は「藤本さん が 保持 したい historical/archive」 前提 で 分類、藤本さん が「これは 消して良い」 判定する 可能性 排除せず
  - Cluster-weighted global estimate は 各 cluster n≤5 で 統計 confidence 不足、幅を honest に示す のみ
- **前提**:
  - Freshness threshold 30 日、min refs 1 の default 設定
  - Reference graph = explicit-path literal 出現 lower bound (semantic reference は 拾わない、既知 device limitation)
  - Consumption-mode heuristic は human-read のみ exclude、archive detection なし (本 review の 主 finding が これ)

## Failure mode (機械学習用 dataset)

- **(ii) 「device precision 高 = 装置有用性 高」誤認**: 100% precision-as-suspicious-flag は 統計的に 正しい が、user judgment (「これは 保持 したい historical」) を 反映しない 限り、cleanup guidance としての 実 utility は 4% (actionable rate) に 落ちる。Prevention: 装置 metric を **precision × user-agreement rate** の 2 段 で 分けて 報告、precision だけを 「装置が useful」 と 早合点しない。
- **(jj) 大 cluster 少 sample の 信頼区間 隠蔽**: sqrt-weighted stratified で external-oss 654 が n=4 sample → 4 全部 B でも「external-oss の A rate = 0%」と 断定できない、95% CI 上限は ~40%。Prevention: cluster-weighted global estimate 提示時 は 各 cluster CI 幅 も 併記、point estimate だけ 出さない。
- **(kk) 「A verdict 少 = 装置 unnecessary」 誤認**: n=50 で A=2 だが、これは 「装置が 有用でない」 意味 では ない、「default excludeModes に archive を 追加 すれば actionable rate が 上がる」 = 装置 refinement の 実 signal。Prevention: A rate 低い場合、原因が (a) 実際に cleanup 候補少ない (b) 装置 category が 粗すぎる の どちら か を区別、後者なら refinement 提案を 出す。

## 詳細参照 (任意)

- Sample tsv: `data/orphan-gauge/step-1947/sample-50.tsv` (50 rows + header)
- Signals tsv: `data/orphan-gauge/step-1947/sample-50-signals.tsv` (basename + size + last commit + semantic grep hits)
- Verdicts tsv: `data/orphan-gauge/step-1947/sample-50-verdicts.tsv` (row × path × verdict × reason)
- 装置本体: `src/aios/orphan-gauge/` 6 files (STEP 1926、commit `0ddff88d6`)
- Full report JSON: session scratchpad `orphan-gauge-full.json` 5.9 MB (session temporary、not committed)
- 関連 STEP: STEP 1909 (chat-Claude simulator origin)、STEP 1926 (device v0.1 land)

## 追加 section — Device refinement recommendation (v0.2 candidate)

**Highest-signal finding**: 96% of RUIN が TOLERATE_ARCHIVE → `consumption-mode` heuristic の `archive` detection 追加 で RUIN count を 大幅 shrink できる。

**Proposed archive detection heuristics** (`src/aios/orphan-gauge/consumption-mode.ts` に追加):

| Pattern | Rationale | Est. shrink |
|---|---|---|
| `**/approved-*.json` | invention approval audit trail | ~15 files |
| `**/publish-log-*.json` | publish audit trail | ~10 files |
| `**/*-YYYY-MM-DD*.json` | dated dataset (search results, experiment output) | ~50 files |
| `data/external-oss/**` | 3rd-party OSS mirror | 654 files |
| `data/external-research/**` | downloaded reference PDF | ~10 files |
| `data/anki/**` | study material | ~5 files |
| `data/lean4-transfer/**` | Lean staging area | ~15 files |
| `data/*/latest.json` when domain has no fetcher scheduled | completed domain snapshot | ~200 files |

**Expected effect**: RUIN 1,136 → ~150-200 (mostly lean4-mathlib stale pilots + 実 dormant assets)、actionable rate 4% → **40-70%**。装置 practical utility が 定量的に 数十 倍 improve。

**Implementation cost** (rough): consumption-mode.ts に regex + directory rule 追加 (v0.1 の 89 行 → 150 行)、test 追加 (n=+15) = ~1 時間 arc。

**藤本さん judgment 領分**: v0.2 refinement 起票するか、または 別 route (chat-Claude CM-0 装置化 swap) に 振り替えるか。**私 推奨** = v0.2 refinement 先 (低 cost + 明確 evidence)、CM-0 swap は refinement 後 に 再評価。

## Meta finding — なぜ 96% が TOLERATE か

Rei プロジェクトは **append-only 学問体系** = 藤本さん の 承認済 record (invention approval) / paper artifact / historical snapshot / external mirror / dated experiment output を **immutable として 保持** する 慣習。これは Rei の 「急がずゆっくりと」原則 (Load-bearing invention #5) + 「知識は 消去 されず 積み上がる」 SEED_KERNEL 前提 の operational consequence。従って RUIN 分類器の 素朴 semantics (「stale + referenced = 疑わしい」) と、Rei プロジェクトの archive 慣習が semantic mismatch を起こしている、というのが 本 review の 深い 発見。装置は 正しい、mismatch は プロジェクト特性、refinement で 装置側を プロジェクト semantics に aligning する のが 適切な action。
