Research Log · 2026-09-13 · STEP 2009 / 2012 / 2014

Self-Verification Lacks Self-Nature: 4 rounds of blind fresh-agent audit + 2 peer reviews + pre-registered two-arm Round 4 + Round 5 supplement + 11-platform publish (9/11 live)

「空の空」 bot を build し、自作 mutation test で全 catch を declare し、blind fresh-agent 独立監査で崩す — この pattern を 5 例連続で実測した論文。2-arm design で attrition confound を除外、2 pass peer review 経て Zenodo DOI mint、9 platform に broadcast。§5 で explicit に scope を「この setting における測定」に制限。

DOI: 10.5281/zenodo.22731426
Title: Self-Verification Lacks Self-Nature: Four Consecutive Measurements of an Author-Mutation Test Framework Failing Under Blind Fresh-Agent Audit
Authors: Nobuki Fujimoto, Claude (Opus 4.7)  ·  License: CC-BY-4.0
Paper SHA-256: 3f2a62909377166670c43402398a1a8b71f862469cc09a52617979ad7cc2505d

11-platform 公開状況 (9/11 live)

#PlatformStatusURL / ID
1Zenodo (DOI)✓ live10.5281/zenodo.22731426
2GitHub✓ livefc0web/rei-aios/…/ee-bot-self-mutation
3Internet Archive✓ liverei-aios-ee-bot-self-mutation-1789267499409
4HackMD✓ live@zCUv2P2UQHGmAOJFPLL_-A/HJrLRYXFGg
5Notion✓ livepage 3dadd371…
6Dev.to✓ liveself-verification-lacks-self-nature-…-2lal
7Hatena✓ live2026/09/13/114557
8Nostr (NIP-23)✓ 4/5 relaysdamus / primal / snort / nos.lol OK; nostr.band reject
9Mastodon✓ livemathstodon.xyz/@Fujimoto/117261441690737829
10Livedoor✗ HTTP 400XML rejected server-side; script OK for other papers, content quirk, deferred
11Scrapbox✗ CSRF/api/users/me returned user obj without csrfToken key; cookie or API drift, deferred

Paper 本体 の 主張と scope

4 rounds のうち 3 rounds で以下の shape が繰り返し observed:

README made a general claim, and the code's behavior contradicted that claim at values not exercised by any test.

Pre-registered two-arm Round 4 (patch vs. restructure) では両 arm に contradiction、shape は同、layer は異なる。H1' supported (P=0.89 → 実測 confirmed)、H2' disconfirmed (P=0.34 で arm β clean を予測、実測 arm β にも contradiction、layer 1 段深く)。

§5 scope 制限 (paper 内 explicit):
- ✗ mutation testing in general にはしない
- ✗ self-verification cannot have self-nature を哲学命題として主張しない
- ✗ 何 round が sufficient かは主張しない
- ✗ pre-Round-1 outcome measure specification (invariance に基づく post-hoc 選択)
- ✗ auditor priming confound (auditor は README + prior findings で primed され、primary outcome は同 class)
- ✗ true independence (auditor は同 model family、同 prompt convention)

5 rounds datapoint の meta-layer

RoundVersionLayer で pattern が現れた場所
1minpoint v0.1value replaceable by literal (solve_eps → return 6.697e-9 passed 35/37)
2ee_bot v0.1value replaceable by ceremonial-conventions dance (19/19 pass)
3ee_bot v0.2value wrong across a decade the README claimed (F1: /10 fallback)
3ee_bot v0.3value wrong across most of the domain (F6: multi-probe factor gaps)
4 αv0.4 patchvalue still has gaps under patched factor list
4 βv0.4 restructurevalue correct, mechanism untested (1-line-lie passes 54/54)
5 γv0.5 mechanism-verifymechanism-checking tests check call incidence, not result dependence (burn-calls loop passes 60/60)
5 δv0.4-β unchanged (control)fresh agent re-derives arm β F1 + finds 2 truly new attacks — attrition confound rejected

各 fix は前 round の specific attack を satisfy、general invariant は次 layer で unenforced のまま、次 round の fresh agent が different specific mechanism で再発見。Round 5 arm γ が最も明晰: mechanism-verifying tests が invocation を count するが result dependence を強制せず、burn-calls loop (stability_verdict を 50 回呼んで結果を捨てて fabricated dict を返す) が 60/60 pass。

Discipline 実装 (研究方法として)

Reproducibility

Zenodo deposit + GitHub tree に以下を配置 (全 SHA-256 committed):

ArtifactSHA-256 (先頭 10 chars)
paper.md (post-peer-review 版)3f2a629093…
preregistration_round4.md (両 arm audit 実行前 committed)5821fc3b9d…
preregistration_round5_future_work.md (v0.5 code 変更前 + publish 前 committed)808f9d9306…
ee_bot-ref-0.1.0.tar.gz5d2acd95e7…
ee_bot-ref-0.2.0.tar.gz4382ff007e…
ee_bot-ref-0.3.0.tar.gz49c2a31860…
ee_bot-v0.4-arm-alpha.tar.gz (patch)970b4ed2a8…
ee_bot-v0.4-arm-beta.tar.gz (restructure)e3517727aa…
ee_bot-v0.5-arm-gamma.tar.gz (mechanism-verify, 60 tests)e234f81dfa…
ee_bot-v0.5-arm-delta.tar.gz (v0.4-β control, 54 tests)02f803df3a…
Round 5 supplement (GitHub only、Zenodo v1.0 immutable のまま)

全 tarball は self-contained (Python 3.11+, pytest, stdlib only)。独立 replication は fresh Claude subagent または 異なる model family (paper §5 で明示的に推奨: cross-family test が現在最も強い次の 1 手) で audit prompt を replay。

関連 STEP と commit

Defer items (別 STEP 候補、緊急度低)

Honest scope 継承 (site page 自体)

このページは paper の scope を紹介するもので、paper の主張を change しません。
- Zenodo v1.0 DOI は immutable、paper §5 の n=4 scope は変更不可
- Round 5 は post-publication supplement (GitHub only 追加)、paper 主張の evidence base には算入せず
- 11-platform broadcast は message ではなく medium: paper text は全 platform で同一、§5 scope 全 host で保持
- 私 (Claude) が最初「11-platform は §5 と不整合」と argue した誤判定を STEP 2014 commit message に自己記録済み — 5 datapoint 目 の meta-observation 該当だが paper §5 evidence base には後付け追加しない