---
name: project-voynich-s2l-v11-v12-layer2-sim-2026-07-28
description: "★★★★★★★ 2026-07-28 Voynich Layer 2 sim (v1.1+v1.2+v1.3+v1.4+v1.5-a) + evening reflection 完全完了。 (C) Parisel Markov 警告完全実現。 (B) v1.2 で VMS A structural 同等。 (v1.3) VMS B structural essentially exact (0.201 vs 0.199 = 1.01×)、 A+B 同時 4/4、 unigram non-reducible gain 0.195。 (v1.4) trigram I(X;Z|Y) VMS A signature -0.007 = word-boundary axis で v1.3 が VMS ceiling 到達済、 Parisel gap は relocate = negative-space finding。 **(v1.5-a) diagnostic sweep で真 gap = word-type richness axis 特定** + extended word grammar (prefix + [m1] + [m2] + suffix) 3 iterations root cause chase → **p1=0.7 で VMS A の 4/5 lexical metrics essentially exact match** (TTR 0.336 / Zipf -0.806 / word length 5.06 perfect / types 3356、 全 VMS A と 1.0-1.07× ratio)、 hapax 差 11pp 残、 Sig1-4 4/4 A+B 保存、 attestation 26-30% 低下 = richness 代償。 **evening reflection**: hapax 絶対数分析で VMS 2200 vs Rei 1980 = **VMS の tail が真に長い**、 3 hypotheses (VMS 固有 / pool wider / 単純 tail 短) 判定不能で Zattera 12-slot 独立検証項目化。 attestation drop は negative claim (意味不要でも Voynich 統計再現可) を強化する方向で agreed caveat wording 固定。 **generator-family lesson 2 pattern 記録** (middle pool overlap → class-flip / duplicate-avoidance resampling → marginal 歪め) 次 generator design 参照点。 藤本さん指示 (C)→(B)→(A)→v1.3→v1.4→v1.5-a→reflection 全完遂。"
metadata: 
  node_type: memory
  type: project
  originSessionId: 34ae4094-d429-47e5-ba3d-7f8d1376a8d3
  modified: 2026-07-28T02:43:04.589Z
---

## 概要

2026-07-28 Voynich Layer 2 simulation 完全 record。 07-27 late evening close 時に用意した v1 spec (二階建て設計 + 質量表 + 3 design 原則) を Rei env で実装、 3 実験 (v1.1 baseline / Parisel-warning discriminator / v1.2 strong coupling) を通した。

## 環境

- Parisel `currier-signatures` repo (github.com/labyrinthinesecurity/currier-signatures) を scratchpad に clone
- `signatures_v26.py` alias 作成 (repo bug: grille.py が v26 参照、 v27 内容同一で cp で解決)
- Python 3.13.2 / numpy 2.2.3 / scipy 1.15.2
- Calibration reproducibility: VMS Currier A = E→S 71.0% / MI 0.5856 / Bilat YES (Se=3, Ee=3) / Zipfian R²=0.863 CV=1.45 / N=9892 (Parisel README 完全一致)
- 全 artifact = `scratchpad/currier-signatures/` (session-isolated)

## v1.0 → v1.1 tuning

**v1.0 (bridge o.prefix 1500 / o.suffix 1200)**: 20 seeds/swing configs 中 5/20 のみ 4/4。 Sig1 E→S 平均 86.6% で上限 86% 超過が主 fail。 root cause: pure final × pure start 構造で E→S 過大 (predicted 0.879 × 0.946 = 0.83)。

**v1.1 (bridge o.prefix 2500 / o.suffix 2000)**: **全 20 configs 4/4 PASS** (swing 0.20-0.35 × 5 seeds)。 E→S mean 77.5% (VMS ref 71%)、 MI chunked mean 0.31 (5-chunk 内部評価)、 Bilat Se=14-17 Ee=10-12、 Shape Intermediate (all seeds)。

**Official verifier.py cross-check (swing=0.30 seed=42)**:
- Currier A: 4/4 PASS — E→S 78.7% / MI 0.4894 (0.84× ref, excellent match) / Bilat YES / Zipfian R²=0.879
- Currier B: 4/4 PASS — E→S 78.7% / MI 0.4894 / Bilat YES / Zipfian (同 generator で偶発的に B も通過)

## (C) Ablation + Shuffle Retention + word-Markov baseline

### Rule 個別 contribution ablation (swing=0.30 seed=42, full corpus MI)

| Config | E→S% | MI_full | MI_shuf | retention% | Shape | 4/4 |
|--------|------|---------|---------|------------|-------|-----|
| v1.1 full (I+II+III+IV) | 78.5 | 0.0929 | 0.0833 | **89.7** | Interm | 3 |
| Rule I only (aiin↔y) | 77.5 | 0.0816 | 0.0825 | 101.2 | Interm | 3 |
| Rule II only (ey↔ar) | 77.4 | 0.0845 | 0.0839 | 99.3 | Interm | 3 |
| Rule III only (dy↔al) | 77.5 | 0.0804 | 0.0826 | 102.7 | Interm | 3 |
| Rule IV only (bridge o) | 78.6 | 0.0841 | 0.0829 | 98.5 | Interm | 3 |
| Swing 0 (no coupling) | 78.6 | 0.0841 | 0.0829 | 98.5 | Interm | 3 |
| **Marginal only** | 77.5 | **0.0796** | 0.0828 | 104.0 | Interm | 3 |

**Rule contribution to MI_full** (delta vs marginal baseline 0.080):
- Rule I: +0.002 / Rule II: +0.005 / Rule III: +0.001 / Rule IV: +0.005
- v1.1 full total: **+0.013** (合計しても baseline の 15%)

### VMS full-corpus MI (direct comparison)

| Corpus | E→S% | MI_full | MI_shuf | retention% | structural MI |
|--------|------|---------|---------|------------|---------------|
| VMS A | 69.0 | 0.209 | 0.125 | **60%** | 0.084 |
| VMS B | 62.8 | 0.281 | 0.081 | **29%** | 0.199 |
| Rei v1.1 | 78.5 | 0.093 | 0.083 | **90%** | **0.010** |

**Rei v1.1 の structural MI は VMS A の 1/8、 VMS B の 1/20**。 chunked 4/4 pass は Miller-Madow bias artifact (chunk 小さいほど MI 過大評価)。

### word-Markov baseline (Parisel 警告の実測)

| Baseline | E→S% chunked | MI chunked | 4/4 chunked | retention% |
|----------|--------------|------------|-------------|-----------|
| Rei S2L v1.1 | 79.7 | 0.489 | **PASS** | 90% |
| **Markov unigram from S2L v1.1** | 79.7 | **0.485** | **PASS** | 103% |
| Markov bigram from S2L v1.1 | 79.4 | 0.525 | PASS | 56% |
| **Markov unigram from VMS A** | 68.0 | **0.509** | **PASS** | 103% |
| Markov bigram from VMS A | 69.7 | 0.605 | PASS | 42% |

**★★★ Parisel 警告完全実現**: word 単純頻度 (unigram) だけで 4-signature 全 pass。 Rei v1.1 の 4/4 は unigram Markov と区別不可能 = protocol は positional/generative mechanism を要求しない。 Rule 群は shuffle 後もほぼ同じ MI = structural に不足。

## (B) v1.2 strong bijective coupling

### 設計 (s2l_v2.py)

- **Bijective suffix → prefix-group mapping** (v1.1 の 3 対称対を各 top-6 suffix に distinct target 割当):
  - y → q (30% prefix mass、 最大対称)
  - aiin → k-group / ey → ok-group / ar → ch-group / dy → d-group / al → s-group
  - Extended: ain → t-group / in → p/f / iin → cph/cfh / ir → ck / edy → c
- **Swing**: multiplicative bias ×(1+swing) for target, ×(1-swing×0.5) for others、 renormalize
- Pool identical to v1.1

### Sweep results (10K words, full-corpus MI + chunked)

| Swing | seed | E→S% | MI_full | MI_chunk | ret% | Shape | 4/4 |
|-------|------|------|---------|----------|------|-------|-----|
| 0.40 | 42-44 | 77-78 | 0.09-0.11 | 0.49-0.51 | 81-88% | Zipfian | 4/4 |
| **0.60** | 42-44 | 77.6-78.6 | 0.12-0.13 | 0.52-0.54 | 64-69% | Zipfian | 4/4 |
| 0.80 | 42-44 | 78.0-79.1 | 0.16-0.18 | 0.56-0.58 | 48-50% | Zipfian | 4/4 |
| 0.90 | 42-44 | 78.3-79.4 | 0.19-0.21 | — | 40-42% | Interm | — |
| 0.95 | 42-44 | 78.5-79.6 | 0.21-0.23 | 0.60-0.62 | 36-40% | Interm | 4/4 |

### Official verifier (v1.2 swing=0.60 seed=42)

- **Currier A 4/4 PASS**: E→S 78.6% / MI 0.5189 (0.89× ref) / Bilat YES / Zipfian R²=0.894 CV=1.81
- **Currier B 4/4 PASS**: E→S 78.6% (borderline, [50,79]% window 際どい slack 0.4pp) / MI 0.5189 (1.04× B ref)

### Official verifier (v1.2 swing=0.95 seed=42)

- **Currier A 4/4 PASS**: E→S 79.5% / MI **0.6045** (1.03× ref、 near-exact match) / Zipfian
- **Currier B 3/4 PASS**: E→S 79.5% > 79% window → Sig1 FAIL (他 3 sig pass)

### structural MI 検証 (unigram Markov discriminator)

| Corpus | MI_full | retention | structural MI |
|--------|---------|-----------|---------------|
| **Rei v1.2 swing=0.95** | 0.210 | **40%** | **0.126** |
| Unigram Markov from v1.2 | 0.087 | 96% | 0.004 |
| Delta (structural gain not surface) | | | **0.122** |
| VMS A | 0.209 | 60% | 0.084 |
| VMS B | 0.281 | 29% | 0.199 |

**Rei v1.2 の structural MI (0.087 at swing=0.60、 0.126 at swing=0.95) は VMS A (0.084) と同等 or 超過**、 VMS B (0.199) には未達。 **Unigram Markov では再現不能** (delta 0.122) = genuine structural coupling。

## 4 「絶対にやらない」 checklist (map late evening 準拠)

- ✗ 閾値下限緩めた — なし (Sig3 [0.293, 1.171] / [0.249, 0.996] そのまま)
- ✗ 方言混ぜて overall と比較 — なし (dialect-specific 使用)
- ✗ Sig3 外して 3/4 で報告 — なし (v1.2 swing=0.95 Currier B の Sig1 fail を 3/4 として honest 記載)
- ✗ MI を上限超過で「達成」 — なし (Sig3 0.51-0.60 for MI target、 下限側 pass、 v1.2 swing=0.95 で ref 直下ぴったり match)

## honest scope (Parisel リスト添付候補)

1. **4-signature protocol は word-unigram Markov で通過可能**。 Rei S2L v1.1 chunked 4/4 pass は Parisel 警告 (surface Markov reproduction) が現実化した実例
2. **v1.2 strong coupling (swing 0.60-0.95) で genuine structural MI 到達**: unigram Markov 対照で non-reducibility 確認 (delta 0.122)、 structural MI 0.087-0.126 は VMS A (0.084) と同等/超過
3. **v1.3 extended coupling + concentration=0.20 で VMS B structural essentially exact**: structural MI 0.201 vs VMS B 0.199 (1.01×)、 MI_full 0.286 vs 0.281 (1.02×)、 retention 30% vs 29% (1.03×)、 5/5 seeds robust、 non-reducible gain 0.1947 = **VMS A と B 両 dialect の structural profile を同時 match**
4. **v1.3 conc=0.20 が全体最適**: A/B 両 dialect で同時 4/4 PASS (5 seeds robust)、 E→S 77.5-78.9% (B upper 79 slack 0.1-1.5pp、 v1.2 swing=0.60 の 0.4pp と同水準)、 Sig3 MI 0.68 は A range [0.293, 1.171] 1.16× / B range [0.249, 0.996] 1.36× 中央付近、 上限超過なし
5. **Attestation v1.3**: A dialect 48.8% / B dialect **61.3%** (v1.1/v1.2 の 47% から B 側大幅向上、 soft diagnostic)
6. **Shape**: v1.3 は全 5 seeds で Zipfian (v1.2 swing 0.90-0.95 で Intermediate になっていたが v1.3 は改善)
7. **Bridge o mass 過剰 (VMS 実測 3.3% に対し 15% 使用)** の妥協は継続。 VMS 実態と structural profile 一致は両立するが lexical shape 逸脱の代償
8. **v1.3 実装ポイント**: concentration=0.20 で 「20% の transitions のみが coupling table を発火、 80% は marginal」 = 一見弱い coupling だが coupling table 内が near-deterministic (target group から weighted sampling) のため実効 MI は大きい。 swing multiplicative model (v1.2) の頭打ちを回避

## Parisel リスト 「S2L (v1.3)」 一行 (draft, updated with bigram caveat)

> **S2L (probabilistic context-sensitive L-system, extended bijective suffix→prefix-group coupling, concentration=0.20)**: 4/4 PASS on both Currier A (E→S 77.6%, MI 0.680) and B (E→S 77.6%, MI 0.680) simultaneously across 5 seeds. Structural profile (full-corpus): MI 0.286, MI_shuffled 0.085, retention 30% — essentially exact match to VMS Currier B (MI 0.281, retention 29%, structural MI 0.199 vs Rei 0.201, ratio 1.01×). Structural gain non-reducible to word-unigram Markov = 0.195 bits. **However, word-bigram Markov trained on VMS Currier B itself reaches structural MI 0.242 (exceeding VMS B's own 0.199 via smoothing), so v1.3's structural coupling is at word-bigram-equivalent expressive power. The 4-signature protocol AND full-corpus structural MI are both reproducible by word-bigram surface reconstruction; the "positional / higher-order / generative mechanism" gap identified by Parisel is not closed by v1.3.** Rei S2L v1.3 is the first tested generator with structural MI profile matching VMS Currier B while passing chunked verification for both dialects, but achieves this at bigram-equivalent power. Attestation: Currier A dialect 48.8% / Currier B dialect 61.3% (soft diagnostic).

## artifacts (scratchpad-only, next-session resume 用)

- `scratchpad/currier-signatures/` — cloned Parisel repo (RF1b-e.txt, verifier.py, signatures_v27.py, grille.py, signatures_A/B.txt)
- `s2l_v1.py` — v1.0/v1.1 generator (Rule I+II+III+IV, rule enable flags)
- `s2l_v2.py` — v1.2 generator (bijective coupling)
- `sweep.py` / `sweep_v2.py` — swing×seed sweep drivers
- `ablation.py` — rule ablation + retention
- `markov_baseline.py` — word-Markov discriminator (both S2L and VMS)
- `vms_full_mi.py` — VMS full-corpus MI reference
- `vms_calibration.json` — cached calibration (E→S/MI/Bilat/Shape for A/B)
- `gen_v11_swing30_seed42.txt` / `gen_v2_swing60_seed42.txt` / `gen_v2_swing95_seed42.txt` — reference outputs
- `sweep_results.json` / `sweep_v2_results.json` / `ablation_results.json` — raw data

**注**: scratchpad は session-isolated、 次 session で必要なら repo re-clone + script は本 memory から再構築。 memory 保存必要は spec (mass 表 + 二階建て + coupling table) と findings (retention profile + Parisel warning 実現) のみ。

## 藤本さん指示履行状況

- ✅ (C) shuffle 対照 + ablation → Parisel 警告完全実現 confirmed
- ✅ (B) v1.2 tuning → structural MI VMS A matched (0.126 vs 0.084)
- ✅ (A) record close → 本 memory + Parisel リスト 一行 draft
- ✅ **v1.3 → VMS B structural essentially exact** (0.201 vs 0.199 = 1.01×、 3 metrics 全一致、 5/5 seeds robust、 A+B 同時 4/4)
- ✅ **bigram-Markov 対照実験** (07-28 続き): v1.3 structural 0.201 vs bigram Markov from v1.3 = 0.256 (bigram が上回る)、 VMS A structural 0.084 vs bigram from A = 0.143、 VMS B 0.199 vs bigram from B = 0.242 = **全 3 corpus で bigram Markov が学習元の structural MI を上回る smoothing 現象**、 v1.3 の structural coupling は本質的に word-bigram レベル = Parisel の「higher than surface」 gap の中には入っていない honest finding。 Parisel が要求した positional/higher-order/generative-mechanism 特有性は v1.3 では未達成 (bigram-equivalent power 止まり)
- ✅ **v1.4 trigram context + I(X;Z|Y) 測定**: v1.4 (s2l_v4.py) で prev-prev.suffix 依存の 2 段目 swing 実装、 trigram_swing sweep 0-0.90。 Sig1-4 全 4/4 A AND B PASS 保存。 I(X;Z|Y) は trigram_swing 上昇で 0.711→0.750 上昇 confirmed。 **★ decisive finding**: **VMS A trigram signature -0.007 (bigram と区別不能) / VMS B trigram signature -0.258 (bigram より constrained)** = VMS 自身が word-boundary MI 軸で bigram-decomposable = **v1.3 が既に正しい metric で VMS 上限到達済**。 Parisel gap は word-boundary MI axis には存在せず、 (a) Zattera 12-slot within-word / (b) longer-range / (c) word-type richness / (d) repetition patterns など 別 metric axis に relocate

## bigram-Markov 対照 finding (07-28 追加)

| Corpus | structural MI | Bigram Markov 学習後 structural MI | 比率 |
|--------|---------------|-----------------------------------|------|
| Rei v1.3 conc=0.20 | 0.2011 | **0.2556** | 1.27× |
| VMS Currier A | 0.084 | 0.143 | 1.70× |
| VMS Currier B | 0.199 | 0.242 | 1.21× |

**説明**: bigram Markov は corpus の bigram 分布を smoothed version として抽出、 sampling で noise-free 版を生成 → structural MI が実質増加。 v1.3 の generative coupling も VMS の観察 coupling も、 抽象化すると word-bigram-equivalent power 内で説明可能。

**帰結**: v1.3 は VMS B structural profile 3 metrics essentially exact match だが、 **「positional/generative mechanism を必要とする」 という Parisel gap の core claim** は未 close。 v1.3 が示したのは 「generative rule set で bigram-equivalent structural MI を出せる」 まで。 「bigram Markov では出せない higher-order structure」 は v1.4 (positional slot 統合 or trigram coupling) の課題として残る。

## v1.4 実装 + trigram signature test (07-28 追加)

**設計** (`s2l_v4.py`): v1.3 base coupling に prev-prev.suffix (H=top-6 mass 39% vs L=other) による subgroup bias 追加、 trigram_swing パラメータで強度制御。 Sig1-4 は preserved (bivariate marginals 不変)。

**Key metric 導入**: `I(X;Z|Y)` = prev-prev.suffix と next.first の conditional MI (prev.suffix 条件付き)。 pure bigram Markov では理論上 0 (実測は Miller-Madow bias で ~0.7 出る)。

### 決定的 finding — VMS 自身が bigram-decomposable

| Corpus | I(X;Z|Y) | Bigram Markov 対照 | trigram signature |
|--------|----------|--------------------|--------------------|
| **VMS Currier A** | 0.7315 | 0.7389 | **-0.007** (ほぼゼロ) |
| **VMS Currier B** | 0.5465 | 0.8040 | **-0.258** (負値: bigram markov より少ない) |
| Rei v1.3 (no trigram) | 0.7110 | — | — |
| Rei v1.4 swing=0.60 | 0.7244 | 0.7001 | +0.024 |
| Rei v1.4 swing=0.90 | 0.7498 | — | +0.039 |

**帰結**:
1. **VMS は word-boundary MI 軸で bigram-decomposable** — trigram signature は observable でない (A) or negative (B、 bigram markov より constrained)
2. v1.4 trigram_swing は **働いている** (I(X;Z|Y) 0.71 → 0.75 順調上昇) が **VMS 側に埋めるべき gap がない**
3. **v1.3 は既に word-boundary MI 軸で VMS 上限到達済** = 正しい metric で ceiling hit
4. **Parisel の「positional / higher-order / generative mechanism 特有性」 gap は word-boundary MI 軸には存在しない**
5. Sig1-4 は全 trigram_swing で 4/4 A AND B PASS 保存

### 未探索 gap axis (v1.5+ 候補)

Bigram-Markov で捕捉できない構造は word-boundary MI **以外** の metric に存在する可能性:

- (a) **Zattera 12-slot within-word structure**: word 内部の positional constraint
- (b) **Longer-range dependency**: sentence/paragraph level correlation
- (c) **Word-type richness**: distinct word count / hapax legomena / Zipf 逸脱
- (d) **Repetition patterns**: 隣接 word 反復性 (VMS B の negative trigram signature の原因候補)

Parisel が主張した gap の実在は本 arc で refute でなく **relocate** された = word-boundary axis 外に探すべし。

## v1.5 diagnostic sweep 実測 (07-28 追加、 4 metric axes)

**目的**: v1.4 後の gap-relocate finding を受け、 4 candidate axes を一度に実測して真の gap 特定。

### 実測 gap table (Rei v1.3 vs VMS baseline)

| Metric | VMS A | VMS B | Rei v1.3 | Gap 大きさ |
|--------|-------|-------|----------|-----------|
| **Hapax rate** | **70.4%** | 68.8% | **28.1%** | **-42%** (最大) |
| **Zipf slope** | -0.79 | -0.90 | -1.29 | -0.4〜-0.5 |
| **TTR** | 0.318 | 0.232 | 0.081 | -0.15〜-0.24 |
| **Distinct types** | 3148 | 4554 | **812** | 4-6× 不足 |
| **Word length mean** | 5.08 | 5.18 | **3.61** | -1.5 char |
| **Word length distinct** | 14 | 14 | 7 | 半分 |
| Adjacent rep | 1.06% | 0.89% | 0.97% | ≈ 0 (match) |
| 2-back rep | 0.97% | 1.06% | 0.95% | ≈ 0 (match) |
| Sentence length | 6.4 | 7.9 | 8.5 | +0.6〜+2 (微差) |

**発見**: 真の Parisel gap は **word-type richness axis** = hapax/TTR/Zipf/word-length で v1.3 は VMS の 1/3〜1/2 水準。 repetition axis は既に match。

## v1.5-a 実装 — extended word grammar (prefix + [m1] + [m2] + suffix)

### 設計 (s2l_v5a.py)

- **word structure**: prefix + optional middle_1 (prob p1) + optional middle_2 (prob p2) + suffix
- **middle pool 選定 3 iterations** (root cause chase):
  - **v1**: prefix/suffix pool 全 overlap → `a.start/end` 比 2.19 で class-flip、 E→S drift 80-86%
  - **v2**: overlap 除去 + SUFFIX_POOL['a'] 400→900 → E→S 79-88% (依然 drift)
  - **v3 (final)**: **duplicate-avoidance resampling 撤廃** + middle pool = `o` (bridge)+`e`+low-mass prefixes (`ckh`, `cth`, `cph`, `cfh`, `ch2`, `sh2`, `cth2`) → **marginal 保存 + E→S 76-79% 収容 ✓**

### v1.5-a v3 実測 (10 configs、 各 seed=42/43)

| Config | V | TTR | hapax% | Zipf | wLen | E→S | 4A | 4B |
|--------|---|-----|--------|------|------|-----|-----|-----|
| VMS A | 3148 | 0.318 | 70.4% | -0.794 | 5.08 | — | — | — |
| VMS B | 4554 | 0.232 | 68.8% | -0.896 | 5.18 | — | — | — |
| **v1.5a p1=0.3 p2=0.1** | 2060 | 0.206 | 53.6% | -1.005 | 4.20 | 77.8 | ✓ | ✓ |
| **v1.5a p1=0.5 p2=0.2** | 2727 | 0.273 | 57.8% | -0.899 | 4.61 | 76.3 | ✓ | ✓ |
| **★ v1.5a p1=0.7 p2=0.3** | **3356** | **0.336** | 59.1% | **-0.806** | **5.06** | 77.3 | ✓ | ✓ |
| **v1.5a p1=0.9 p2=0.4** | 3729 | 0.373 | 60.2% | -0.757 | 5.46 | 78.7 | ✓ | ✓ |

**★ sweet spot: p1=0.7 p2=0.3** — VMS A profile を **4/5 lexical metrics で essentially exact match**:
- Distinct types 3356 vs 3148 (**1.07×**)
- TTR 0.336 vs 0.318 (**1.06×**)
- Zipf slope -0.806 vs -0.794 (**1.02×**)
- Word length 5.06 vs 5.08 (**1.00× ← perfect**)
- Hapax 59% vs 70% (0.84×、 差 11pp 残存 — future gap)

### Official verifier confirm (v1.5a p1=0.7 p2=0.3 seed=42)

- **Currier A: 4/4 PASS** — E→S 77.3% / MI 0.6653 / Bilat YES / Zipfian R²=0.855 CV=1.57
- **Currier B: 4/4 PASS** — E→S 77.3% ∈ [50,79] (slack 1.7pp) / MI 0.6653 ∈ [0.249, 0.996]
- Attestation: A dialect 25.7% (v1.3 48.8% から低下) / B dialect 29.9% (v1.3 61.3% から低下) — **richness を得た代償**、 soft diagnostic

### 累計到達点 (v1.1 → v1.5-a)

| 軸 | v1.1 | v1.2 | v1.3 | v1.4 | **v1.5-a** |
|----|------|------|------|------|------------|
| Sig1-4 (chunked) | 4/4 A | 4/4 A+B | 4/4 A+B | 4/4 A+B | **4/4 A+B** |
| structural MI (word-boundary) | 0.010 | 0.126 (VMS A) | **0.201 (VMS B)** | 0.201 | 0.201 |
| unigram non-reducible | 0.008 | 0.122 | 0.195 | 0.195 | 0.195 |
| trigram signature | — | — | — | -0.02 (VMS -0.007) | 同 |
| **TTR** | 0.08 | 0.08 | 0.08 | 0.08 | **0.336 (VMS 0.318 match)** |
| **Zipf slope** | -1.29 | -1.29 | -1.29 | -1.29 | **-0.806 (VMS -0.794 match)** |
| **Word length** | 3.6 | 3.6 | 3.6 | 3.6 | **5.06 (VMS 5.08 match)** |
| Hapax rate | 28% | 28% | 28% | 28% | 59% (VMS 70% 差 11pp) |

**v1.5-a により 4/5 lexical richness metric が VMS A essentially exact match**。 Hapax rate のみ 11pp 残存 gap = future v1.6 target (rare token pool 拡張 or Zattera 12-slot 統合)。

## 07-28 evening 藤本さん reflection turn (整理 + agreed caveat)

藤本さん本 arc 到達点 valuation:
- **v1.4 negative-space finding**: 「gap を close するのでなく relocate する」 = 探索としては失敗に見えて実際は最も情報量の多い結果。 word-boundary MI 軸で VMS 自身が bigram-decomposable と分かった以上、 そこを追う理由がなくなった (research direction 削除は positive contribution)
- **v1.5-a 意味**: prefix + middle + suffix という単純な word grammar だけで TTR / Zipf slope / word length / distinct types 4 metrics を VMS A と 1.0-1.07× で揃えつつ、 Sig1-4 を両 dialect 4/4 維持 = **「意味を持たない生成機構でも Voynich の統計を再現できる」 側の主張をかなり強くする**
- **残 2 concerns 特定**: hapax 11pp 差 + attestation drop

### Hapax 11pp 差 — 絶対数分析

| Corpus | Types | Hapax rate | Hapax count (絶対) |
|--------|-------|-----------|----------------------|
| VMS A | 3148 | 70.4% | **~2200** |
| Rei v1.5-a | 3356 | 59.1% | **~1980** |

Types 総数は v1.5-a の方が多いが、 hapax count は VMS の方が多い = **VMS の tail が真に長い**。

**3 candidate hypotheses (現状判定不能)**:
1. VMS に generator にない mechanism (typo / scribe hand 差 / multi-word glued together / 意図的 rare form) がある = **VMS 固有**
2. VMS の underlying pool が我々の modeled prefix/suffix pool より wider = **pool sampling 問題**
3. 我々の generator の tail が短いだけで rare token を足せば埋まる = **単純 engineering gap**

**独立検証項目**: Zattera 12-slot 系 generator を別実装、 それでも同 hapax 差 11pp が残るなら 「generator-class 共通 gap」 で hypothesis 1 (VMS 固有) の可能性上昇。 v1.5-a 単体では判定不能。 藤本さん agreed: 「hapax が本当に閉じないなら、 それこそが VMS 固有の何かかもしれません」。

### Attestation trade-off — publish caveat 合意版

| version | A dialect | B dialect | 特徴 |
|---------|-----------|-----------|------|
| v1.3 | 48.8% | 61.3% | 少 vocabulary で VMS pattern 高忠実 |
| v1.5-a | 25.7% | 29.9% | VMS 未見 combinations 70% |

**藤本さん指摘**: 「richness を未 attested な word form で買った部分がある。 publish 時には caveat として明示が要りそうです」

**方向性の意味 (Rei 追記 + 藤本さん agreed)**: attestation drop は negative claim を強化する方向:
- Sig1-4 & lexical statistics を match するのに VMS-attested vocabulary は必要ない
- **「意味を持たない機構でも Voynich 表面統計は再現できる」 の主張がむしろ強くなる**

**Publish caveat 合意 wording** (agreed 07-28 evening):
> "Attestation drops from v1.3 48-61% to v1.5-a 26-30% because vocabulary expansion introduces novel combinations. This trade-off strengthens the negative claim: 4 lexical richness metrics + Sig1-4 can all be matched even though 70% of generated tokens are VMS-unattested vocabulary."

### 3 iterations root cause chase — generator-family lesson patterns

藤本さん指摘: 「同じ罠を次の軸でも踏むはずなので」 (次の generator design で同じ marginal-preservation 問題が起きる予想)。 memory 化した 2 教訓 pattern:

**Pattern 1: middle pool overlap → class-flip**
- middle として追加する glyph が既存 prefix/suffix pool と overlap すると、 その glyph の start/end count 比が変化 → class 分類が flip
- v1.5-a v1 で `a` が ambig (ratio 2.19) から start に flip、 E→S drift 80-86% 誘発
- **対策**: middle pool は overlap を避けるか、 overlap 時は overlap 元の suffix mass を追加調整して比を保存
- **一般化**: 任意の generator で「中間要素追加」 する場合、 その要素の既存 classification (Sig1 の start/end/ambig 等) への影響を pre-check 必須

**Pattern 2: duplicate-avoidance resampling → marginal 歪め**
- 「suffix が middle と duplicate したら resample」 のような naive avoidance ロジックは、 その要素の marginal 分布を歪める (bridge o/a 等 popular token が under-represented)
- v1.5-a v3 で resample 撤廃、 duplicates 許容 → marginal 保存 → E→S 76-79% window 収容成功
- **対策**: duplicate は許容する。 「不自然に見える」 が marginal 保存の方が優先
- **一般化**: 任意の generator で resampling / rejection sampling 実装時、 rejection criterion が本来 marginal を歪めていないか check 必須。 特に stratified sampling でない場合。

これらは Voynich domain 外の generator design (chat-Claude collab や別 project) にも適用可能。 次に同 issue に当たった時、 本 record を参照して同 chase を repeat せずに済むよう保存。

### future work items (v1.5-a arc close 時点)

- **(v1.5-b) Zattera 12-slot 統合**: within-word positional structure = hapax 11pp gap の hypothesis 1/2/3 判定に必要な独立検証
- **(v1.6) rare token engineering**: hapax closing の 直接 approach (hypothesis 3 の場合効く)
- **(publish 準備)** Parisel リスト添付、 caveat 明示版で finalize
- **(別 topic 転換)**: 3-site 見直し等

## Related

- [[reference-voynich-prior-art-map-2026-07-27]] — Parisel 2026 v2 基軸マップ、 v1 spec (質量表 + 二階建て) 出典
- [[project-session-2026-07-27-voynich-prior-art-arc-full]] — 前日 arc の full record、 v1 spec 決定経緯
- [[feedback-prior-art-list-grep-verify-each-entry]] — 9 rule 体制
- [[feedback-rule-add-requires-fire-audit]] — 型 X (先に出すべき数字) discipline、 本 sim で全 sweep 結果は型 X 準拠
- [[feedback-projection-self-audit-pattern]] — 5 rule 体制
- [[feedback-evaluation-symmetry-principle]] — v1.1 の 4/4 pass を inflate せず、 v1.2 の VMS B 未到達も deflate せず両方記録
- [[feedback-no-rush-publication]] — 急がずゆっくりと、 Parisel リスト添付は spec finalization 後
