Template-heavy domain family — 系統的探索 設計 v0.1
藤本さん note 2026-07-26 ヴォイニッチ手稿 Fable 5 統計解析 arc の 3 発見 (綴り 規則性 / 4 語で 決まり文句 消失 / p<10⁻⁸) を 圧縮 観点 で 再読 → 「short-range 高規則 + long-range 独立」 の domain family。 cmix / modern LLM の 主戦場 (long-range dependency) の 逆側 domain = domain-specific compressor で 系統的 攻めうる。 12 candidate registry + 5 特性測定軸 + 4 phase plan の 設計案 land。
1. Motivation
cmix v21 (1.17 bpc、 STEP 1837 corrigendum 確定値) + modern LLM Shannon bound 0.60-0.80 bpc (STEP 1847 Gemini audit) の 主戦場 は long-range dependency。 ヴォイニッチ 型 domain は その 逆側 (short-range のみ、 long-range 独立) で cmix 過剰能力 領域。 domain-specific compressor family として 系統化 する 副次 direction。
2. Candidate domain registry (v0.1、 12 members)
| # | domain | 分類 | 予想 template ratio |
|---|---|---|---|
| 1 | ヴォイニッチ手稿 (EVA transcription) | anchor | 極 高 (~10%?) |
| 2 | JSON schema instance | 形式 code | 高 |
| 3 | XML template | 形式 code | 高 |
| 4 | Structured log | log format | 中 - 高 |
| 5 | Hash / UUID sequence | random-uniform | 極 低 (max entropy) |
| 6 | Numeric ID sequence | monotone | 極 低 |
| 7 | Constructed language corpus | conlang | 中 - 高 |
| 8 | DNA sequence | biological | 中 |
| 9 | Protein sequence | biological | 中 |
| 10 | Chess PGN | game record | 高 |
| 11 | LaTeX source | mixed markup | 中 |
| 12 | Regex pattern collection | pattern lang | 極 高 |
3. 5 特性測定軸 (Voynich signature 導出)
| axis | 定義 | Voynich 予想 | enwik8 (一般 SOTA 主戦場) 予想 |
|---|---|---|---|
| A. char bigram entropy | H(chart+1 | chart) | 低 (~2-3 bit) | 中 (~4-5 bit) |
| B. n-gram uniqueness decay | 「1 回 のみ 出現」 の 最小 n | n=4 (Voynich signature) | n=8-12 |
| C. long-range mutual info | I(chart; chart+1000) | 極 低 (~0) | 中 (~0.1-0.3 bit) |
| D. template ratio | 「top 10% n-gram で cover される token 割合」 | 高 (>50%) | 中 (~20-30%) |
| E. compressor ratio gap | (gzip / zstd / cmix) の 相対 gap | 小 (cmix 優位薄) | 大 (cmix 30-40% 優位) |
Voynich signature 定義: axis B ≤ 5 AND axis C ≈ 0 AND axis D > 30% を 満たす domain。 「一般 SOTA compressor が 過剰能力」 の 指標。
4. Phased research plan (4 phase)
- Phase 1 (2-4 STEP): §2 registry × §3 5 軸 の measurement script + 12 domain sample corpus 収集 → CSV output。 deliverable = signature match list
- Phase 2 (2-3 STEP): signature-positive short-list 特定 (3-5 domain 予想: hash/UUID は max entropy、 chess PGN 高 template、 conlang / log / regex が signature 候補)
- Phase 3 (5-10 STEP): 1 domain 選択 → template dictionary + n-gram hybrid compressor prototype。 成功基準 = 「選択 domain で cmix 超え」 (一般 SOTA では ない、 1 domain のみ)。 失敗許容 = 「主戦場 で ない domain」 実測 が 系統論文 素材
- Phase 4 (1-2 STEP): 理論枠組 paper draft。 Rei stack comparison-basis contract の domain-selection 変種 論
5. Prior art (novelty 主張なし)
| 既存研究 | 本 arc との 関係 |
|---|---|
| LZMA / LZ4 dictionary preset | 既存 template-based 圧縮 実装、 直接借用可 |
| PPMii (Shkarin) | short-range context model、 baseline |
| Grammar-based coding (Kieffer-Yang 2000) | context-free grammar 抽出、 template dictionary の 数理背景 |
| Tokenizer-based compression (BPE for compression) | subword 特化、 conlang 系 で 直接適用可 |
| Bulatov 2024 Recurrent Memory Transformer | long-range 依存 vs local pattern 分離、 Voynich は local のみ |
| Deletang 2026 "Language Modeling Is Compression" | model-compressor duality、 Voynich 型 は 「小 model + local template」 側 evidence |
本 arc の 唯一の 貢献 candidate 2 点: (1) Voynich signature を characterization 軸 に 昇格、 (2) Rei stack comparison-basis contract を domain-selection に 適用 (「basis 空 で 肯定的 verdict 禁止」 = 「domain 外 で SOTA 主張 禁止」 の 同型)。
6. Honest scope
- 一般 SOTA (enwik8 cmix 打破) は 目標 ではない。 主戦場 の 逆側 domain family の 明示 が 目的
- 1 domain で cmix を 超えても 「一般 SOTA 突破」 では ない (domain-specific compressor は well-known 系統)
- Voynich 単体 圧縮率 は 実験的 に 出せる 見込み だが 「未解読 文書」 で 実用性 なし = artifact 価値
- domain 選択 の 恣意性: §2 registry は 私 (Rei-AIOS session) 主観、 系統性 は Phase 1 measurement で validate
- v0.1 は 設計案 のみ、 実装 + measurement は 別 STEP、 藤本さん explicit go 後
7. Rei stack 統合点
- comparison-basis contract v0.2 の
null-comparedkind に 「domain 分布 vs uniform random」 対比 として 適用可能 - Phase 1 characterization は STEP 1846 R2 「separatingPower を CI から 導出」 と 同型 (metric axis も CI 判定)
- Phase 4 paper は Paper 179 の 「boundary identification」 pattern 継承
8. 対 arc: STEP 1852 (Paper 180 concept)
本 STEP と 同 seed (ヴォイニッチ arc 対話) から 派生 する 2 direction の もう 一方: STEP 1852 Paper 180 concept — 「statistical significance ⊥ semantic content」 の contrastive 2-case study (Voynich Case A + Paper 179 Case B)。 本 STEP 1851 は 「domain-specific compressor 系統」 direction、 STEP 1852 は 「対 論文 direction」。
9. 次 step (藤本さん judgment)
- v0.1 設計 承認 → Phase 1 (registry + characterization) 着手 STEP 予約
- domain 追加/削除 の 指示 (§2 の 12 domain pin or 拡張)
- STEP 1852 (Paper 180 concept) と の 順序 判断
- v0.2 送り: Phase 3 「1 domain 選択」 基準 事前登録、 Phase 4 paper scope (single-domain vs family survey)、 Rei stack 統合 具体化
私 の 推奨: Phase 1 (measurement script + registry validation、 2-3 STEP 想定) 先行、 Phase 2-4 は Phase 1 結果 で recalibrate。
10. 詳細参照
- 設計 doc:
docs/template-heavy-domain-family-v0.1-DRAFT.md - Seed: 藤本さん note 2026-07-26 ヴォイニッチ手稿 Fable 5 統計解析
- 関連 arc: STEP 1846 v0.2、 STEP 1848 bug fix、 STEP 1837 cmix corrigendum、 STEP 1847 Gemini audit
- 対 arc: STEP 1852 Paper 180 concept
- Notepad: 2026-09-07T00-46 STEP 1851
- Fragment: notes/steps/STEP_1851.md
- ← Tools index に 戻る