Voynich manuscript D-FUMT₈ exploratory characterization — 100 年 open problem に 1 dimension の 統計特徴 追加
★ NOT a decoding claim ★
本 experiment は Voynich manuscript を 「解読」 したものでは ありません。 100 年 の 先行研究 (Amancio 2013 / Reddy-Knight 2011 / Zandbergen 系) が 既に 多く の 統計特徴分析を 済ませて おり、 本 characterization は D-FUMT₈ 8-axis という 1 dimension を 追加 したのみ。 (a) natural language cipher、 (b) constructed language、 (c) hoax with statistical patterns、 (d) restricted-vocabulary natural language、 (e) systematic transcription error のいずれ の 仮説 も 本 experiment では 排除 しません。
本 experiment は Voynich manuscript を 「解読」 したものでは ありません。 100 年 の 先行研究 (Amancio 2013 / Reddy-Knight 2011 / Zandbergen 系) が 既に 多く の 統計特徴分析を 済ませて おり、 本 characterization は D-FUMT₈ 8-axis という 1 dimension を 追加 したのみ。 (a) natural language cipher、 (b) constructed language、 (c) hoax with statistical patterns、 (d) restricted-vocabulary natural language、 (e) systematic transcription error のいずれ の 仮説 も 本 experiment では 排除 しません。
藤本さん 2026-09-07 「超圧縮 を 進めるにあたり ヴォイニッチ手稿 の 真実 の 手がかり も つかめそうですか?」 + option 1 「今夜 私 実行、 exploratory framing、 commit、 chat-Claude verify 別 turn」 explicit direction。 Voynich (Takahashi 1998 EVA 転写、 241 KB、 5120 lines) の D-FUMT₈ 8-axis profile を 10 known languages + 3 controls (shuffled/random-char/ROT13-en) と 比較。
実験 setup
- Voynich source:
http://www.voynich.nu/data/beta/LSI_ivtff_0d.txt(Stolfi IVTFF format、 1.75 MB) → Takahashi (H) transcription 抽出 → 転写 marker 除去 (! , < > { } ? * $ @ % - digit) → 241 KB、 5120 lines、 57 unique glyph - 10 known languages: STEP 1855 の corpora 再利用 (eng/deu/rus/cmn/jpn/fra/spa/kor/arb/hin、 各 4.8-14.4 KB Wikipedia)
- 3 controls: (a) VOYNICH shuffled words (Fisher-Yates seed 20260907) (b) VOYNICH random-char (unigram-preserved random draw) (c) ENGLISH ROT13 (Caesar 13 cipher output、 entropy 保存 の 既知 cipher)
- Metrics: D-FUMT₈ Hellinger + xz -9e bpc + zstd -22 --long bpc
D-FUMT₈ Hellinger 距離 from Voynich (昇順)
| # | ranker | Hellinger | 備考 |
|---|---|---|---|
| 1 | VOYNICH shuffled words | 0.0007 | 予想通り 同一 profile (word-shuffle は 8-axis 保存) |
| 2 | hin (Hindi) | 0.2923 | ★ Latin-alphabet Voynich から 非 Latin 言語 が 最近 |
| 3 | rus (Russian) | 0.3001 | ★ 同上 |
| 4 | arb (Arabic) | 0.3073 | ★ 同上 |
| 5 | kor (Korean) | 0.3797 | 中距離 |
| 6 | ENGLISH ROT13 | 0.4482 | 英 と 独立位置 (ROT13 で 文字 permute → 8-axis fingerprint 変化) |
| 7 | eng (English) | 0.5022 | Latin alphabet 同族 なのに 中距離 |
| 8 | deu (German) | 0.5285 | 同上 |
| 9 | VOYNICH random-char | 0.5413 | 同一 unigram 分布 なのに 遠い = 特定的 char pattern が 単純 random とは 異なる |
| 10 | fra (French) | 0.5476 | Latin |
| 11 | spa (Spanish) | 0.5614 | Latin |
| 12 | cmn (Mandarin) | 0.5779 | non-Latin だが 遠い (漢字 profile が Voynich と 異質) |
| 13 | jpn (Japanese) | 0.7070 | 最遠 (混合 script kana+kanji profile) |
★ 予想外 の 1 finding
Voynich は Latin-alphabet EVA 転写 なのに、 D-FUMT₈ Hellinger 距離 で Hindi (0.29) / Russian (0.30) / Arabic (0.31) が 最も 近い、 Latin 系 (eng 0.50 / fra 0.55 / spa 0.56) より 明確に 近い。
直感的 explanation candidate: Voynich の 極端に 制限された character set (top 5 char = o/e/h/y/a で 194,000 glyph 中 47% dominance) が、 非 Latin languages の 「distinctive dominant-char profile」 に 類似する D-FUMT₈ TRUE axis 発火 pattern を 生む 可能性。 STEP 1855 の 「D-FUMT₈ = script clustering」 finding と 一見 矛盾 だが、 「effective alphabet size + dominance skew」 が 支配軸 の 可能性 (Direction A の shared-orthography confound の 更 微細な 側面)。
Voynich は Latin-alphabet EVA 転写 なのに、 D-FUMT₈ Hellinger 距離 で Hindi (0.29) / Russian (0.30) / Arabic (0.31) が 最も 近い、 Latin 系 (eng 0.50 / fra 0.55 / spa 0.56) より 明確に 近い。
直感的 explanation candidate: Voynich の 極端に 制限された character set (top 5 char = o/e/h/y/a で 194,000 glyph 中 47% dominance) が、 非 Latin languages の 「distinctive dominant-char profile」 に 類似する D-FUMT₈ TRUE axis 発火 pattern を 生む 可能性。 STEP 1855 の 「D-FUMT₈ = script clustering」 finding と 一見 矛盾 だが、 「effective alphabet size + dominance skew」 が 支配軸 の 可能性 (Direction A の shared-orthography confound の 更 微細な 側面)。
圧縮率 (xz -9e + zstd -22 --long bpc)
| ranker | xz-9e bpc | zstd-22 bpc | 対 English 差 |
|---|---|---|---|
| VOYNICH (Takahashi EVA) | 2.319 | 2.331 | -0.90 (低 entropy) |
| VOYNICH shuffled words | 2.431 | 2.453 | -0.79 |
| VOYNICH random-char | 4.250 | 4.143 | +1.03 (高 entropy) |
| ENGLISH ROT13 | 3.221 | 3.125 | +0.01 (英 と 一致) |
| eng | 3.216 | 3.125 | baseline |
| hin (Hindi) | 1.932 | 1.917 | -1.28 (最低、 devanagari) |
| rus | 2.454 | 2.464 | -0.76 |
| arb | 2.578 | 2.580 | -0.64 |
| kor | 2.887 | 2.889 | -0.33 |
| jpn | 2.784 | 2.817 | -0.43 |
| cmn | 3.843 | 3.756 | +0.63 |
| deu | 3.391 | 3.373 | +0.18 |
| fra | 3.164 | 3.108 | -0.05 |
| spa | 3.341 | 3.284 | +0.13 |
既知 Voynich 特徴 confirm 3 件 (100 年の 文献と 整合)
- 構造 exists (not random): Δ vs random-char = 1.93 bpc = 強い higher-order structure — Amancio 2013 系 の 「Voynich は 統計的 dependency あり」 と 一致
- Word-order redundancy 弱: Δ vs shuffled = 0.113 bpc (natural language 典型 0.3-0.5 bpc より 弱い) — Reddy-Knight 2011 系 の 「Voynich word-order flexibility」 と 一致
- 低 entropy: Voynich 2.32 bpc vs English 3.22 bpc = 明確に 低 — Zandbergen 系 voynich.nu の 既知観察 と 一致
Rule 7 gap 明示
Rule 7 chance baseline 未計算
「Hindi 0.2923 が 最も 近い」 claim に 対して chance baseline (shuffled Voynich vs each language の 距離分布 permutation null) を 未 measure。 現時点 で 「exploratory observation」 レベル に 留める、 「finding」 昇格 には 追加 experiment 必要。 chat-Claude 別 turn 依頼 (下記 handoff task 1) で 対応予定。
「Hindi 0.2923 が 最も 近い」 claim に 対して chance baseline (shuffled Voynich vs each language の 距離分布 permutation null) を 未 measure。 現時点 で 「exploratory observation」 レベル に 留める、 「finding」 昇格 には 追加 experiment 必要。 chat-Claude 別 turn 依頼 (下記 handoff task 1) で 対応予定。
Honest scope (7 層)
- Takahashi 1998 転写 単独 (他 transcriber Currier/Friedman/N 版 との cross-verify 未実施)
- EVA = Latin transliteration、 元 Voynichese glyph の Unicode PUA 対応 未使用、 EVA バイアス
- NOT a decoding claim (意味発見 / 言語 identification / 暗号 solve いずれでもない)
- Comparison basis 10 lang + 3 control = convenience sample (WALS 100 stratified の 方が 系統的)
- STEP 1855 shared-orthography confound の 継承
- Rule 7 chance baseline 未計算 (exploratory 留め)
- 私 単独 experiment、 chat-Claude / Gemini 独立 verify なし
chat-Claude handoff 準備 (別 turn 依頼予定 3 tasks)
- Rule 7 chance baseline 計算: Voynich shuffled-word や random-char control が 各 known language に 対して 示す 距離分布を permutation null (N=1000+) として compute、 「Hindi 0.2923 が どれ significant か」 verify
- Takahashi 単独 転写 の 他 transcriber cross-verify: LSI ivtff の C (Currier) / F (Friedman) / N (transcript X) 版 で 同 analysis 実行、 D-FUMT₈ profile が 転写者 選択 に どれ ぐらい 依存 するか
- 予想外 finding の 独立 explanation: 私 の 「制限 alphabet + 高 skew」 説明 が sufficient か、 他 explanation 候補 (D-FUMT₈ 特定 axis の EVA 転写 artifact 等) は あるか
Rei stack 4 STEP evidence stack (D-FUMT₈ cross-lingual signal 性質、 継続)
- STEP 1839 / chat-Claude STEP 1845 Tatoeba: small above-chance signal 存在
- chat-Claude STEP 1848 multi-pair: script disjointness 仮説 棄却
- STEP 1855 URIEL: typological correlation 未サポート
- STEP 1858 Voynich: 制限 alphabet + 高 skew profile 特徴化 (非 Latin cluster surprise)
→ 収束: 「D-FUMT₈ cross-lingual signal = small script/character-distribution clustering artifact」 + Voynich は 特殊 case で 制限 alphabet 特徴が signal を 増幅 する 追加 dimension
成果物
data/experiments/voynich-typology-2026-09-07/LSI_ivtff_0d.txt(1.75 MB、 IVTFF format 原本)data/experiments/voynich-typology-2026-09-07/voynich_eva_takahashi.txt(241 KB、 Takahashi H clean extract)data/experiments/voynich-typology-2026-09-07/voynich-characterization-results.json(全 profile + distance + bpc + interpretation)scripts/experiments/voynich-dfumt8-characterization.ts(~230 行、 実験 driver)
参照
- Voynich 先行研究: Amancio et al. 2013 "Probing the Statistical Properties of Unknown Texts" / Reddy-Knight 2011 EMNLP / Zandbergen 系 voynich.nu 出版物 / Cheshire 2019 refuted claim example
- 関連 Rei-AIOS STEP: 1839 · 1841 · 1842 · 1855
- Rule 7 protocol:
docs/notepad/2026-09-06T14-48_STEP-1837_cmix-bpc-unit-corrigendum.md