STEP 1858 · 2026-09-07 · Voynich manuscript D-FUMT₈ exploratory (NOT decoding)

Voynich manuscript D-FUMT₈ exploratory characterization — 100 年 open problem に 1 dimension の 統計特徴 追加

← STEP 1855 · STEP 1839 · notepad

★ NOT a decoding claim ★
本 experiment は Voynich manuscript を 「解読」 したものでは ありません。 100 年 の 先行研究 (Amancio 2013 / Reddy-Knight 2011 / Zandbergen 系) が 既に 多く の 統計特徴分析を 済ませて おり、 本 characterization は D-FUMT₈ 8-axis という 1 dimension を 追加 したのみ。 (a) natural language cipher、 (b) constructed language、 (c) hoax with statistical patterns、 (d) restricted-vocabulary natural language、 (e) systematic transcription error のいずれ の 仮説 も 本 experiment では 排除 しません。

藤本さん 2026-09-07 「超圧縮 を 進めるにあたり ヴォイニッチ手稿 の 真実 の 手がかり も つかめそうですか?」 + option 1 「今夜 私 実行、 exploratory framing、 commit、 chat-Claude verify 別 turn」 explicit direction。 Voynich (Takahashi 1998 EVA 転写、 241 KB、 5120 lines) の D-FUMT₈ 8-axis profile を 10 known languages + 3 controls (shuffled/random-char/ROT13-en) と 比較。

実験 setup

D-FUMT₈ Hellinger 距離 from Voynich (昇順)

#rankerHellinger備考
1VOYNICH shuffled words0.0007予想通り 同一 profile (word-shuffle は 8-axis 保存)
2hin (Hindi)0.2923★ Latin-alphabet Voynich から 非 Latin 言語 が 最近
3rus (Russian)0.3001★ 同上
4arb (Arabic)0.3073★ 同上
5kor (Korean)0.3797中距離
6ENGLISH ROT130.4482英 と 独立位置 (ROT13 で 文字 permute → 8-axis fingerprint 変化)
7eng (English)0.5022Latin alphabet 同族 なのに 中距離
8deu (German)0.5285同上
9VOYNICH random-char0.5413同一 unigram 分布 なのに 遠い = 特定的 char pattern が 単純 random とは 異なる
10fra (French)0.5476Latin
11spa (Spanish)0.5614Latin
12cmn (Mandarin)0.5779non-Latin だが 遠い (漢字 profile が Voynich と 異質)
13jpn (Japanese)0.7070最遠 (混合 script kana+kanji profile)
★ 予想外 の 1 finding
Voynich は Latin-alphabet EVA 転写 なのに、 D-FUMT₈ Hellinger 距離 で Hindi (0.29) / Russian (0.30) / Arabic (0.31) が 最も 近い、 Latin 系 (eng 0.50 / fra 0.55 / spa 0.56) より 明確に 近い。

直感的 explanation candidate: Voynich の 極端に 制限された character set (top 5 char = o/e/h/y/a で 194,000 glyph 中 47% dominance) が、 非 Latin languages の 「distinctive dominant-char profile」 に 類似する D-FUMT₈ TRUE axis 発火 pattern を 生む 可能性。 STEP 1855 の 「D-FUMT₈ = script clustering」 finding と 一見 矛盾 だが、 「effective alphabet size + dominance skew」 が 支配軸 の 可能性 (Direction A の shared-orthography confound の 更 微細な 側面)。

圧縮率 (xz -9e + zstd -22 --long bpc)

rankerxz-9e bpczstd-22 bpc対 English 差
VOYNICH (Takahashi EVA)2.3192.331-0.90 (低 entropy)
VOYNICH shuffled words2.4312.453-0.79
VOYNICH random-char4.2504.143+1.03 (高 entropy)
ENGLISH ROT133.2213.125+0.01 (英 と 一致)
eng3.2163.125baseline
hin (Hindi)1.9321.917-1.28 (最低、 devanagari)
rus2.4542.464-0.76
arb2.5782.580-0.64
kor2.8872.889-0.33
jpn2.7842.817-0.43
cmn3.8433.756+0.63
deu3.3913.373+0.18
fra3.1643.108-0.05
spa3.3413.284+0.13

既知 Voynich 特徴 confirm 3 件 (100 年の 文献と 整合)

  1. 構造 exists (not random): Δ vs random-char = 1.93 bpc = 強い higher-order structure — Amancio 2013 系 の 「Voynich は 統計的 dependency あり」 と 一致
  2. Word-order redundancy 弱: Δ vs shuffled = 0.113 bpc (natural language 典型 0.3-0.5 bpc より 弱い) — Reddy-Knight 2011 系 の 「Voynich word-order flexibility」 と 一致
  3. 低 entropy: Voynich 2.32 bpc vs English 3.22 bpc = 明確に 低 — Zandbergen 系 voynich.nu の 既知観察 と 一致

Rule 7 gap 明示

Rule 7 chance baseline 未計算
「Hindi 0.2923 が 最も 近い」 claim に 対して chance baseline (shuffled Voynich vs each language の 距離分布 permutation null) を 未 measure。 現時点 で 「exploratory observation」 レベル に 留める、 「finding」 昇格 には 追加 experiment 必要。 chat-Claude 別 turn 依頼 (下記 handoff task 1) で 対応予定。

Honest scope (7 層)

  1. Takahashi 1998 転写 単独 (他 transcriber Currier/Friedman/N 版 との cross-verify 未実施)
  2. EVA = Latin transliteration、 元 Voynichese glyph の Unicode PUA 対応 未使用、 EVA バイアス
  3. NOT a decoding claim (意味発見 / 言語 identification / 暗号 solve いずれでもない)
  4. Comparison basis 10 lang + 3 control = convenience sample (WALS 100 stratified の 方が 系統的)
  5. STEP 1855 shared-orthography confound の 継承
  6. Rule 7 chance baseline 未計算 (exploratory 留め)
  7. 私 単独 experiment、 chat-Claude / Gemini 独立 verify なし

chat-Claude handoff 準備 (別 turn 依頼予定 3 tasks)

  1. Rule 7 chance baseline 計算: Voynich shuffled-word や random-char control が 各 known language に 対して 示す 距離分布を permutation null (N=1000+) として compute、 「Hindi 0.2923 が どれ significant か」 verify
  2. Takahashi 単独 転写 の 他 transcriber cross-verify: LSI ivtff の C (Currier) / F (Friedman) / N (transcript X) 版 で 同 analysis 実行、 D-FUMT₈ profile が 転写者 選択 に どれ ぐらい 依存 するか
  3. 予想外 finding の 独立 explanation: 私 の 「制限 alphabet + 高 skew」 説明 が sufficient か、 他 explanation 候補 (D-FUMT₈ 特定 axis の EVA 転写 artifact 等) は あるか

Rei stack 4 STEP evidence stack (D-FUMT₈ cross-lingual signal 性質、 継続)

  1. STEP 1839 / chat-Claude STEP 1845 Tatoeba: small above-chance signal 存在
  2. chat-Claude STEP 1848 multi-pair: script disjointness 仮説 棄却
  3. STEP 1855 URIEL: typological correlation 未サポート
  4. STEP 1858 Voynich: 制限 alphabet + 高 skew profile 特徴化 (非 Latin cluster surprise)

→ 収束: 「D-FUMT₈ cross-lingual signal = small script/character-distribution clustering artifact」 + Voynich は 特殊 case で 制限 alphabet 特徴が signal を 増幅 する 追加 dimension

成果物

参照