Hacking at the Voynich manuscript - Side notes 101 Preparing clean samples of various other languages Last edited on 2026-07-23 18:21:13 by stolfi SUMMARY Here we prepare text samples in English, Latin, and other languages, comparable in size to the Voynichese reference sample, for the statistical analyses that will go into the "lexeme structure" technical report. SETTING UP THE ENVIRONMENT Links: ln -s ../tr-stats/dat ln -s ../tr-stats/tex ln -s ../../../work ln -s ../../../langbank ln -s work/wds_to_tlw.gawk ln -s work/update_paper_include.sh ln -s work/format-words-filled format_words_filled.sh ln -s work/compute_freqs.gawk ln -s work/compute_freqs.gawk NAMING THE SAMPLE TEXTS A sample is identified by a pair {smp = "{LANG}/{BOOK}"} where {LANG} is the general language and writing system ("engl" for English in standard spelling, "chip" for Mandarin Chinese in pinyin, etc.) and {BOOK} is the source document ("wow" for War of the Worlds, "ptt" for the Pentateuch, "voa" for Voice of America broadcasts, etc.). A sample may be divided into sections or sub-samples for the purpose of studying statistical variations within the same document. A section is identified by a name {sec = "{TAG}.{N}"} where {TAG} is a descriptive string ("gen" for Genesis, "exo" for Exodus, "hea" for Herbal-A) and {N} is a serial number in case the section is split into separate segments (as the VMS Herbal-A is split into "hea.1", "hea.2", etc.). Every sample must have a "tot.1" section (the whole sample), which must be processed only after any other sections. The samples, sections, and their attributes are specified in the table "sample-sections.tbl". FORMAT OF THE SAMPLE TEXTS For each sample {smp="{LANG}/{BOOK}"} and each section {sec="{TAG}.{N}"} (including "tot.1"), we produce a file called "dat/{smp}/{sec}/whole.tlw" that contains the corresponding "raw" tokens, suitably encoded, filtered, and tagged as "good" or "bad" for linguistic analysis. The file "whole.tlw" is derived from a reference file "LANG/BUK/main.wds" from my Linguistic Sample Bank ("/home/staff/stolfi/projects/langbank/"). This file is also linked as "dat/{smp}/org/main.wds". The "LANG/BUK" of the source may be different from the "LANG/BUK" of the sample {smp}; for example, the Langbank source "engl/wow" is used to produce the samples "engl/wow" (full text in lowercase), "engl/wnm" (proper names only), and "envg/wow" (Vigenère-coded text). A truncated version "dat/{smp}/{sec}/trunc.tlw" of the "whole.tlw" file is also created, containing a specified maximum number of "good" tokens. The roughly uniform sizes of these files makes them more suitable for certain comparative analyises, such as Zipf law plots. For compatibility with the VMS samples, each ".tlw" file must EXCLUDE any blanks, embedded comments, alignment fillers, or punctuation (including line and paragraph breaks); or any other tokens that are semantically equivalent to them. On the other hand, the ".tlw" files must INCLUDE any symbols that stand for tokens in the text, such as numbers, abbreviations, and "*"-surrogates for illegible or omitted tokens. In particular, each ".tlw" file must EXCLUDE any "null-like intrusions", i.e. undesirable sub-sections of the selected section that are syntactically equivalent to a null or blank space -- such as margin notes, footnotes and footnote marks, etc.. On the other hand, it must INCLUDE any "symbol-like intrusions", i.e. undesirable sub-sections that play a syntactic role in the text -- such as formulas, tables, poems, foreign phrases, etc.. However, the tokens of a symbol-like intrusion should be replaced by "*" symbols to clearly distinguish them from valid tokens; and, if the intrusion has two or more consecutive tokens, only the first and last one should be kept. The entries of a ".tlw" file have the format "{TYPE} {LOC} {TOKEN}", where: {TYPE} is either "a" or "s"; {LOC} is the full location of the token (section ID plus line number) with components delimited by braces (e.g. "{b1}{c3}{sA}{tx}{73}"); and {TOKEN} is the token or symbol in question. The type "a" denotes plain ("gud") tokens of the language that are suitable for letter-level linguistic analysis, such as letter frequencies and correlations, word length distribution, etc. The type "s" denotes anomalous ("bad") tokens that ought to be excluded from such analyses, such as numerals, abbreviations, symbols, foreign-language intrusions, unreadable tokens, etc.. Note that the "bad" tokens cannot be discarded, because they are relevant for some investigations, such as lexeme-pair correlations, concordances, etc.. The conversion of the source "main.wds" to the sample file "whole.tlw" and "trunc.tlw" is done by {wds_to_tlw.gawk} and consists of the following steps: (1) Read the tokens from "main.wds", which must have been roughly classified as comments (type="#"), alpha tokens ("a"), symbol tokens ("s"), punctuation chars ("p"), section starts ("$") and line starts ("@"). (2) Test each input token with a sample-specific procedure {smp_reclassify_token} from the library {smp}/sample_fns.gawk. This procedure should re-assign type "x" to any tokens that lie outside section ${SEC} or are symbol-like intrusions ???NEW: including "p" entries which are parag-like breaks :WEN???; and as type="n" any null-like intrusions. (4) Based on the reclassified type, discard all entries except "a", "s" and "x". (5) To the remaining tokens, apply a sample-specific global token transformation, e.g. map upper to lower case, map Chinese ideograms to pinyin, change the alphabet, delete vowels or diacritics, split hyphenated tokens, etc. This is also a chance to delete any undesired tokens that could not be discarded by {smp_reclassify_token}. This step uses a sample-specific word substition table {smp}/word_map.tbl followed by a sample-specific function {smp_fix_token} from the library {smp}/sample_fns.gawk. (6) Insert the {LOC} field, prepend "*" to any "x" tokens (to avoid confusion), and squeeze any long runs of the latter. (7) Re-classify each "a" and "s" token as "gud" or "bad", using the predicate {smp_is_good_token} from the same library; "x" tokens are automatically "bad". Replace the type tag by "a" for the "gud" tokens, "s" for the bad tokens. Write both types the "all.tlw" file. The "whole.tlw" file is then truncated ???NEW: at both ends :WEN??? to a specified number of "gud" records, producing the "trunc.tlw" file. The file "trunc.tlw" for each sample and section is copied to "raw.tlw" for compatibility with other Notes workbooks, and then split into dat/{smp}/{sec}/{sizeopt}/gud.tlw - the "gud" subset. dat/{smp}/{sec}/{sizeopt}/bad.tlw - the "bad" subset. From these files are also created the derived files dat/{smp}/{sec}/{sizeopt}/raw.wfr - lexeme occurrence counts in "raw.tlw". dat/{smp}/{sec}/{sizeopt}/gud.wfr - lexeme occurrence counts in "gud.tlw". dat/{smp}/{sec}/{sizeopt}/bad.wfr - lexeme occurrence counts in "bad.tlw". dat/{smp}/{sec}/{sizeopt}/raw.wdf - tokens from "raw.tlw", w/o locations, line-filled. dat/{smp}/{sec}/{sizeopt}/gud.wdf - tokens from "gud.tlw", w/o locations, line-filled. dat/{smp}/{sec}/{sizeopt}/bad.wdf - tokens from "bad.tlw", w/o locations, line-filled. dat/{smp}/{sec}/{sizeopt}/raw-wds-summary.tex - TeX include file for tech report dat/{smp}/{sec}/{sizeopt}/gud-wds-summary.tex - TeX include file for tech report dat/{smp}/{sec}/{sizeopt}/bad-wds-summary.tex - TeX include file for tech report The files {raw,gud,bad}.{tlw,wdf,wfr} are temporarily created also for the full sample "{smp}/{sec}/{sizeopt}/whole.tlw", but are then overwritten for the "trunc" version. do_note_101.sh 1 RESULTS As of 2026-07-23: # Counts for gud text (whole) # Counts for gud text (trunc) # sample/sec tokens lexemes unique # sample/sec tokens lexemes unique # -------------- ------- ------- ------- # -------------- ------- ------- ------- arab/qcs/tot.1 74212 15873 9603 arab/qcs/tot.1 35027 9025 5649 arab/qph/tot.1 77845 17380 10742 arab/qph/tot.1 35027 9434 6044 arab/qud/tot.1 77455 15314 9109 arab/qud/tot.1 35027 8531 5245 arab/quf/tot.1 77394 19852 12911 arab/quf/tot.1 35027 10935 7353 arab/quv/tot.1 77411 19530 12595 arab/quv/tot.1 35027 10762 7187 chin/ptn/deu.1 35370 1463 368 chin/ptn/deu.1 35027 1457 367 chin/ptn/exo.1 40159 1450 305 chin/ptn/exo.1 35027 1439 321 chin/ptn/gen.1 49305 1555 317 chin/ptn/gen.1 35027 1380 312 chin/ptn/num.1 39792 1308 294 chin/ptn/num.1 35027 1254 288 chin/ptn/tot.1 193319 2266 318 chin/ptn/tot.1 35027 1399 289 chin/ptt/exo.1 35252 1424 275 chin/ptt/exo.1 35027 1424 277 chin/ptt/gen.1 45081 1503 301 chin/ptt/gen.1 35027 1376 276 chin/ptt/num.1 36843 1303 312 chin/ptt/num.1 35027 1291 310 chin/ptt/tot.1 174364 2177 278 chin/ptt/tot.1 35027 1396 285 chin/red/tot.1 706889 4271 585 chin/red/tot.1 35027 2420 663 chin/voa/tot.1 58813 1886 376 chin/voa/tot.1 35027 1616 348 chip/voa/tot.1 59476 930 114 chip/voa/tot.1 35027 830 98 chrc/red/tot.1 706889 4271 585 chrc/red/tot.1 35027 2420 663 engl/cul/her.1 112695 5685 2402 engl/cul/her.1 35027 3399 1551 engl/cul/tot.1 120721 6065 2562 engl/cul/tot.1 35027 3555 1656 engl/twp/tot.1 81498 6799 3465 engl/twp/tot.1 35027 4202 2225 engl/wow/tot.1 60273 6788 3244 engl/wow/tot.1 35027 4870 2466 engp/mkw/tot.1 49126 22479 19990 engp/mkw/tot.1 35027 16728 14907 engp/mkx/tot.1 49126 22479 19990 engp/mkx/tot.1 35027 16728 14907 engp/mky/tot.1 49126 22479 19990 engp/mky/tot.1 35027 16728 14907 engp/mkz/tot.1 41194 25475 23740 engp/mkz/tot.1 35027 21908 20454 enrc/wow/tot.1 60293 6790 3245 enrc/wow/tot.1 35027 4870 2466 envg/wow/tot.1 60273 19113 13033 envg/wow/tot.1 35027 12911 9130 envt/wow/tot.1 42123 1709 286 envt/wow/tot.1 35027 1638 286 fran/tal/tot.1 54061 8102 4555 fran/tal/tot.1 35027 6223 3698 germ/sim/tot.1 184498 18556 10020 germ/sim/tot.1 35027 6826 4223 grek/nwt/tot.1 66183 8301 4163 grek/nwt/tot.1 35027 5436 2824 hebr/tad/tot.1 66311 19556 12807 hebr/tad/tot.1 35027 11856 7842 hebr/tav/tot.1 66311 20976 14023 hebr/tav/tot.1 35027 12487 8498 ital/psp/tot.1 216969 18965 9671 ital/psp/tot.1 35027 6623 4085 latn/nwt/tot.1 58741 7990 3946 latn/nwt/tot.1 35027 5740 2948 latn/ock/tot.1 37263 5774 2996 latn/ock/tot.1 35027 5589 2926 latn/ptt/tot.1 96870 13946 7568 latn/ptt/tot.1 35027 6633 3875 port/csm/tot.1 64602 9032 5081 port/csm/tot.1 35027 6267 3772 russ/pic/tot.1 45915 11831 7936 russ/pic/tot.1 35027 9761 6659 russ/ptr/tot.1 111824 12032 5925 russ/ptr/tot.1 35027 5520 2910 russ/ptt/tot.1 111824 12034 5926 russ/ptt/tot.1 35027 5521 2911 span/qvi/one.1 177061 14247 7466 span/qvi/one.1 35027 5452 3237 span/qvi/tot.1 364837 22475 11175 span/qvi/tot.1 35027 5575 3385 span/qvi/two.1 187776 16023 8543 span/qvi/two.1 35027 5698 3558 tibe/ccv/tot.1 88620 1155 292 tibe/ccv/tot.1 35027 846 196 tibe/pmi/tot.1 143289 2932 666 tibe/pmi/tot.1 35027 1963 515 tibe/vim/tot.1 53287 1469 389 tibe/vim/tot.1 35027 1300 370 viep/mky/tot.1 39293 3471 1161 viep/mky/tot.1 35027 3341 1174 viet/nwt/tot.1 91019 2735 684 viet/nwt/tot.1 35027 2011 570 viet/ptt/gen.1 42099 1793 430 viet/ptt/gen.1 35027 1690 421 viet/ptt/num.1 37097 1485 363 viet/ptt/num.1 35027 1459 367 viet/ptt/tot.1 169480 2684 489 viet/ptt/tot.1 35027 1623 399 chin/ptn/lev.1 28693 1169 274 chin/ptn/lev.1 28693 1169 274 chin/ptt/deu.1 31494 1433 336 chin/ptt/deu.1 31494 1433 336 chin/ptt/lev.1 25694 1095 261 chin/ptt/lev.1 25694 1095 261 engl/cpn/tot.1 541 400 322 engl/cpn/tot.1 541 400 322 engl/cul/rec.1 8026 1464 770 engl/cul/rec.1 8026 1464 770 engl/wnm/tot.1 831 194 100 engl/wnm/tot.1 831 194 100 geez/eno/tot.1 17736 6274 4193 geez/eno/tot.1 17736 6274 4193 geez/gok/tot.1 34291 12272 8344 geez/gok/tot.1 34291 12272 8344 grek/nwt/joh.1 15919 2586 1422 grek/nwt/joh.1 15919 2586 1422 grek/nwt/luk.1 19887 4609 3015 grek/nwt/luk.1 19887 4609 3015 grek/nwt/mat.1 18745 3958 2350 grek/nwt/mat.1 18745 3958 2350 grek/nwt/mrk.1 11632 2898 1842 grek/nwt/mrk.1 11632 2898 1842 hebr/tav/deu.1 12007 5455 3972 hebr/tav/deu.1 12007 5455 3972 hebr/tav/exo.1 13870 5711 3882 hebr/tav/exo.1 13870 5711 3882 hebr/tav/gen.1 17211 7212 5100 hebr/tav/gen.1 17211 7212 5100 hebr/tav/lev.1 9650 3860 2609 hebr/tav/lev.1 9650 3860 2609 hebr/tav/num.1 13573 5306 3697 hebr/tav/num.1 13573 5306 3697 latn/ahl/tot.1 6534 1353 828 latn/ahl/tot.1 6534 1353 828 latn/nwt/joh.1 14026 2523 1377 latn/nwt/joh.1 14026 2523 1377 latn/nwt/luk.1 18004 4406 2743 latn/nwt/luk.1 18004 4406 2743 latn/nwt/mat.1 16431 3911 2278 latn/nwt/mat.1 16431 3911 2278 latn/nwt/mrk.1 10280 2913 1810 latn/nwt/mrk.1 10280 2913 1810 latn/ptt/deu.1 18502 4466 2815 latn/ptt/deu.1 18502 4466 2815 latn/ptt/exo.1 20060 4701 2790 latn/ptt/exo.1 20060 4701 2790 latn/ptt/gen.1 25217 5713 3485 latn/ptt/gen.1 25217 5713 3485 latn/ptt/lev.1 13775 3233 1909 latn/ptt/lev.1 13775 3233 1909 latn/ptt/num.1 19316 4340 2595 latn/ptt/num.1 19316 4340 2595 russ/ptr/deu.1 20988 3913 2238 russ/ptr/deu.1 20988 3913 2238 russ/ptr/exo.1 22960 4084 2112 russ/ptr/exo.1 22960 4084 2112 russ/ptr/gen.1 28445 4897 2703 russ/ptr/gen.1 28445 4897 2703 russ/ptr/lev.1 16901 2659 1305 russ/ptr/lev.1 16901 2659 1305 russ/ptr/num.1 22530 3952 2142 russ/ptr/num.1 22530 3952 2142 russ/ptt/deu.1 20988 3913 2238 russ/ptt/deu.1 20988 3913 2238 russ/ptt/exo.1 22960 4084 2112 russ/ptt/exo.1 22960 4084 2112 russ/ptt/gen.1 28445 4899 2704 russ/ptt/gen.1 28445 4899 2704 russ/ptt/lev.1 16901 2659 1305 russ/ptt/lev.1 16901 2659 1305 russ/ptt/num.1 22530 3952 2142 russ/ptt/num.1 22530 3952 2142 viep/grs/tot.1 31200 7760 3216 viep/grs/tot.1 31200 7760 3216 viet/nwt/jhn.1 21872 1289 428 viet/nwt/jhn.1 21872 1289 428 viet/nwt/luk.1 27637 2117 750 viet/nwt/luk.1 27637 2117 750 viet/nwt/mat.1 25615 1818 564 viet/nwt/mat.1 25615 1818 564 viet/nwt/mrk.1 15895 1572 556 viet/nwt/mrk.1 15895 1572 556 viet/ptt/deu.1 31361 1614 439 viet/ptt/deu.1 31361 1614 439 viet/ptt/exo.1 33760 1649 368 viet/ptt/exo.1 33760 1649 368 viet/ptt/lev.1 25163 1207 339 viet/ptt/lev.1 25163 1207 339 voyp/grm/tot.1 708 307 204 voyp/grm/tot.1 708 307 204 voyp/grs/tot.1 1950 635 365 voyp/grs/tot.1 1950 635 365 # END