[quote='rikforto' pid='80054' dateline='1771453968']null hypothesis explains quite well, however, why 主 appears 7 times to daiin's 5 and you had to recruit two more words to get a loose fit that doesn't click any other words in the paragraph into place.[/quote] The fact that daiin occurs 5 times in the longest parag may be compatible with the "null hypothesis". But that is not the evidence. The most general Null Hypothesis (NH) would be "there is no relation between the SBJ and the SPS". That hypothesis already fail to explain why the number of parags Np(SBJ) of the SBJ is within the estimated range of Np(SPS). But I admit that I picked the SBJ precisely because I knew that it had 365 parags. So the Null Hypothesis must be modified to "the SBJ is a text with Np(SBJ) ≈ Np(SPS) that is unrelated to the SPS". That NH still fails to explain why the minimum and average number of words per parag of the SBJ, Wmin(SBJ) = 8 and Wavg(SBJ) = 37, are so close to Wmin(SPS) = 6 and Wavg(SPS) = 34. I did not know those numbers when I picked the SBJ as the candidate for the ur-text of the SPS. But I learned those numbers last year, when I posted the parag-size histograms. If the numbers had been too different, I would have abandoned the SBJ. Since I kept insisting, lets take that as being cherry-picking too. So let's change the NH to "The SBJ is a text with Np(SBJ) ≈ Np(SPS), Wmin(SBJ) ≈ Wmin(SPS), Wavg(SBJ) ≈ Wavg(SPS) that is unrelated to the SPS" This new NH still does not quite explain why the max parag sizes Wmax(SBJ) = 92 and Wmax(SPS) = 73 are relatively close. The Wmax(SBJ) could have been anywhere from 35 to a thousand or more. But let's assume that the distribution of parag sizes or random medieval texts follow some special two-parameter distribution, so if the min and avg match, the max would roughly match too. Then let's give the new NH a pass on that test. In that longest parag of the SPS the string daiin occurs 5 times. That is just a fact, and there is no probability to be computed; it just is so. In the longest parag of the SBJ there is a Chinese character that occurs 7 times, 5 of them as a compound with another specific character. The next common character after those 2 occurs 3 times. I haven't computed the probability of that situation under the NH, but it seems likely enough given Zipf's law. So let's give the NH a pass on that too.