Is a Small Change Evidence of Anything? Measuring the Indus Script
A side project with ChatGPT, Claude and Gemini on the undeciphered Indus script: why a small change in a popular statistic is not evidence until you calibrate it, and three headlines we took back.
Igor Ramos ·

The Indus Valley civilisation left thousands of short inscriptions, mostly on small stamp seals like the one above, and nobody can read them. There is no bilingual text, no Rosetta Stone, and nobody knows what language they record. The inscriptions are also tiny: the corpus we used has 2,543 of them, averaging 4.44 signs each.
So people turn to statistics. The best-known example is a 2009 paper in Science (Rao et al.) that measured the script’s conditional entropy and argued that it sits with natural languages rather than with non-linguistic symbol systems. Others have argued that the script may not encode language at all (Farmer, Sproat & Witzel, 2004), and that statistics like this cannot tell language apart from non-linguistic symbol systems (Sproat, 2014).
With three AI collaborators, I took on a narrower question. We deciphered nothing. What we found is smaller, and I think more useful.
- Apply the documented differences between two Indus sign catalogues and conditional entropy moves by less than 1%. That looks like robustness. It isn’t: two large effects cancel.
- At this corpus size, more than half of the measured “structure” between signs also shows up in shuffled data that has none.
- Checked against corpora where the answer is known, the real corpus lands in between: taken together, the variants the catalogues disagree about are partly, but only partly, predictable from their neighbours.
- This is a first draft, not peer reviewed. The code and data are on GitHub.
What we measured
Conditional entropy answers one question: once you’ve seen a sign, how uncertain are you about the next one? It’s measured in bits. Zero bits means the next sign is certain, one bit is a coin flip, and two bits is a pick among four equally likely options.
Average that uncertainty over every sign in a corpus and you get a single number. Rao et al. compared the Indus number with those of known languages and non-linguistic systems.
Its twin is mutual information: how many bits knowing one sign saves you when you guess the next. It is the uncertainty you would have with no clue, minus the uncertainty left once you know the previous sign. The two are tied together by that equation, but as you’ll see, they behave very differently here.
The question
Every one of those statistics is computed over a sign list, which is a decision about which marks count as the same sign. The sign lists disagree. Mahadevan (1977) records roughly 417 signs, and Wells (2015) considerably more. Much of the gap comes down to allographs: whether a small variation in a mark makes a different sign, or is just a different way of writing the same one.
So: if you apply the documented disagreements between catalogues, how far do the statistics move, and does that movement mean anything? It matters because a statistic offered as evidence about the script should be measuring the script, not the catalogue, so the first thing to know about such a number is how much of it the catalogue can move.
We used an open digitisation that maps each sign in Asko Parpola’s catalogue to the Wells signs it groups together. Of those groupings, 67 “merge families” can be tested in our corpus (the other nine have fewer than two variants that actually occur in it). Merging a family means relabelling all of its variants as one sign, the way Parpola’s catalogue groups them, and recomputing. Merging all 67 collapses 208 sign types into 67 and shrinks the vocabulary from 592 signs to 451.
The obvious answer, and why it’s wrong
Merge everything and conditional entropy moves by under one percent (−0.91%). The first headline wrote itself: the statistic is robust to catalogue choice.
It isn’t. A merge does two things at once. It makes the next sign easier to guess, because there are fewer options to choose from. And it throws away clues, because two signs that predicted different things become one. The two effects pull in opposite directions:
The small number was cancellation, not stability.
Even that number is fragile. Read the pairs in the other direction and the change is −0.28%. Divide by the largest value it could take for this vocabulary (log₂ of the sign count, a common normalisation), and the sign flips to +3.51%. Mutual information is immune to both, which is one reason we switched to it.
The switch gave us a better-looking headline: the merge destroys about 8% of the information between neighbouring signs, while conditional entropy barely moves.
Then our adversarial reviewer, an AI agent whose only job was to find what’s wrong with our results, shuffled the corpus. It scrambled every sign across the corpus while keeping every inscription’s length and every sign’s frequency, which destroys all sequential structure. The result:
Why would data with no structure show any? Because the corpus is small for its alphabet. With 592 signs there are about 350,000 possible pairs, and we observe only 8,737. Most pairs turn up once or never, so chance coincidences look like rules. Statisticians call this estimator bias. Here, data with no structure at all reproduces more than half of what looks like dependence between signs.
Merging shrinks the alphabet, so it shrinks the bias as well as any real structure, and the raw number can’t tell you which one it removed.
What works: calibrate it
If the raw change can’t be read on its own, compare it with corpora where you know the answer. It’s the same idea as checking a kitchen scale with weights you already know. We built three reference corpora with the same number of inscriptions, the same lengths and the same pair counts, and applied the identical merge to each:
- Context-free allographs: within each family, which variant appears is random, so the variants carry no information.
- Context-determined variants: which variant appears is fixed by the sign before it, so the variants carry real information.
- Shuffled: no order at all, as above.
As a control, we also applied random merges of the same shape (same number of families, same sizes, signs of similar frequency) to the real corpus.
Two results come out of this.
First, the same merge moves conditional entropy anywhere from −2.9% to +4.4%, depending only on the corpus it is applied to. An uncalibrated change is not evidence of anything.
Second, the real corpus sits between the two constructed extremes and close to neither. If the documented families were true allographs, conditional entropy would have fallen about three times further. If the variants were set by their context, it would have risen. The catalogue’s merge families are partially context-distinguishable: the variants aren’t interchangeable, but they aren’t fully determined by their context either. Put plainly, which variant turns up is partly, but only partly, predictable from its neighbours.
One comparison survives the bias problem, because both of its sides suffer the same bias: the documented merges against the random merges, both shrinking the vocabulary from 592 signs to 451. The documented merges lose markedly less mutual information: about 0.26 bits, against 0.44 for the random ones. That gap is 9.8 standard deviations of the simulated random baseline (1,000 draws, so the p-value floors at 0.002). It persists, with a similar effect, after duplicate inscriptions are removed.
In plain terms, the catalogue groups signs that tend to turn up in similar places, not arbitrary ones. That is evidence the grouping tracks something real, but it doesn’t prove the grouped marks are the same sign (more on that below).
What this means
The contribution is a method, not a verdict on the Indus script. For anyone computing statistics on a small corpus of undeciphered signs, it comes down to three cheap checks:
- Calibrate before you interpret. Run the same change on corpora where you know the answer. Here the same merge produced anything from −2.9% to +4.4%.
- Measure the floor. Shuffle the corpus and see how much “pattern” survives. Here it was 57% of the mutual information.
- Say which way you read and how you normalised. Reversing the pairs changed the result here, and normalising flipped its sign.
For the Indus debate specifically: how a conditional-entropy figure responds to the choice of sign catalogue can’t be read without calibration. Such figures should come with that sensitivity measured, not just the catalogue named.
What this doesn’t show
- It is not a decipherment. Nothing here says what any sign means.
- It is not a refutation of any published work.
- It is not proof that any two marks are the same sign. When merged signs cluster in one slot of a formula, “two spellings of one sign” and “two different values that alternate in that slot” (numbers, say, like the stroke signs above) look identical to these statistics.
- The evidence is narrow: one corpus, one merge set and one digitisation. That digitisation is one person’s work in progress, not a validated concordance.
What’s next
- Compare with a recent preprint (Nair, 2026) that very likely uses the same underlying data.
- Put the “shared slot” reading in front of Indus epigraphers, who may already know it well.
- Get a DOI and turn the draft into a formal preprint.
- Check the catalogue’s decisions against the marks themselves. That needs better images than exist today: in our small check of the published plates, many of the relevant marks were too unclear to classify.
If you work on Indus epigraphy, I’d especially like to know whether the “shared slot” reading, that these merge families are marks alternating in one slot of a formula, is already old news to you. Find me on LinkedIn.
Three headlines we took back
This is the part I learned the most from. The project retracted three headline claims, and each one looked like a result until someone tried to break it.
- “The statistic is robust to catalogue choice.” The small change came from two large effects cancelling out.
- “The merge destroys about 8% of the sequential information.” The shuffle test showed that number can’t be separated from estimator bias: a corpus with no structure at all loses more. It was withdrawn.
- “Five specific merge families are statistically significant.” The random baseline for a two-sign family has only 172 to 782 possible configurations, so a reported p-value of 0.0002 was below anything the baseline could actually express. When we enumerated every configuration, the flagship example failed. All per-family results are withdrawn.
One of the reviewer agents also caught a bug nobody had asked it to look for. Wells sign IDs are zero-padded (W017) and the corpus codes are not (G17), so a naive conversion silently dropped 52 of 498 signs and nine testable families.
Who did the work
Me and three AIs, which I call Chatinho, Claudinho and Geminho: Brazilian Portuguese for “little Chat,” “little Claude” and “little Gemini.” (Brazilians will notice that chatinho also means “a bit annoying.” For the one whose job was to criticise us, that’s a compliment.)
- Claudinho (Claude, Anthropic) did the programming, the data preparation, the statistical design, an independent numerical checker and the drafting. It also ran the adversarial reviewer agents that caught its own mistakes.
- Chatinho (ChatGPT, OpenAI) was the outside critic. It pointed out that a small net change can hide large opposing terms, brought in the data-processing inequality that bounds how much information a merge can really destroy, and insisted that the random baselines match each sign’s frequency.
- Geminho (Gemini, Google) provided the opening research directions and the sensitivity-analysis framing we adopted. Its literature summaries also contained errors, including a corpus-size figure that turned out to be a misreading of a different statistic, so we didn’t rely on them.
I directed the work, refereed when they disagreed, and I’m responsible for every claim here. The most valuable part of the whole project was an AI whose only job was to prove the rest of us wrong.
Status
This is a version 1 working draft. It has not been peer reviewed or submitted anywhere. The code, the data we’re allowed to share, the draft paper and the full retraction log are on GitHub (GPL-3.0). A standalone checker, written without importing the analysis code, recomputes the baseline entropies, the mutual information and the cancellation from the raw inputs; separate scripts regenerate the calibration and the random-baseline comparisons. Questions and corrections are welcome on LinkedIn.
Sources
- Code and data: isramos/indus-script-sign-merging (GPL-3.0).
- Corpus:
indus_website_real_corpus.csvvia joyboseroy/indus_decipher, derived from yajnadevam/indus-website. We use only the sign sequences; nothing here depends on, or evaluates, any decipherment claim. - Merge families: mayig/indus-valley-script-corpus (MIT).
- Seal photograph: The Metropolitan Museum of Art, 49.40.1, Open Access (public domain).
- Farmer, S., Sproat, R., Witzel, M. (2004). The collapse of the Indus-script thesis. Electronic Journal of Vedic Studies 11(2).
- Mahadevan, I. (1977). The Indus Script: Texts, Concordance and Tables.
- Nair, A. (2026). How non-linguistic is the Indus sign system? A synthetic-baseline scorecard. arXiv:2604.17828.
- Parpola, A. (1994). Deciphering the Indus Script. Cambridge University Press.
- Rao, R. P. N. et al. (2009). Entropic evidence for linguistic structure in the Indus script. Science 324(5931).
- Sproat, R. (2014). A statistical comparison of written language and nonlinguistic symbol systems. Language 90(2).
- Wells, B. (2015). The Archaeology and Epigraphy of Indus Writing. Archaeopress.