Compelling, comprehended reading is one of the most powerful and best-researched paths to proficiency. But what makes a text comprehensible?
A lot of things go into a given learner’s comprehension of a text. Many of those things are hard to measure, but one objective statistic stands out as especially useful: unique word count versus total word count, or how many different vocabulary items appear in a text compared to the overall length of the text.
Why Does Unique Word Count Matter So Much?
Picture two stories that each have a total word count of 1000. One of them has 400 unique words and the other has 100 unique words. The story with 400 unique words spread across 1000 total words is likely to contain many words that appear only once or twice in the story, meaning that a.) those words might not be very common words in the language and b.) readers don’t get enough exposure to those words to acquire their meanings.
By contrast, in the story with 100 unique words, each word is repeated, on average, 10 times. (In practice, some words will be repeated many more times, and some may be repeated just a few times.)
Which story is likely to help a student read more fluently and retain word meanings long-term?
Even though someone reading the second story will meet fewer different words, they will significantly boost their reading fluency and will “lock in” their understanding of more words and expressions, making learners more likely to be able to use these in their own speech and writing. They will also read the story much faster than they would have read the first story. All of this has other positive effects for our students and our classes, such as increased confidence on students’ parts and increased chances that they will want to read more.
By extension, unique word count is important for authors and publishers. Texts with low unique word counts are highly prized, partly because there are more novice and early intermediate learners in schools than there are later intermediate and advanced learners, and partly because texts with low unique words are harder to find “in the wild,” where there is practically an infinity of texts with high unique word counts.
But there are two practical problems:
-
The fewer unique words one is willing to use, the harder it is to write an interesting text.
-
Unique word count is surprisingly hard to determine.
For this post, we’ll leave the first problem to authors and focus on the second one, along with its implications for World Language teachers and learners.
Why Is Unique Word Count Tricky to Determine?
It seems like it shouldn’t be hard to figure out how many different words appear in a text. There are free sites where you can paste in a text and have it report what it considers the unique word count to be. But a look at almost any paragraph of any text reveals that the question is much more complex. For instance, how many unique words are in the following 16-word sentence?
Before I went to the University of Michigan, I wanted to go to two other universities.
Did you say 11? Did you say 14? Something in between? There are 13 differently spelled items in the sentence. But some are different forms of the same word (went/go, University/universities), while other items that are spelled the same don’t mean the same thing—check out the word to on either side of the word go.
If the point of calculating unique word count is to give a sense of how easy or difficult it might be for a learner to understand a text, is it helpful or unhelpful to count go and went—which don’t have a single letter in common—as the same word, based on the fact that they are both forms of the word [to] go? What if it were go and goes? What about goes and going?
Similar complications can arise in any language. Here are just a few other cases in languages commonly taught in the United States:
English
- run–ran (one is the past tense of the other, but this isn’t a common pattern—fan isn’t the past tense of fun, nor ban of bun or pan of pun.)
- ring–rang–rung (and don’t forget about ring [jewelry on a finger] and rung [step on a ladder]!)
- beat (present tense) – beat (past tense)
- too (also) – too (excessively)
Spanish
- doy – das – da etc.
- ir – va– iban – fuimos – fuese – yendo (!)
- oro (gold) – oro (I pray)
- nada (nothing) – nada (swims)
French
- (je) parle – (tu) parles – (il/elle) parle etc.
- aller – vont – allant
- œil – yeux
- la pêche (peach) – la pêche (fishing)
German
- geht – ging – gegangen
- bin – ist – war
- Maus – Mäuschen
- Bank (bench) – Bank (place that deals with money)
Latin
- dīcō – dīcis – dīcit etc.
- fert – tulit
- lapis – lapidem
- quod (interrogative) –quod (relative) – quod (conjunction)
So, if unique word count is hard to determine, but significant for learners—and for teachers selecting books for their students—how can we proceed?
Three Possible Approaches
From strictest to loosest:
Option 1
For two words in a text to count as the same word, they must have the same spelling and the same meaning. In this system, run and run are two different words if in one instance it means “move fast on one’s feet” and in the other it means “lead” (as in “run a company”).
Option 2
For two words in a text to count as the same word, they must at least have the same stem and have basically the same root meaning, though they may represent a different number, gender, person, tense, or mood. In this system, actor and actresses are counted as the same word, as are go, going, and gone (except when gone means “depleted”).
Option 3
For two words in a text to count as the same word, they must pertain to the same word family. In this system, not only would see, sees, and saw count as one unique word, but that same “word” would also include sight, insight, overseer, and sightseeing.
As you can imagine, these three approaches can yield rather different unique word counts. This means, among other things, that a book advertising 130 unique words might actually have more than another book advertising 170 unique words, if the publisher of the first book used Option 2 or 3 and the publisher of the second used Option 1.
If you’ve guessed that it gets even more complicated, you’re right.
Cognates
Most publishers (and many independent authors) of language learner literature don’t include cognates in the official unique word count. In other words, they determine the unique word count, usually using Option 2 or 3 above, and then subtract all the words whose meanings they think will be clear to learners due to similarities to words in another language that the learner knows. For instance, if Option 3 yields a unique word count of 280 for a Spanish book, and the publisher thinks that the meanings of 50 of those words will be pretty obvious to English speakers, the back cover of the book will often say 230 unique words.
To understand the major problems with this approach, we need to understand a few things about cognates.
Cognates: Definition
Technically, any words with a common ancestor are cognates. Sometimes, the connection is clear. For instance, the Latin word interesse (“to be between” and thus “to make a difference”) is the source of English interesting, French interessant, German interessant, Spanish interesante, Swedish intressant, and Russian interesnyy (интересный), even though not all these languages themselves are descendants of Latin. (Note: the Latin word interesse itself is not considered a cognate of the words it gave rise to, just as your parent is not considered your sibling. In case you’re interested[!] in the terminology, linguists call interesse the etymon [from the Greek for “true” or “actual”] of all these words.)
Cognates: Complications
But many words that don’t look that similar are cognates, too. Bafflingly for non-linguists, English heart, Latin cor, and Russian serdtse (сердце) are all cognates, derived from the same Indo-European root. Those words at least all mean the same thing; many cognates do not. This is true both for cognates that look near-identical and for ones that don’t. German Gift means not “gift” (present) but “poison,” though both words have the same root. (It’s a long story.) Spanish constipado means “having a cold”—a different kind of being stuffed up. Meanwhile, Spanish hermano (“brother”) and English germane (“pertinent”) are cognates, both derived from the Latin germānus (“having the same parents”), but knowing one doesn’t mean you’ll understand the other in context. The English words chapter, cape, capital, chef, and chaperone are cognates of each other, all from Latin caput (“head”), as are the Spanish cabeza (“head”) and jefe (“boss”)—and, yes, this means that the Spanish word jefe and the English word chapter are cognates—but knowing one doesn’t give you the meaning of the others.
Then there is the opposite: words that look the same or similar, but are not cognates. Usually, these have different meanings. Spanish sed (“thirst”) is unrelated to Latin sed (“but”). French blesser means not “to bless” but “to injure” and has a different origin. Fascinatingly, there are cases where non-cognate lookalikes happen to have similar meanings. You may be surprised that the Spanish word mucho is historically unrelated to its English translation “much.” The two words come from different Indo-European roots and their similar spelling is a handy coincidence.
Cognates: Implications
This leads us to at least five problems with excluding cognates from unique word counts, the first four of which arise from the above complications:
-
It’s often unclear to teachers or even to authors and publishers what constitutes a cognate.
-
Even if a target-language word is truly a cognate of a word in a language that the publisher has in mind (such as English), a person reading the text may not speak that language.
-
Even if a target-language word is truly a cognate of a word in a language that the reader does know, the reader may not recognize it as such.
-
Even if a reader recognizes that a word in the text is a cognate of a word they know, the in-context meaning may not be clear.
The fifth problem is that, even if the reader understands every single one of the cognates that the publisher has subtracted from the unique word count, the resulting number is still misleading as an indicator of how much repetition is in the text, and therefore of how useful the text is to a learner.
Imagine a 200-word story with 50 unique words, but 30 of these are cognates, so the publisher advertises 20 unique words. This number suggests that each word appears in the text an average of 10 times (200/20). But, in fact, each word in the text occurs an average of only four times (200/50), meaning less than half the exposure to the 20 “new” words than the advertised unique word count suggests.
A Solution: Transparency
Because there is no one correct method of determining unique word count, what ends up mattering the most is clarity about what method is used for a given text, along with any related information that could help a teacher or learner gauge the suitability of the text. Fortunately, there are some excellent models of this. Adriana Ramírez’s María Cano, la flor del trabajo has this on the back cover:
Esta historia está escrita para estudiantes de español de nivel dos. Tiene 395 familias de palabras, de los cuales 145 son cognadas.
This story is written for students in Spanish level two. It has 395 word families, of which 145 are cognates.
Level 2 can mean different things to different people, but the concrete information about the number of word families and cognates is helpful, especially when there are other books that might simply have reported 250 different word families (395 minus the 145 cognates), and some that would have just said “words” instead of “word families,” while still counting different forms of a word as being the same word. It would be great to know the total word count, so as to estimate the amount of repetition, but one can roughly guess the total word count by counting the number of words on one page, multiplying by the number of pages containing text, and rounding down a bit because some pages are only partially full of words.
Publishers and authors could achieve even more transparency by describing in detail how they determine unique word count, and explaining why they use their chosen method. This explanation could be accessed through a link or QR code in the front matter of each book.
In the meantime, don’t be afraid to contact a publisher or author to request not only information about how they calculate unique word count, but also any totals that they may not have included in or on the book, such as the total word count.
You can read about Teacher’s Discovery’s approach to word counts and reader levels in this guide.