This is just from my recollection in the late 1990s, not from original documents of the research on readability.
The first readability formulae were based on a small set of passages that had been rated as appropriate for specific grade levels by a single person in the 1930s. That set of passages, and the associated grade level ratings continued to be the benchmark against which the validity of all readability formulae were measured lasting into the 1990s. Most existing formulae, therefore, continue to be based on a single person's ratings of grade level appropriateness from 70+ years ago.
There are two basic measures in the vast majority of readability formulae:
- Vocabulary load, typically measured as (1) the percentage of words in a passage that do not appear on lists of frequently used words, or (2) the average length of words in a passage (in letters or syllables). Other measures are used as well. Each of these options has its own set of problems.
- Syntactic complexity, typically measured as (1) sentence length (in words), or paragraph length (in sentences). Other measures are used as well. Each of these also has its own set of problems.
When I was working for the BYU English Language Center on an educational software product called SoftRead, I programmed over 30 different readability indices into the software. There were about as many different measures of vocabulary load and syntactic complexity as there were readability formulae.
For the stats geeks like me, the typical methodology, and an oversimplified methodology at that, is to run a linear or non-linear regression of the benchmark readabilities (the criterion for validation is always an expert human judgment of readability) on the measures of vocabulary load and syntactic complexity. One has to wonder why readability formulae are used at all for legal purposes, since the ultimate validation criterion is expert human judgment of grade level.
In very simplistic terms, this is using a statistical procedure to estimate what is already an estimate given by a human. Pretty difficult to defend.
There are some more sophisticated approaches to readability, which essentially use the statistical sub-specialty of psychometrics using cloze tests (passages with a few words strategically removed where readers have to supply or select an appropriate word to fill in the blank) where passages are rank ordered statistically to identify harder passages and easier passages. Some of the problems with this is that the selection of words removed is very important and still remains a subjective portion of the procedure. One pro of this approach is that it does not rely on human judgment of passage level, even though it does rely on human judgment of what words should be removed. I tend to be more accepting of this type of readability formula, though still quite skeptical.
I thought I was going to do my dissertation on readability, until I decided that it was all smoke and mirrors, and I had better not base my future career on something I thought was not sufficiently valid to have a strong place in education.
Some of the reasons I came to that conclusion were the following:
- The genre of reading (e.g. story, information, poetry, test items, instructions) is typically not varied in the development of readability formulae. It is typically narrative (story) passages that are used to create the formulae, but then the formulae are applied to all kinds of genres. This is a very strong assumption that the relationship of vocabulary load and syntactic complexity can be extrapolated from the sample of passages to a completely different type of passage.
- The literature indicates that high vocabulary load and syntactic complexity are less important in predicting comprehension than people's level of interest in the subject of the passage.
- The literature indicates that at best readability identifies a rough level of comprehensibility, but should not be used to guide writing. When readability indices are used to guide writing, very complex writing may be done with short words and short sentences, be nearly incomprehensible, and show up as a very low level text (and vice versa).
- The literature is rife with applications of readability formulae to short passages, yet the literature is also full of cautions against using readability formulae to evaluate the reading level of short passages.
At the same time, I did notice something interesting about the measures of readability. Almost all measures of readability use in their formulae measures of the average (the average word length, the average number of words a common word list, the average sentence length, the average paragraph length).
None of the formulae look at variation in these measures (e.g. the standard deviation of word length, sentence length, paragraph length). It seems to me that changes in vocabulary load and syntactic complexity across different portions of a text would also change the comprehensibility of a text. I never carried out the study I had planned to include these things because I became convinced that even if I could achieve better relationships between computed (estimated) readability and expert ratings using these new indicators, I still could not overcome the issues described above.
Despite this, I think readability results probably give some useful information. I just wouldn't want to use it for anything important. Using readability formulae to satisfy curiosity is probably safe. Any type of evaluation really seems to be taking it too far for me.
I wonder how readable this post is. D'oh!
Huh?
ReplyDeleteJust for kicks I plugged your blog address back into the thingy and it said junior high level. But when I plugged in the address on JUST this post.... Genius Level!!!
ReplyDeleteI had to send Elijah away just to be able to concentrate enough to grasp any of this.
ReplyDeleteThis increases my blog rating. Yes!
Wait...
ReplyDeletewhat?