Zipf's law is an empirical statistical pattern observed in many kinds of ranked data, most famously the frequency with which different words occur in a natural-language text, stating that the frequency of any given item is roughly inversely proportional to its rank in the frequency table: the most common item occurs about twice as often as the second most common, three times as often as the third most common, and so on. The law is named for linguist George Kingsley Zipf, who documented and popularized the pattern extensively across languages in his 1949 book Human Behavior and the Principle of Least Effort, though similar earlier observations of comparable rank-frequency patterns had already been made by others, including French stenographer Jean-Baptiste Estoup and German physicist Felix Auerbach. Beyond word frequency, closely similar rank-size patterns have been found to hold, at least approximately, across a range of other domains, including the population sizes of cities within a country and the sizes of business firms within an industry. The precise underlying causal mechanism that produces this recurring pattern across such varied domains remains a subject of ongoing theoretical debate, with proposed explanations ranging from simple random text-generation processes to preferential-attachment and least-effort optimization arguments.
Connections
Reader Challenges (0)
No disputes yet. Spotted an error or a better source? Open the first one.
Sign in to dispute this or suggest a correction.