Typological analysis of word order in phrases
Word order in a phrase is the linear sequence of grammatical elements within a noun, verb, or prepositional phrase that a language establishes as the primary order. Typological analysis compares such sequences across hundreds of unrelated languages and searches for recurring relationships: if a language places an object before a verb, how likely is it to have postpositions rather than prepositions? Answers to these questions were provided by two seminal works — Joseph Greenberg’s 1963 article and Matthew Dreier’s corpus studies, as well as the WALS database, which covers over 1,300 languages.
2 Conceptual apparatus
3 Frequencies of orders in world languages
4 Greenberg correlations and their testing
5 The Dreyer Method: Genders instead of languages
6 Branching direction theory
7 Adjective and noun
8 Order in a verb group
9 Procedural explanations
10 Exception languages
11 Flexible order and morphology
12 Areal and genealogical distortions
13 Testing on the cases
Sources: Greenberg’s sample
In 1963, American linguist Joseph Greenberg (1915–2001) published " Some Universals of Grammar with Particular Reference to the Order of Meaningful Elements," in which he formulated 45 linguistic universals. Of these, 28 concern the relative positions of syntactic units — that is, the order of elements. The formulations were implicational: not "all languages are structured this way," but "if a language has property X, then property Y occurs with greater than random frequency."
The sample was based on 30 languages: Basque, Serbian, Welsh, Norwegian, Modern Greek, Italian, Finnish, Yoruba, Nubian, Swahili, Fula, Maasai, Songhai, Berber, Turkish, Hebrew, Burushaski, Hindi, Kannada, Japanese, Thai, Burmese, Malay, Maori, Loritja, Mayan, Zapotec, Quechua, Chibcha, and Guarani. By today’s standards, the sample is small. By the standards of 1963, it was bold: for the first time, word order was measured in languages from all continents, not just Indo-European ones.
Greenberg knew the limits of his base. Of the 45 universals, he declared 33 to be valid without exception; the rest were statistical tendencies, where violations were possible but rare. The very phrase "with overwhelmingly greater than chance frequency" became a textbook example and was incorporated into Universal #4 on postpositions in SOV languages.
Conceptual apparatus
Typology works with pairs of elements: subject and predicate, object and verb form, attribute and noun, possessive marker and noun. Each pair is assigned an order — for example, OV for "object + verb" and AdjN for "adjective + noun." The English " large dogs" yields AdjN, while the French " chiens grands" yields NAdj.
The central distinction is between phrase order and intraphrase order. The sentence "the dog saw the cat" is described by the SVO schema at the clause level, but its noun phrases conceal their own schemas: an article before a noun, an adjective before a noun, and a dependent noun before a principal noun. Greenberg and his followers found that these levels correlate less strongly than expected, and part of the discipline grew out of attempts to explain precisely this gap.
A separate issue is "dominant order." Many languages allow multiple sequences, and then the dominant order is considered either the only possible one or the statistically dominant one. WALS uses a rule of thumb: an order is considered dominant when it occurs more than twice as often as the alternative; any less than that is considered dominant, and the language is declared to have no dominant order.
Designations and abbreviations
Standard notations: S is the subject, O is the direct object, V is the verb, G is the genitive phrase (possessor), A is an adjective, N is a noun, Adp is an adposition without specifying the side. The combinations are read from left to right: AdjN means "adjective before noun," Pr are prepositions, Po are postpositions. For verb phrases, notations like AuxV and VAux are accepted, for prepositional phrases, AdpN and NAdp.
Frequencies of orders in world languages
The WALS database for feature 81A provides a distribution of six logically possible orders S, O and V across 1376 languages:
| Order | Number of languages |
|---|---|
| SOV | 564 |
| SVO | 488 |
| VSO | 95 |
| VOS | 25 |
| OVS | 11 |
| OSV | 4 |
| No dominant order | 189 |
Together, the two orderings — SOV and SVO — cover more than two-thirds of the sample; all three patterns, where the object comes first, are rare. Greenberg’s Universal #1 ("in declarative sentences with a nominal subject and object, the subject almost always precedes the object") is clearly confirmed by these figures.
The picture is different for adjectives and nouns. WALS covers 1,367 languages based on feature 87A: AdjN occurs in 373, NAdj in 879, and in another 110 languages, both orders are used without a clear advantage. The "noun + adjective" order is more than twice as common as the reversed order. This is surprising to the school-age intuition of a European accustomed to the previous definition.
The order of the exponent and adjective divides languages almost equally: in 227 languages, the exponent precedes the adjective, while in 192 languages, it follows it. This even distribution itself argues against the universal strength of this correlation.
Greenberg correlations and their testing
Greenberg’s implicational universals linked clause order with orders within phrases. Let’s give verifiable examples.
Universal #3: Languages with a dominant VSO always have prepositions. Universal #4: Languages with a normal SOV overwhelmingly have postpositions. Universal #5 restricts the condition: if in an SOV language the genitive phrase follows the main noun, then the adjective follows the noun. Universal #25 links the levels of analysis within a clause: if the pronominal object follows the verb, then the nominal object follows it.
The general pattern, later statistically formalized, is as follows: OV languages typically pair with postpositions, while VO languages pair with prepositions. The pairs "genitive + noun" and "relative clause + noun" behave similarly to adpositions: in OV languages, the dependent clause usually precedes the principal clause.
Tests on larger samples yielded a mixed picture. A team that ran Greenberg’s universals against the tree banks of the Universal Dependencies project (over 70 languages) confirmed universals #19 and #25, while others yielded weaker support. The current literature formulates a cautious conclusion: most of Greenberg’s implications are statistically significant, while a few are not.
The Dreyer Method: Genders instead of languages
Directly increasing the sample size creates a trap. Three hundred languages of Niger – Congo or Australia may reflect not a universal pattern, but a common heritage or contact. Matthew Dryer, in a 1992 paper, proposed a solution: count not languages, but taxonomic genera — groups of closely related languages. His sample consisted of 625 languages, distributed across genera and macroareas.
The logic is simple at the everyday level. If forty languages share a single ancestor with postpositions, then forty "votes" for postpositions are effectively one vote. Gender as a unit of record mitigates genealogical imbalances; areal stratification mitigates regional imbalances. Only after such a cleansing can correlations be considered evidence of the structure of a language in general, rather than the fate of specific families.
Using this sample, Dreyer determined which pairs of elements correlate with the verb-object order and which do not. Adpositions, the genitive, and the relative clause were among the correlated pairs. Adjective-noun order was not.
Branching direction theory
To explain these correlations, Dreyer proposed the Branching Direction Theory. It is based on a distinction between "verb patterners" and "object patterners": elements of the first group behave like a verb, while the second act like an object. Verb patterners are single words that do not form a phrase; object patterners are entire phrases that can grow.
The theory states that languages tend to branch consistently in one direction — either entirely to the left or entirely to the right. Mixed branching is difficult to handle, and frequency correlations reflect this difficulty. An OV language with postpositions branches to the right uniformly; English with SVO and prepositions also branches to the left uniformly. A language that mixed types would be more expensive for a parser.
The theory has a weakness. In flat component structures, without deep nesting, some correlations cannot be derived from it — the order of the article and noun, for example, requires a separate explanation, as proposed by John Hawkins. Critics have also noted another point: the theory does a worse job of describing where correlations exist than it does of explaining why they exist.
Adjective and noun
The most famous negative result of word order typology concerns the adjective. A common belief associated AdjN with OV languages and NAdj with VO languages. In his works of 1988 and 1992, Dreyer demonstrated that this is not the case: NAdj is more common than AdjN among both OV and VO languages.
The WALS results for feature 97A confirm this conclusion. The combination OV and AdjN is found in 216 languages, OV and NAdj in 332, VO and AdjN in 114, and VO and NAdj in 456; another 198 languages do not fall into any of the four cells. All four logically possible types are frequent. There is no correlation between the features — there is only a general preponderance of NAdj across the planet.
So where does the myth of AdjN’s connection to the OV order come from? Dreyer provides an areal answer: AdjN is more common among OV languages in Eurasia, and early samples, overloaded with Eurasian languages, mistook this regional feature for a universal one. Turkic, Mongolian, Tungusic, Hindi, and Japanese — all with AdjN — created the impression of a pattern that does not exist outside the continent.
An example of Udege
In the WALS materials, the "OV and AdjN" type is illustrated by Udehe, a Tungusic language of Siberia, in which both the object precedes the verb and the attribute precedes the noun. This example is indicative precisely because of its areal location: a language from the Eurasian zone of AdjN orders, not from Africa or America, where NAdj would predominate under the same typological conditions.
Order in a verb group
A verb phrase is more compact than a noun phrase. Auxiliary verbs tend to be positioned at the edge of the phrase: before the main verb in VO languages, after it in OV languages. Negative markers, modal particles, and adverbs of manner are attached to the verb according to the same rules of lateral placement.
Adpositions are the most consistent part of the picture. The association of OV with postpositions and VO with prepositions is confirmed in Dreyer’s sample and in WALS. A historical explanation suggests working with etymology: adpositions in many languages originate from nouns, so the order of adposition naturally follows the order of the genitive phrase. The English preposition " in front of" was grammaticalized from a nominal construction; routes of this kind were repeated independently in different families.
Greenberg’s Universal #4, formulated as a tendency, has survived tests better than many categorical ones. It’s the categorical formulations that suffer the most: VSO languages indeed always use prepositions, while postpositions are only statistically dominant among OV languages.
Procedural explanations
The statistics require explanation. Why does a human parser prefer uniform branching and proximity of dependent elements?
John Hawkins developed a two-pronged answer. The principle of head proximity describes the tendency of languages to avoid separating adjectives and genitive phrases between the head noun and the clause verb. In terms of classification, languages are divided into V-initial, SVO, and SOV, and the prohibition on "sandwiching" dependents between heads explains why the possible noun phrase orders are so uneven.
The second approach is the Minimize Domains principle. The idea is intuitively simple: if two elements share a syntactic or semantic dependency, processing that dependency is facilitated by their proximity; word order tends to minimize the average distance between dependent pairs. Heavier components are moved to the end of the group — hence the well-known English "heavy NP shift," familiar to anyone who has read phrases like "gave to the committee the document everyone had been waiting for."
Process explanations are attractive because they are testable: they predict gradient effects in psycholinguistic experiments, not just table frequencies. Critics object that the convenience of parsing alone does not explain why languages remain in awkward states for centuries. The answer lies in diachronic mechanisms: phrase order is transmitted through grammaticalization, and the legacy of this transmission is stronger than the moment.
Exception languages
Exceptions are the typology’s working material. German combines OV order in subordinate clauses with prepositions and, to some extent, NAdj noun structure — a mixed profile that branching theory would predict to be unstable. Persian maintains the verb at the end of clauses with prepositions and layered adjectives. These languages have been around for centuries, and no parser can "correct" them.
Among languages without dominant order, WALS has 189 positions, and Russian grammar is among them. The relative order of subject, object, and verb in Russian sentences is free: case endings take over the job that position does in English. "The cat ate the fish," "the fish was eaten by the cat," and "the cat ate the fish" differ in communicative structure, not in the identity of the participants in the event.
Morphology and word order act as interchangeable tools for marking relationships. A rich case paradigm reduces the need for rigid linearity; rigid linearity compensates for the paucity of endings. The Russian "lisya nora" and the English "fox’s den" solve the same problem through different means — the former through case, the latter through position and the possessive morpheme.
Flexible order and morphology
Freedom of order is not chaos. In languages with discernible case marking, Russian sentences are ordered according to the rules of theme and rheme, not syntactic position. In Australian and many American languages, clauses are rearranged for similar informational reasons, with an almost complete absence of fixed patterns.
WALS captures the "double dominance" phenomenon as a separate feature: 29 languages exhibited a combination of SOV or SVO as equal base orders, 14 exhibited VSO or VOS, and 13 exhibited SVO or VSO. This mixed behavior is not a sampling flaw, but a type of linguistic organization with its own logic.
Flexibility is also characteristic of the phrasal level. The Russian adjective appears before the noun by default and after it in terminological and poetic contexts; the WALS order "both without preponderance" is noted in 110 languages based on the adjective feature. The position of the attribute often carries the semantic load of distinguishing between a permanent and a temporary feature, as seen in pairs like "a sick person" and "a sick person."
Areal and genealogical distortions
Any frequency table of the world’s languages incorporates biases of two kinds. The first is genealogical: the 1,376 languages in WALS are unevenly distributed across families, and large families can triple the weight of a single ancestor solution. The second is areal: neighboring languages copy each other’s structures over centuries, and entire continents acquire common features without any shared kinship.
The case of AdjN in Eurasia is a textbook example of the second kind. If the sample had been taken only from languages from the Atlantic to the Sea of Japan, the conclusion about the connection between OV and AdjN would have seemed firm. African and American OV languages undermine this connection. Dreyer’s method of genera was designed precisely to counter such pitfalls.
The area-based approach also yields positive results. A comparison of areas revealed that the combination of SVO with prepositions and NAdj dominates in Africa, SOV with postpositions dominates in Eurasia and large parts of the Americas, and verb-initial orders are concentrated near the Pacific Ocean and in Africa. The frequency map itself is a fact about human migrations and contacts, read through grammar.
Testing on the cases
The current stage of verification moves analysis from questionnaires to annotated corpora. Universal Dependencies tree banks annotate syntactic dependencies in texts in dozens of languages using a unified scheme, allowing Greenberg’s universals to be tested in real usage rather than in grammar descriptions.
The initial results are moderately encouraging. Universals #19 and #25 have been empirically confirmed across more than 70 languages. Others have performed differently: some implications hold, while others are eroded by the vibrant variability of texts. Corpus-based testing is also important because it measures dominant order by corpus frequencies rather than informant judgment, and the "twice as often" threshold from WALS becomes a directly calculable value.
The quantitative approach has also taken shape theoretically: work on implication universals is moving toward quantitative ones, where the strength of a connection is expressed numerically rather than in a binary "confirmed or refuted" sense. For phrasal orders, this means moving from the question of "does AdjN correlate with OV" to the question of "how weakly." The answer to the second question is already known: so weakly that in a sample of 1,316 languages, all four possible combinations are frequent.
- “Размышления о первой философии” Рене Декарта, краткое содержание
- Что означает слово “инфографика”
- Игорь Дрёмин: Слова и вещи
- Современные концепции красоты: философский анализ
- “Возвращённый рай” Джона Мильтона, краткий анализ
- “Мы в порядке” Нины Лакур, краткое содержание
- Кристофер Вул: Картины со словами
- Таинственный Рембрандт: рентгеноструктурный анализ выявил детали скрытой картины
- Несколько слов о портрете
You cannot comment Why?