Chapter 15.5 PIE Morphology

The basic morphemic structure of PIE has many similarities with Modern English: a root
morpheme to which is added various derivational and then inflectional morphemes. The major
difference is the productive use of what we call ablaut or gradation,
a technique the traces of which are still found in Modern English, though no longer in
productive use. Ablaut is a regular system of vowel variation, as seen in the
Modern English root sing. If we change the vowel of this verb we produce other grammatical
forms, though the semantic meaning remains the same: the present tense has one “grade” of the vowel,
which appears as sing, while the past tense has a different grade, and appears
as sang. The past participle originally had a zero-grade, that is, no vowel at all, and appears
in Modern English as sung. Finally, there is a noun version which originally had
yet another grade. This appears now as song. Throughout all the vowel grades, the basic
form of the word has remained as S_NG. We therefore can represent the basic structure of a PIE morpheme
as CVC (consonant-vowel-consonant), with the understanding that the internal vowel can change.

The PIE ablaut series alternates between an e-grade (the origin of sing) or an o-grade
(the origin of song) or a zero-grade (∅-grade) (the origin of sung) if there is
no vowel. There can also be grades with lengthened versions of the vowels: an ē-grade and
an ō-grade (the origin of song).

An example from PIE is the root morpheme *wed-, which means “wet.” This morpheme has
many descendents in Modern English, each of which looks very different today depending
on which vowel grade of the root was used or which derivational suffixes were added to it.
The root can appear in an e-grade as *wed-, a lengthened ē-grade as *wēd-, or it can appear
with an /o/ vowel, as *wod- or *wōd-; it can also appear with in zero-grade as *wd-. The problem
with this form is there is no vowel, so the semivowel /w/ converts to the vocalic form /u/, and
the root becomes *ud-. The following list presents Modern English words based on
different variations of the root *wed- from PIE:

  • o-grade with the noun suffix –r: *wod-r-. This is Modern English water,
    a noun formed from the adjective of being wet, meaning “the thing that is wet.”
  • o-grade with a reflexive suffix –sk- (meaning “to do to oneself”) and a verb suffix,
    turning the adjective into a verb. In Germanic this appears as *wat-sk-anan,
    “to make oneself wet.” This becomes Old English wascan, Modern English, “wash.”
  • zero-grade with a noun suffix relating to animals: *ud-ro- or *ud-ra-. This
    becomes udra in Greek (Modern English hydra) and in Germanic it
    becomes otor (Modern English otter). Both otters and hydras
    are etymologically “water animals.”
  • e-grade with a nasal infix -n- *we-n-d-, with the Proto-Germanic suffix *-ruz added to
    get Proto-Germanic *wintruz, “the wet time of year,” Modern English “winter.”
  • zero-grade with a nasal infix and a noun suffix: *u-n-d-a-, becomes Latin unda,
    “wave” (compare Modern English words like undulate, “to act like a wave.”)
  • zero grade with the suffix –skio: *udskio- becomes Old Irish uisce, Modern
    English whisky.
  • o-grade with the suffix noun –a: *wod-a-. Becomes Russian voda, “water” or with
    the diminutive suffix –ka, vodka, “little water, i.e. vodka.”

In sum, PIE morphology worked in very much the same way as Modern English, and all other descendents
of PIE. Root morphemes can appear with different grades of vowels and derivational suffixes can be
added to them to create new meanings. As such, these languages are quite different
from agglutinative languages like
Finnish, Japanese, or Nahuatl, which add strings of unchanged morphemes together.

PIE Inflectional Endings

PIE also had inflectional suffixes that were added to morphemes to give grammatical information.
Different functions of a noun within a sentence each take a different inflectional
ending. These functions are called cases. Modern English nouns have two basic cases.
Subject/object and possessive. Consider the two sentences:

  1. The boy carried a book.
  2. The boy’s books were heavy.

If we look at the first sentence, we see that the words boy and book
have no inflectional endings, even though one is the subject of the sentence and
the other is the object of the sentence. If we look at the second sentence, we see
that boy’s has the inflectional ending –‘s to signify that it is in
the possessive case, and books has the inflectional ending -s to signify a plural. These are the only case endings that remain in Modern English for nouns.

Pronouns in Modern English preserve three cases: subject, object and possessive.

  1. I saw the boy.
  2. The boy saw me.
  3. The boy saw my book.

In addition to the first person pronouns above, we have he, his and
him; we, our, and us; they, their, and them,
and so on. Some pronouns, however, only preserve two distinct forms: she and
her; you and your.

Modern English also preserves two numbers,
singular and plural. In most nouns, the plural
inflectional ending is s. Pronouns also preserve a distinction in number, with
I and we and he, she, it and they.

PIE had eight different cases, each with a distinct inflectional ending. It also
distinguished between singular and plural and also dual, indicating that there
were two of an object. Thus, a single noun could have up to 24 different inflectional
endings depending on which case it was and whether it was singular, dual or plural.
In fact, it had far more than these 24 endings, for different nouns took different kinds
of endings. We still have some remnants of this in Modern English. Our usual plural
suffix is –s, but we also have –en as in oxen and brethren.
We also have a zero-ending as in deer or sheep (i.e., people do not say *deers or *sheeps). If we add to
those the plural inflectional endings from foreign languages like Latin (cactus, cacti)
and Greek (phenomenon, phenomena), and so on, you will see that
English is not as simple in its inflectional endings as it at first appears.

Let’s look at the PIE root morpheme *eku-. This is probably related to the adjective *ōku “swift” (which has a
lengthened o-grade). If we add a basic noun suffix *-o onto the end of *eku-, we get the noun *ekw-o-, which would
mean “a thing that is swift.” This is the
origin of Latin equus, “horse,” from
which we get Modern English terms like equine,
equestrian
, etc. To put this noun
into a sentence, we must add inflectional endings to the root *ekwo-. If it is the subject of the sentence, we add
*–s, making *ekwos (Latin equus). If it is the direct object of the sentence,
we add *–m, making *ekwom (Latin equum). If it is possessive, we add *-syo, making
*ekwosyo. Other forms are *ekwobhyos,
*ekwod, *ekwoysu, and so on.

In Modern English verbs, we have one distinct ending in the present tense, for
the 3rd person singular, and we also have a past tense inflectional ending, –ed.
Many modern languages have different inflectional endings for 1st, 2nd, and
3rd person, both singular and plural. PIE had inflectional endings for all these
and in numerous different tenses and aspects (While tenses govern time, such
as future, present, and past, aspect governs the way in which the action was completed: a
habitual action, a completed action, an action that represents an eternal state, and so on).

In conclusion, PIE had a vast number of inflectional endings, not only far more
than Modern English, but more than one finds even in languages like ancient Greek
or Sanskrit. As PIE evolved throughout the millennia, it has lost many of these
endings, simplifying them into fewer and fewer cases or tenses.

Syntax

Because Modern English has lost case endings to
signify the grammatical function of nouns, we must rely heavily on word order (syntax) to make our sentences
clear. In the sentence “the
cowboy rode the horse,” we know “the cowboy” is the subject of the sentence
because it comes before the verb and we know “the horse” is the direct object
because it comes after the verb. If we
reversed the order, “The horse rode the cowboy,” we get a sentence that means
something entirely different. Languages
that rely heavily on word order are called analytic
languages. PIE on the other hand does not need to rely
on word order. It would not matter where
*ekwos appeared in a sentence in PIE because the final –s tells us that it is
the subject of the sentence. Similarly,
even if the word *ekwom appeared as the first word of an utterance in PIE,
speakers would recognize it as the direct object because of the final –m. Languages that rely heavily on inflectional
suffixes to signify grammatical structure are called synthetic. We will see in
the next few weeks how English has changed over the last 1000 years from a
highly synthetic language to a highly analytic language.

Because of the inflectional endings, a language
like PIE could be completely free in its word order. We still see this today in languages that are
heavily synthetic. Latin, for example,
has a very free word order. Linguists
have determined, however, that although PIE could
have used fairly free word order, it is most likely that PIE speakers preferred
to put the verb last. Thus, they would
have the subject of the sentence first, the direct object next, and the verb
last. This makes what we call an SOV (subject-object-verb) word order,
or more simply an OV word
order. English, on the other hand, is an
SVO language, or more simply, a VO language.
Further details of PIE syntax are beyond the scope of this course, but you can read
much more about it in this online edition of Winfred
Lehmann’s Proto-Indo-European Syntax.




Chapter 15.4 PIE Phonology

Consonants

Modern English has the following consonant phonemes:

One of the questions for the historical study of English is where and how these
phonemes originated. To start that process, examine the set of consonant phonemes in PIE:

The charts are quite similar in some aspects, though there are some major differences.
Both PIE and Modern English have the set of voiceless and voiced stops /p/ /t/ /k/
and /b/ /d/ /g/, but PIE divides its velar stops into three distinct types. Instead
of a simply /k/ for instance, there is a palatalized /k’/, a standard velar /k/,
and then a labiovelar /kw/. The first of these is similar to the type of /k/ that
we would pronounce in the words leak or king, where the velar
/k/ is what we would pronounce in the words lock or kong. The
labiovelar is a /k/ with rounded lips, similar to our “q” sound as in queen.

A further difference between PIE and Modern English consonant phonemes is that
PIE has a set of voiced aspirated stops: /bh/, /dh/, /gh/. These are similar to
/b/, /d/, and /g/, but with a heavy breath of air. In Modern English the voiceless
stops /p/ /t/ and /k/ are aspirated at the beginning of a stressed syllable,
but this aspiration is not phonemic; on the other hand, Modern English voiced stops /b/
/d/ and /g/ are not aspirated. In PIE, voiced stops were phonemically distinguished
into aspirated and non-aspirated varieties. To approximate the voiced varieties,
try saying the words abhor or adhere.

Finally, a number of Modern English phonemes do not occur. PIE only has one fricative,
/s/, and there are only two nasal phonemes, /m/ and /n/; /ŋ/ does not appear.

Laryngeals

The last set of sounds to discuss in the consonant set are the so-called laryngeals,
a set of sounds that sometimes acted as vowels, sometimes as consonants. They were
first proposed by Ferdinand de Saussure in a paper he wrote (when he was 21!) to
explain some idiosyncrasies in the Greek phonetic system. To put it extremely simply,
Saussure showed using the comparative method that there were some sounds in Greek
words that should have been different sounds if all the proper developments of
sound laws were followed. He hypothesized that the reason the “wrong” sounds appeared
in a few places in Greek could be explained by assuming that there must have been
other sounds in PIE that had entirely disappeared but had affected the sound changes
before they disappeared. It was difficult for him to prove the existence of something that
had completely disappeared, and therefore many people did not accept his theory when it
first appeared. Several decades later, however, some clay tablets were excavated
in Turkey written in a previously unknown language around the 16th century BCE.
The language, which scholars named Hittite, was eventually deciphered and recognized
an Indo-European language. In fact, it is the earliest Indo-European language that
was ever written down. As the language was documented, linguists realized that it
contained sounds which corresponded exactly to the laryngeals Saussure had proposed
and which occurred in the same places where he said they should! They are written today
as h1, h2 and h3. It is not entirely clear what
the distinctions between these were, and the details are not necessary to go into here.
Simply put, these sounds disappear in almost all PIE languages (except for the Anatolian
branch), but they do leave traces. The h1 laryngeal is a neutral one
and has no effect on the vowel around it. The h2 laryngeal has an “a-coloring”
to it. This is why we used it when we reconstructed the PIE word for father,
*ph2ter-. PIE only had an /e/ and an /o/ vowel, and so the /a/ vowel in the words
for father must have come from the h2 laryngeal. Finally, the
h3 laryngeal probably included lip rounding as well, since it often led to an /o/
type vowel.

The laryngeals disappear in Germanic and so we will not discuss them in any more detail.

Vowels

There are major differences in the vowels of PIE compared to Modern English. Modern English has
a set of five front vowels, some tense and some lax, one central vowel, and six back vowels,
again tense and lax varieties. According to some scholars, PIE in its earliest
stages only had one vowel /e/. Most scholars, however, find it easier to work with
the notion of two vowels: /e/ and /o/, each of which could occur in long and short varieties,
making four vowel phonemes total: /e/, /ē/, /o/, /ō/. Note that these are both mid vowels,
differing from each other in frontness and in roundness of the lips. Note also that in PIE vowel
length was a phonemic difference.

We can add to these four vowels the two semivowels /w/ and /j/. When these appear in an environment
in which there is no other vowel, they convert to the appropriate vowel /u/ and
/j/. Thus, the root *wed- (which we will discuss in the section on morphemes) has
a version in which there is no /e/ vowel. It would appear as *wd-, but this is
unpronounceable, and so the /w/ converts to its corresponding vowel, /u/, giving us
*ud-.

Although not classified as true vowels, the resonant sounds (i.e., /r/, /l/, /m/, /n/) could
appear as the nucleus of a syllable as well. This is similar to a Modern English word
like bottle, where the final /l/ acts as a syllable to itself, in the
same way a vowel would. In other words, to pronounce that word, we do not say
[bɑtəl] but rather [bɑtḷ]. The dot underneath the [ḷ] signifies that it is acting
as the nucleus of a syllable in the way that vowels normally do. We do the same
with nasals in English, such as in the words bottom or button.




Chapter 15.3 Linguistic Paleontology

The reconstruction of a PIE word for “father” tells us something that is quite
important about the speakers of PIE. Because they had a word for “father,” they
must have had fathers. Now of course this is obvious, since all humans have fathers,
but there are numerous other words to describe objects and concepts that are not
as obvious as fathers. By examining the entire lexicon of PIE, we can reconstruct
their society, how they lived, what kind of religion they had, what kind of farming
techniques they used, their foods, their economic systems, etc. This process
is called linguistic paleontology.

To give a brief example, many modern Indo-European languages have the same word
for “field”: Latin ager, Greek agros, Sanskrit ajras,
English acre. This suggests that the Indo-Europeans were farmers and herders
(as opposed to nomads or hunter-gatherers), since this word suggests not a wild
plain but a tilled field. We also find numerous cognates in modern Indo-European
languages for different livestock.

English Latin Greek Sanskrit
fee pecu paśu
ewe ouis ois avi-
eoh (OE) equus hippos aśva
cow bos bous gauḥ
swine sus sus

The words above are all cognates, showing that the words can be traced all the way
back to PIE. We can therefore assume that the Indo-Europeans had access to these
animals and probably herded them; otherwise they would not have had words for them.
Compare the following list:

English Latin Greek Sanskrit
donkey asinus onos khara
chicken pullus alektōr kukkuṭa-
camel camelus kamelos shṭra

These words are clearly not cognates. Take the words for donkey. The words donkey
and asinus are not related, and neither look very close to Sanskrit
khara. If we look at the etymologies of these words in the OED, we can see that
the origin of “donkey” is in fact a mystery. The OED suggests that the first part
of the word might be from the word dun, meaning “dark-colored,” or perhaps
the word is a nickname of Duncan. The Latin word asinus, the origin of
modern English ass (as in donkey, not buttocks), is probably a loanword
from a Semitic language. This makes sense since the native homeland of the donkey
was in the Middle East. So now we can hypothesize that the original Indo-Europeans
did not live in the Middle East, since they had no knowledge of donkeys.

The same is true for the various words for chicken. Chickens are originally from
Africa, and since the Indo-Europeans had no common word for chicken, we can
surmise that they did not originally live in Africa. The word chicken seems to have originated as
an affixed form of “cock,” which was probably an echoic word for the sound of a
cock crow (compare Modern English “cockle doodle do”). The Latin word pullus
is related to a general Indo-European word for the young of an animal, but not
specifically a chicken. The same word survives in English as foal,
a young horse. The words for camel in English, Latin and Greek do
look the same, but not because of shared origin. Instead, all of them were borrowed
from a Semitic word, as we would expect for an animal originally from Northern Africa
and Western Asia.

Based on the words we can reconstruct that were common to Proto-Indo-European,
we can tell that the Indo-Europeans were farmers, raising several types of grain,
as well as herders of livestock. They worshipped a paternal god whose name is
related to words for the sky or shining (Greek Zeus-pater; Latin Ju-piter; Sanskrit
Dyaus-pitar; Germanic Tiw, as in English Tuesday, or Tiw’s Day). They lived in
wooden houses (Latin dom-us; English tim-ber) with doors (Latin for-is, English door).
They travelled by means of some type of wheeled transport, probably wagons and
chariots (Greek cycle; English wheel). They were also able to travel by boat (Latin
navis) with oars, but apparently did not have windsailing technology, nor did
they have a word for ocean. Besides the grain that they farmed, one of their favorite
foods was honey, from which the alcoholic beverage mead was made.

Horses and the Kurgan Culture

One of the most important words for the reconstruction of Indo-European society is
the word for horse. The Indo-Europeans seem to have been one of the earliest cultures
to have domesticated the horse, and that would have given them a dramatic technological
advantage over other neighboring cultures. Horses would have extended the range of the
territory they could herd livestock on and increased the amount of land that
could be plowed. They would also have allowed them to enhance considerably the
distance they could travel. Furthermore, horses would have been a formidable addition
to their ability in warfare. In fact, early Indo-European cultures seem to have
ridden into war on chariots and could have easily dominated enemies fighting on foot.

The identification of the horse as a central animal for the Indo-Europeans has
led to the hypothesis that the Indo-Europeans can be identified with the
Sredny Stog culture (ca. 5000 BCE-3500 BCE) and the subsequent
Yamna culture (ca. 3500 BCE -2200 BCE). These peoples lived in
the area north of the Caspian Sea, in what is modern day Russia and Kazakhstan. The
earlier Sredny Stog culture is the earliest culture known to have domesticated
the wild horse. The Yamna culture buried its dead in large mounds known in Russian
as kurgans. Some of these graves even contain horses that were
interred with the dead—the teeth of the horses show wear from bits, a clear sign
of domestication. These mounds have led to the so-called Kurgan Hypothesis
of Indo-European origins, first developed by the archaeologist Marija Gimbutas in the
1950s. Most Indo-Europeanists today favor a modified version of this hypothesis
(Gimbutas’s original hypothesis was bound up with the notion that the original
Europeans had been a matriarchal society who worshipped a peaceful mother-goddess
and these had been taken over and subjugated by the warlike and patriarchal Kurgan
society. Scholars today tend to favor the notion of the Kurgan society as the original
Indo-Europeans but downplay the notions of patriarchy vs matriarchy).

The location of the Kurgan peoples fits well with what we know of the people who
spoke Proto-Indo-European. A number of river names in the area preserve old PIE words,
the flora (like beech and birch trees) and fauna (bears, salmon, wolves, etc.) of
that area at the time fit well with the animals whose names we can reconstruct for
PIE, and geographically it is plausible given the later spread into both Europe
and southeast Asia. For more information, including maps, see the wikipedia pages on
the Indo-Europeans,
and the Kurgan Hypothesis.

How the Indo-Europeans spread

How does a culture that lives in a fairly restricted area—the Russian steppes
north of the Caspian sea—spread out west to cover all of Europe and south into
India? It is important to realize that Europe was well populated in the third
millennium BCE. There were numerous cultures at that time to whom archaeologists
have given names, often based on distinctive pottery shapes (the Corded Ware Culture,
the Bell Beaker Culture, the Funnelbeaker Culture, etc.), each of which had its
own language. We can say almost nothing of these languages but it is assumed that they
were unrelated to PIE.

Throughout the third, second and first millennia BCE the Indo-Europeans started
spreading outwards from their original homeland, and presumably
encountered many of these other cultures on the way. We do not know exactly how
such contact proceeded. Gimbutas’s theories involved the notion of violent warfare
and domination. This stands in stark contrast to another theory of Indo-European origins,
the Anatolian hypothesis which suggested that the spread of the Indo-Europeans was a
slow gradual spread of farmers introducing their farming techniques to cultures
one by one. There is probably some truth to both of these ideas. Warfare must have
been involved here and there but we need not assume that it is the only method of
expansion. Advanced technology in farming or other aspects of culture could lead to
the spread of Indo-European speakers, or even just the language, with little to no
displacement of the actual population. Even if we only envision a small but elite
group moving into one area, it would be quite possible to see how the language
of that elite group could take over in the new area and displace the indigenous
languages–consider the global spread of English today.

There are in fact several language families in Europe today that pre-date the arrival
of the Indo-Europeans. Finnish and Hungarian are two related languages that both
derive from a common original called Finno-Ugric that was spoken
over a large part of Europe before the spread of Proto-Indo-European. Also,
Basque, spoken today in the mountainous area of northern Spain
and southern France, is a non-Indo-European language. But all other languages of Europe are
Indo-European in origin.

This spread of PIE was gradual, as people, or just the language itself, moved further
and further from the original homeland. The end result is that after a few millennia,
that is, by around 2500 BCE, it is hard to consider the Indo-Europeans a single culture
since they had moved so far apart. Speakers of PIE were now living throughout Europe,
down into the Greek and Italian peninsulas, up into Scandinavia, and farther east
into parts of Asia and south into India. With such a large geographical spread,
causing the isolation of people speaking one dialect of PIE from those speaking
other dialects of PIE, the differences in the dialects grew greater and greater
until they were no longer comprehensible to each other. This is the origin of the
so-called “daughter languages” or language families of PIE.




Chapter 15.2 Linguistic Reconstruction

The problem with studying a language like PIE is that it was spoken before the
development of writing. Thus, we must “reconstruct” the language. This is done
through what is called the comparative method, a process developed
by early linguists such as Rask, Bopp, and Jakob Grimm (one of the famous brothers
who gathered together Grimm’s Fairy Tales). The comparative
method, as its name suggests, is simply based on comparing different languages
to find their similarities and differences. Through this comparison, we can
reconstruct what their ancestor language must have looked like. Linguistic reconstruction
can be an incredibly complicated task, since it requires knowledge of the
phonology, morphology, and lexicon of all Indo-European languages, but the
principle behind it is simple enough: the comparison of the phonology and morphology of languages.

As an example, look at the words for “father” in English and some other Germanic
languages like German, Swedish, Norwegian, and Icelandic.

English father
German Vater
Swedish far
Dutch vader
Icelandic faðir
Norwegian far
Faroese faðir
Danish far
North Frisian faaðer

It should be clear immediately that all these words are very similar to each other.
These are the contemporary words for “father” in the Germanic family (Scaliger’s
Godt group). We call them Germanic because they are all descended from a common
language spoken by the Germanic tribal peoples who lived in northern Germany and
southern Scandinavia about 2500 to 2000 years ago. Although the Germanic tribes did begin
writing a few inscriptions in the Runic alphabet around the first or second century CE, much
of the language that they spoke was never written down. Thus, if we want to know what
the Germanic language was like before the Germanic tribes broke up, we must reconstruct
it from the surviving words. First, we must gather together all the words spoken in all these
languages. We could use contemporary versions, like those for “father” listed above,
but it’s much easier if we use the oldest versions of the words possible. In fact,
we have examples that were written down only a few centuries after the Germanic
tribes broke up.

Gothic faðar
Old Icelandic faðir
Old English fæder
Old Saxon fadar
Old High German fater

Note the similarities in these words. Each word begins with an [f] and ends with an
[r], making it very likely that the original Proto-Germanic word also began and
ended like this. Thus we can begin our hypothetical construction like this:

*f _ _ _ r

Next, the first vowel in all the words is either [a] or [æ]. Only Old English has [æ],
and if we compare other words in these languages we will notice that Old English always
has [æ] where the others have [a]. It is therefore most likely that
[a] was the original vowel and that the [æ] is the result of a later sound
change that affected English only. So we
can reconstruct the first vowel as [a], giving us:

*fa _ _ r

For the medial consonant two languages have [d], two
have [ð], and one has [t]. The details of reconstructing which of these sounds
was the original is complex, and so I will not go through all the details here,
but if we compare these sounds with what we find in other words in these languages
we can see that the original sound was [ð] and that the [d] and the [t] were the
products of later sound changes in English and German:

*fað_r

The next vowel is also difficult to explain without a much longer and more complex
discussion than is necessary here, but it seems to have been [ē] (that is, a long [e];
we could also write it [eː]). Thus, the reconstructed form of the word for
“father” in Proto-Germanic is

*faðēr

This is in fact almost exactly how we still pronounce the word today in Modern
English, showing how stable languages can be even over such a long period of time.

(Note: the symbol * is used before a word or phrase that is not attested,
meaning that there is no witness to it in written records. In contemporary English, this refers to word
forms or phrases that we consider unacceptable, but it is also used for
all reconstructed forms to signify that these forms are hypothetical
and not found in any written text, i.e., there are no witnesses to them. Since
Proto-Indo-European was not written down, every single PIE word should be
preceded by *).

This reconstruction allows us to formulate rules of certain sound changes. Consider the
[d] in the Old English word fæder and the [t] in the Old High German word
fater. From this one example we can extrapolate that whenever we have a
[d] in Old English, we should expect to find a [t] in German – of course this
will not be true every time because of other sound changes that interfere,
but it is often true. Consider the following words in Modern English and Modern High German:

English German
good gut
door Tür
day Tag
middle mittel

Observations like this are what allow linguists to figure out the sound changes
that have occurred so that they can reconstruct languages. They are also what allow us
to discuss how historical languages like English have changed over the years.

The last point to make about these Germanic words is that they are not different words. They
are cognates, related words that spring from a common source.
It is somewhat misleading to speak of the English word father as a different
word from the German Vater or from Icelandic faðir. They are
not different words, but the same word that now happens to be pronounced slightly
differently depending on whether one lives in Iceland or Germany or the US. In
fact, the pronunciation differences between these languages are probably no more
varied than the differences across the English speaking world.

Now that we have reconstructed the Proto-Germanic word for father, let’s
go even further back, to about 2500 B.C. when the Indo-Europeans were still
living as a linguistically unified group of peoples. We can start with other words for father that
you may know: Spanish padre, French père, and Italian padre.
All these words descend from the Latin word pater. Latin is a well-known and
well-attested language; it was spoken around the same time as Proto-Germanic, but the speakers of
Latin were literate, and so we have many attestations of the word pater
(notice we do not have to write it with an *).

But if you look closely at the Latin word pater and compare it with the English
word father or the Proto-Germanic word *faðēr, you
should see how similar the words are. They both begin with labial sounds [p] and [f],
they both have a low back vowel [a], they both have a medial dental consonant, [t] or
[ð], and they both end with [r]. Immediately we recognize that father and
pater are cognates; they both descend from a common original. That original is
Proto-Indo-European. So what was the PIE word for “father”? We can reconstruct it
in the same way. First, we gather up all the earliest attested forms in the various
Indo-European languages:

Latin pater
Greek patēr
Sanskrit pitar
Old Irish athair
Gothic faðar

As with the Germanic words, the similarities here should be noticeable, though
less obvious. Three of the languages, the three oldest by far, have an initial [p],
where the Germanic has an initial [f] and the Old Irish has no initial
consonant at all. All the languages end in [r] and have a medial dental sound.
Vowels are harder to reconstruct and those in this word are difficult, but linguists
who study PIE have reconstructed the original root morpheme for father in PIE as
*ph2tēr- (where h2 represents a sound called a laryngeal that became
an [a] like vowel in many later languages).

(Note: Because the reconstruction of all PIE words are root morphemes, we usually
add a hyphen (-) at the end of them, showing that any derivational or inflectional
suffixes may be added to them and that we should not take these reconstructed root
morphemes as actual “words”).

The examples above give a very simplified illustration of how
linguistic reconstruction is done. To do this thoroughly and correctly, all the
evidence must be compared, that is, all the words in all the earliest surviving
Indo-European languages. For further reading, you should look at the wikipedia page
on the Comparative Method.




Chapter 15.1 The Discovery of Proto-Indo-European

If you have learned a language such as Spanish, French, Italian, or German,
you may have noticed that there are many words that seem very similar to those
in English, such as Spanish “silencio” and English “silence” or German “Hand”
and English “hand.” This is not a new phenomenon; even the ancient Romans knew
that their language had many similar words with Greek. Se we must ask
why there are similarities. There are three possible explanations:

  1. coincidence
  2. borrowing
  3. shared origin

Very rarely coincidence is the reason for such a similarity.
For example, in the extinct Australian Aboriginal
language Mbabaram, the word for “dog” is “dog.” There is no connection between
English and Mbabaram that could otherwise explain the similarity between these words,
and so we must assume that it’s due to coincidence.

Borrowing is much more common than coincidence. English
words like “canyon,” “bronco,” “burrito,” and “rodeo,” were all borrowed
from Spanish. And “dance,” “government,” “joy,” and “villain,” were all
borrowed from French. English has borrowed a vast number of words from other languages, so
much so that the majority of words in English today are non-native, meaning they are
not part of the original word-stock of the Germanic languages. We call these words
loanwords. One of the main areas of studying the history of
a language is studying the historical factors that led to borrowing from other
languages.

The final explanation for similarity in words, and grammar as well, is
shared origin, a relatively recent idea, only a few
hundred years old. In the early modern period, up to the eighteenth century,
language was often connected to ethnicity. Thus, the origins of languages were
tied to the origins of different ethnic groups, all of which could ultimately be
traced back to the story from Genesis 11 of three sons of Noah (Shem, Ham, and Japheth) as the progenitors
of the entire human race. Shem, or Sem, was considered to be the ancestor of the Jews and Arabs
(Semites), Ham was the ancestor of the sub-Saharan Africans (Hamites), and
Japheth was the ancestor of the Europeans (Japhethites). When the offspring
of Noah’s sons were building the Tower of Babel, God sent confusion down to
them so that they all started speaking different languages, which then spread
throughout the world. Today Hebrew and Arabic are still called Semitic languages,
and in the seventeenth and eighteenth centuries European languages were called “Japhetic languages.”

Within these three large families, there was still some attempt to relate individual
languages to each other in smaller sub-groups. One of the greatest intellectuals of the
Renaissance, Joseph Scaliger (1540-1609), divided European languages
into various groups based on their words for “god.” He sorted them into the
Deus group (Latin deus, French dieu, Spanish dio,
Italian dio), the Godt group (English god, German Gott,
Dutch god, Swedish gud), the Theos group (Greek θεός) and finally the
Slavic boge group (Russian bog, Polish bog, Czech buh).
Scaliger, however, did not try to establish any connection between these four groups.

The first modern advancement in relating the languages was made
by Sir William Jones. “Sir William Jones was one of
the greatest polymaths in history. At the time of his early death, in 1794, he knew
13 languages thoroughly and another 28 moderately well. But languages were for him
only a means of reaching a deeper understanding, in contrasting cultures, of law,
history, literature, music, botany, and other disciplines. Elected at the age of
26 to [Samuel] Johnson’s Literary Club and knighted at 37, Jones was a close friend
to many leading English luminaries of his time. He was called “Oriental Jones” by
some, and his study of middle-eastern cultures, his championship of American independence,
and finally his appointment as high court judge in Calcutta, made him a truly universal
figure” (from the description
of the book by Alexander Murray, Sir William Jones 1746-1794: A Commemoration (Oxford
University Press, 1999)).

When Jones became a judge in Calcutta, he began to learn Sanskrit, the classical
language of India, and the language in which the laws of the country were written.
He noticed that there were very strong similarities between the grammar of Sanskrit
that of both Latin and Greek. It was no surprise to Europeans at the time that Latin
and Greek shared numerous features, something that had long been known, and which
had been attributed to cultural borrowing, but the inclusion of Sanskrit could
not be explained at all by cultural contact.

In a speech on Persian antiquities which Jones gave to the
Asiatic Society in 1786, he included this statement:

The Sanskrit language, whatever be its antiquity, is of a wonderful structure;
more perfect than the Greek, more copious than the Latin, and more exquisitely refined
than either, yet bearing to both of them a stronger affinity, both in the roots
of verbs and the forms of grammar, than could possibly have been produced by
accident; so strong indeed, that no philologer [i.e., linguist] could
examine them all three, without believing them to have sprung from some common
source, which, perhaps, no longer exists
; there is a similar reason,
though not quite so forcible, for supposing that both the Gothic and the Celtic,
though blended with a very different idiom, had the same origin with the Sanskrit;
and the old Persian might be added to the same family.

Jones did not hypothesize what language this “common source” may have been or seek
to discover if it still existed, but the hypothesis of its existence was groundbreaking
and prompted several scholars to begin the search. Two scholars in particular,
Rasmus Rask and Franz Bopp, began learning all the
relevant languages and figuring out the precise connections between them. In doing
so, we can say that they invented the modern field of linguistics.

Throughout the nineteenth century linguists worked to compile a list of languages that were related by shared origin,
all descended from an ancestor language that we now call Proto-Indo-European (PIE). Jones had suggested Sanskrit (and with it all
the current spoken languages of India), Greek, Latin, Germanic languages (like
Gothic and English), Celtic, and Persian. Others were soon added. Below is a list of the major families of Indo-European languages:

Indo-European languages

We use a biological metaphor to group these languages into families with PIE as the
mother language. Over time, PIE changed and evolved but not always in
the same way. People who lived in one part of the Indo-European homeland may have made
one sound change while people living in another part made a different sound change.
People in the northern part of the homeland may have started using one particular
grammatical construction while people in the southern part used a different one
and people over in the east started using an even different one. (Just think of all
the variations within American English, and then add 1000 years of continual change
to it). Eventually, so many changes added up that the different Indo-European dialects
were no longer intelligible to each other. Therefore, they became new languages;
we can think of them as daughter languages. These new languages also eventually underwent
their own internal changes, leading to even more new languages (we could
perhaps call them grand-daughter languages, although this is not a term that is
ever used). I will list now the major daughter language families and then the
“grand-daughter” languages that they have broken into. (Note: There are many Indo-European
language families I am omitting from this list; for a complete list and discussion,
see Philip Baldi, An Introduction to the Indo-European Languages).

  • Germanic. This language was spoken by the Indo-Europeans who
    settled in northern Germany and southern Scandinavia. It eventually further divided into a North
    Germanic branch, which includes Swedish, Norwegian, Icelandic, and Danish; an
    East Germanic branch which includes Gothic, an extinct language spoken by the
    Goths; and West Germanic, which includes English, German, Frisian, and Dutch.
    (Do not confuse the name Germanic with German). The earliest written Germanic languages are
    Runic inscriptions from the second century CE.
  • Celtic. This language group was spoken by peoples who lived
    throughout central Europe, north of the Italian and Greek peninsulas and south
    of the Germanic tribes. They eventually spread all the way from Galatia in modern
    day Turkey to Britain and Ireland in the west. The main surviving Celtic languages
    are Irish, Welsh and Scottish Gaelic. (Note, the words Celt and Celtic are pronounced
    with an initial “k” sound). Our earliest attested examples of Old Irish come from
    inscriptions of the fifth century CE written in the Ogam alphabet, but we do have earlier continental
    Celtic inscriptions from as far back as the sixth century BCE.
  • Italic. These languages were spoken throughout the Italian peninsula.
    Only Latin has survived, which then developed into the Romance languages (Italian,
    French, Spanish, Portuguese, Romanian). The oldest Latin inscriptions we have
    are from the seventh century BCE.
  • Hellenic. This group was spoken in Greece and survives only as Greek.
    The earliest Greek texts we have are written in the Linear B script on tablets from
    the fifteenth century BCE.
  • Balto-Slavic. This family includes the
    languages of the Baltic countries (Lithuanian and Latvian) and of the Slavic countries
    (Russian, Polish, Czech, Slovak, Serbian, Croatian, etc.).
  • Anatolian. This language family was spoken in what is modern-day
    Turkey and part of the Middle East. When the tablets containing these languages
    were discovered and deciphered in the early 1900s, scholars were unsure what
    language they were in or who created them and so they named them after the Hittites, a tribe mentioned in the
    Bible. None of the Anatolian languages are still spoken today, but they are
    the earliest written Indo-European languages and thus are incredibly valuable
    for our knowledge of what PIE was like. The earliest tablets were written
    in the seventeenth century BCE, although we have a few words from even earlier.
  • Indo-Iranian. This language family includes Persian
    (Farsi) as spoken in modern-day Iran, and Sanskrit and most of the languages spoken
    on the Indian sub-continent (except for Tamil and the other Dravidian languages of
    southern India).

Click here for a complete chart
of all
the Indo-European languages
(surviving and extinct).