Showing posts with label Linguistics. Show all posts
Showing posts with label Linguistics. Show all posts

December 24, 2009

Computational Linguistics

Ever have difficulty deciding whether material should be classed in 006.35 Natural language processing or in 410.285 Computational linguistics? (It would seem so, since many works have been classed in both numbers.) Since we have also found it difficult to distinguish clearly between the two numbers, we decided to take advantage of a recent major gathering of computational linguists at ACL-08: HLT (ACL = Association of Computational Linguistics; HLT = Human Language Technology) to get their feedback on the treatment of computational linguistics and natural language processing in the DDC.

Language variations around the world

Funny how one simple little question (“Where do comprehensive works on British slang belong?”) caused us such consternation. Our proposed response affects dozens of records. (OK, so the British slang—that is, English slang in the British Isles—solution involves only three records, but then that solution has to be propagated to languages throughout the 400s.)

December 14, 2009

Languages appear to take on lives of their own

Many of us at some time have probably looked at a dictionary and wondered, 'how many words are in there and how many do I know?' With new words being coined all the time and old words slipping out of use, it's impossible to know for sure how many words are in our language. No dictionary contains all the words. Estimates vary but they range, for English, from about half a million to about a million words. That million, though, includes about half a million highly technical/scientific words seldom found in standard dictionaries.

A good standard dictionary that serves us extremely well will likely contain less than 10 per cent of all English words and less than 20 per cent of all non-scientific words. My 'Concise Oxford' boasts 1,696 pages and '140,000 meanings.' I assume this indicates the actual number of words is perhaps only about 60,000 or fewer since many have two or more meanings. We can be sure the Oxford people chose that word 'meanings' carefully rather than using 'words defined.' Making such distinctions is their job!

How many?

How many words a person knows, of course, varies by individual, but a good estimate can be made. Studies show on average it is about 50,000 words actively (we use them and don't have to think too hard to do so), and about another 12,500 passively (we don't use them, but know them when we hear them): a total of about 62,500 words. Is it coincidence this is roughly the same number a good store-bought dictionary contains?

Adult English-speakers, studies show, on average speak about 15,000 words a day and about 370 million in a lifetime (this includes many repeats, each use being 'a word'). And we all know someone who's trying to increase the average, don't we!

Where from?

Where do words come from? Beyond the obvious (many somebodies invented them), the answer of how it works is being vigorously researched. Recent columns here have outlined several scientific efforts under way to get to the real roots of language. Many linguists now believe that once speech arises in some incipient form, language itself evolves via a process that seems to occur 'naturally.'

Professor Simon Kirby, chair of language evolution at the University of Edinburgh, Scotland, is doing intriguing research with fascinating results. It doesn't 'prove' anything yet, but it gives us good glimpses into how it probably works, subject to more research.

Indications

Professor Kirby designed an experiment he can do in an afternoon. He invented an 'alien language' consisting of a series of made-up words for 'alien fruit.' Nothing more. It isn't really a language, just a random batch of invented names for a batch of fictional fruit. A subject is shown fanciful drawings of the non-existent fruit and must learn the names of each. Then they are tested. Because it is at first random words with no connections between them, the subjects do poorly.

But Prof. Kirby is sneaky. During the test he introduces a picture of a new alien fruit the subject hasn't seen before. They have no word for it. Most subjects don't even notice! They simply make up a word for it that seems consistent with the 'alien' words they did learn (trying to recall the word they think they must surely have been given).

Next, the word or words the first subject unknowingly invents are included in the words taught to the second subject, and so on. Each subject believes they are simply giving back the words they learned, but in fact with each person the incipient language is changing over time, little by little. And over time the 'language' not only grows, but converts itself from a random, chaotic language to one with an identifiable structure with combinations that can be more easily learned! After a mere nine 'generations' many of the words can be broken into parts, each with a different meaning; parts that can be recombined for new meanings.

"This happens gradually, it happens without anyone knowing and yet it is the exact feature of language that makes it unique, makes language different from everything else in nature," says Prof. Kirby.

Conclusions

This suggests strongly that nobody designed English or other languages; nobody began with any established rules, rather these evolved. Second, the research gives a glimpse of how language changes over time. It suggests that language itself is a living entity with a life of its own. Prof. Kirby says this is one thing that makes it so fascinating, but also so difficult to pin down. It's a complex process.

The last word

Here is Professor Simon Kirby on what he takes from his own research:

"No one designed English, no one sat down and said it would be useful if it had relative clauses or grammar, it just happens, happens by this blind, unconscious process of transmission."

Babies develop crying 'accents'

LEIPZIG, Germany, Nov. 9 (UPI) -- French infants cry in a different way than German babies in their first days of life, researchers in Germany and France found.

Researchers from the Max Planck Institute for Human Cognitive and Brain Sciences in Leipzig, the Centre for Pre-language Development and Developmental Disorders at the University Clinic Wurzburg, both in Germany, and the Laboratory of Cognitive Sciences and Linguistics at the Ecole Normale Superieure in Paris compared recordings of 30 French and 30 German infants 2-5 days old.

The researchers found the French newborns more frequently produced rising crying tones, German babies cried with falling intonation.

In the last trimester of the pregnancy, human fetuses become active listeners, the researchers said.

"The sense of hearing is the first sensory system that develops," Angela Friederici, one of the directors at the Max Planck Institute said in a statement.

"The mother's voice, in particular, is sensed early on, what gets through are primarily the melodies and intonation of the respective language."

The crying patterns of the German infants mostly began loud and high and followed a falling curve while the French infants more often cried with a rising tone -- similar to the intonation of the babies' mother tongue. This early sensitivity to features of intonation may later help the infants learn their mother tongue, the researchers said.

The findings are published in Current Biology.

Family GUy Linguistics

The most recent Family Guy episode ("Family Gay": Season 7, Episode 8 ) relied on some interesting linguistics in two separate jokes.

First, when Peter enters his brain damaged horse, 'Till Death, in a race, it runs amok in the crowd and the announcer on the loud speaker says the following (about the 5:50 mark on Hulu):

What's this? It looks like 'Till Death has taken a right turn and is heading into the stands. Dear god! I could describe the horror I am witnessing but it is so unfathomably ugly and heart rendering that I cannot bring myself to do so, although I do posess the necessary descriptive powers. Haw, well at least the horse raced past the class of visiting deaf second graders...oh no! Dear god he's going back oh I know you can't hear any screams but I assure you they are signing frantically just as fast as their little fingers can shape the complicated phonemes necessary to convey dread and terror.

Having never studied the linguistics of sign language, my first reaction was to ask, are there truly phonemes in sign langauge? In spoken language, a phoneme is a conceptual clustering of phonetic segments into a single group. For example, in English the segment /p/ can occur with a little extra burst of air called aspiration (typically at the beginning of words), or not, like at the end of words (try saying the words "pat" and "tap" with your hand in front of your lips and, if you're a native speaker if English, you should be able to feel the little burst of air that accompanies the /p/ in "pat" but not in "tap"). So, there is a difference in how we articulate /p/ depending on where it occurs in a word. Nonetheless, we still consider both versions of /p/ to be "the same sound." We say there is a single phoneme [p] with two phonetic realizations, aspirated and unaspirated.

But this use of "phoneme" is based on spoken language. How does this relate to signed languages like ASL? After a quick bit of Googling, I've discovered that the term "phoneme" is in fact used to refer to segments of signed language by various sign language scholars, though it is used more as a conceptual borrowing than as a term referring to sound. The most relevant discussion I found was in an abstract for the paper "Sign language phoneme transcription with PCA-based representation" by Kong, W.W. and Ranganath, S. (from the National University of Singapore). They "first apply a semi-automatic segmentation algorithm which detects minimal velocity and maximal change of directional angle to segment the hand motion trajectory of signed sentences. We then extract feature descriptors based on principal component analysis (PCA) to represent the segments efficiently. These high level features are used with k-means to cluster the segments to form phonemes." According to this approach, phonetic segments are roughly approximated to sign language as feature sets composed of "minimal velocity and maximal change of directional angle" and phonemes are approximated as k-means clusters of those feature sets. Cool stuff, for sure. But it's not clear to me if this computational approach is consistent with the natural way humans actually perceive and analyze sign language segments. I'm still looking for more on that topic.

Nonetheless, there remains the issue of the Family Guy writers getting the nature of phonemes fundamentally wrong. Phonemes convey no meaning (excusing for the moment the weak possibility of sound symbolic associations). The writers, who were clearly willing to do a little research (even if a very little), could easily have substituted "morphemes" for "phonemes" in the script and would have had the same joke without the error. And just how complicated are the signs for dread and terror anyway?

(pssst, I've clearly spent too much time reading linguistics because when I first read "the horse raced past the class of visiting deaf second graders" I assumed it was a garden path sentence similar to Bever's famed example "The horse raced past the barn fell." It took me several reads to realize that, nope, "raced" is not a reduced relative clause, but rather a run-of-the-mill past tense main verb. A nice example of construction priming, eh? I'm primed to read any "X raced past Y" clause as being a reduced relative).


Second, the writers went out of their way to construct a joke not only based on conversational pragmatics, but based on EXPLAINING Gricean maxims (about the 9:50 mark on Hulu.com).

Lois: Peter, what exactly did they inject you with?

Peter: Oh all sorts of things. Hepatitis vaccine, a couple of steroids, the gay gene, calcium, a vitamin B extract...

Lois: What did you just say?

Peter: The gay gene. I assume that's the one you meant even though it wasn't literally the last thing I said when you said what did you just say, it's just that clearly (it) was most unusual... (note: the pronoun "it" was reduced to near imperceptibility).

In this exchange, Peter explains that, under normal circumstances, after listing a set of items and someone asks "what did you just say" he would interpret "what" as referring to the most recent item in the list (presumably because of the semantics of "just"). But in this case, one earlier item was more "unusual" than the others.

Let's re-explain this using conversational pragmatics and Gricean maxims, okay?

Peter lists five items. He believes that one of the five items is controversial while the other four are not. He believes Lois believes this too. The controversial item is in the middle of the list. Peter believes he articulated each item clearly such that Lois could properly hear all items. He believes Lois believes this too. So, when Lois asks "what did you just say," Peter believes 1) that she heard the most recent item clearly and 2) that this item has little informational value. He believes Lois believes this too. Peter believes Lois is not flouting conversational norms. He believes Lois believes this too. Therefore, Peter believes Lois is trying to make her contribution (her question) informative (maxim of quantity). He believes Lois believes this too. Peter believes that repeating a well heard, uncontroversial item has no information value. He believes Lois believes this too. Thus, he infers that "what" must refer to some item other than the last one. He believes Lois believes this too. Peter believes there is only one item on the list that meets the information value requirement. He believes Lois believes this too.

This is a long-winded way of saying the same thing Peter did, but we linguists have to make things complicated and technical. It's our job.

Is There A GEnder Gap in Linguistics?

The interwebs is abuzz with the latest girls can't do science scuttlebutt (Alex Tabarrok at Marginal Revolution has a useful overview). This got me to wondering why this kind of speculation never seems to be applied to linguistics.

First, of course, is the fact that linguistics is a small field (I'm near certain that Language Log had a post in the last few months regarding the small size of linguistics compared to other fields, but I've failed miserably to find it). There simply aren't enough of us to cause any controversies, except when someone wants us to prove that "X" is not really a word and we refuse (see Zwicky's relevant post here).

Second, there are many easily recognizable female linguists who have been highly influential. Off the top of my head I can easily think of Barbara Partee, Adele Goldberg, Joan Bresnan, Joan Bybee, Eve Sweetser, and Eve Clark , amongst a great many others (hehe, that list TOTALLY marks me as a Buffalo functionalist).

At least as interestingly, the highly technical sub-field of computational linguistics/NLP is brimming with examples of influential female scholars, such as Ann Copestake, Paola Merlo, Bonnie Dorr, Tanya Reinhart, and Jan Weibe to name a just a small few.

The only sub-field of linguistics that seems to be male dominated is syntactic theory invention. I can think of many males who are strongly associated with the invention of a theory of syntax/grammar (though, in all honesty, no one person truly invents a theory alone), but only Goldberg & Michaelis's construction grammar comes to mind as a theory of grammar designed by female linguists (I happily solicit examples of my ignorance). Whereas, a list of male-invented grammatical frameworks/theories is easy to compile:

* N. Chomsky --> Minimalism, GB, Transformational Grammer

* G. Gazdar --> Generalized Phrase Structure Grammar

* Pollard & Sag--> HPSG

* Van Valin & LaPolla --> Role and Reference Grammar

* D. Perlmutter --> Relational Grammar

* C. Fillmore --> Case Grammar

* R. Langacker --> Cognitive Grammar

Is there a gender gap in grammar theorization?

The Preposition "from"

Having just now discovered I missed National Preposition Day, I offer a post relating to the preposition from and my dissertation.

It has long been noted that the English preposition from most typically occurs with Source (Huddleston and Pullum 2002; Van Valin and LaPolla 1997; Jolly 1991; Clark and Carpenter 1989a; Clark and Carpenter 1989b; Quirk 1985; Vestergaard 1977; Wood 1967).

(1) a. Chris returned from California.
b. Hide took the book from Atsuko.
c. Mike drove from Buffalo to Toronto.

There are some uses, however, where it appears to occur with Goals and Themes as well.

(2) a. The fence blocked the car from the driveway.
b. The tent shielded the kids from the rain.

When from occurs with verbs denoting barrier events like bar, ban, block, shield it can marks NPs representing either the unattained goal of the entity being blocked, or the restrained theme which failed to attain its goal. Interestingly, it can also mark VPs as in (3):

(3) The judge barred the journalists from entering the courtroom.

In my first qualifying paper (SUNY Buffalo’s linguistics department uses qualifying papers in lieu of a master’s thesis) I argued that from acts not as a preposition, but rather as a complementizer with barrier verbs.

The class of English “barrier verbs” (as originally sketched by Len Talmy) are negative causative object control verbs which encode the relationships between a goal directed participant (or “agonist”), its goal and a barrier participant (or “antagonist”). They are negative verbs in the sense of Laka 1994: they encode the negation of an event. Think of the verb neglect. If you neglect to do X, then X did not happen. With barrier verbs, if you ban X from doing Y, then Y did not happen.

These verbs fall into the following general constructional template:

NP1 verb NP2 from NP3/VP.


In this construction, the subject of a barrier verb (NP1) acts as the barrier (either directly or indirectly) to the achievement of a goal event (NP3/VP) by a goal-directed participant NP2).

I have recently discovered that Idan Landau has a detailed analysis of negative verbs in Hebrew which, extended to English prevent, suggests the complementizer interpretation as well. Though we have very different theoretical frameworks, I think we share some conclusions.

I’m interested in a variety of the phenomenon associated with barrier verbs (including the potential for a coercion analysis of the NP replacement of complement VPs)

Google Linguistics

I have used Google repeatedly to find instances of constructions that I could not find using standard corpus linguistics methods with hand compiled corpora like the BNC. Typically I’m looking for any instance, just to prove people really do say the thing I’m claiming is possible. For example, I needed to find some examples of passivized complements embedded under 60 different barrier verbs following this pattern:

a. I banned John from being examined by the doctor.
b. I banned John from getting examined by the doctor.

Many of the verbs I wanted to search for are low frequency in the BNC (e.g., barricade, derail, hamper, etc) so the likelihood of finding examples of passivized complements using say a Tgrep2 search is low. So, I ventured into the scary land of Google Linguistics. I used the search query “verbed * from being” and “verbed * from getting” Within a short time, I had multiple examples for most of the verbs I was looking for. I can’t imagine performing this task more efficiently with any other tool. Google really worked well under those circumstances.

Let me note that I have not used Google hit counts or page counts to derive any statistics regarding frequency of occurrence, though. When I do this sort of thing, I’m careful to use my common sense to decide if a return is from a native speaker or not, and often what I do is skim a page to see if there are any obvious ESL errors. Also, I use my own intuition regarding the acceptability of a usage (by pure coincidence, Peter Ludlow from U. Toronto will be here in Buffalo this week giving a talk on the role of linguistic intuitions).

One of the more thorough discussions of the use of search engines in linguistics research is Adam Kilgarriff’s “Googleology is bad science”, a squib from Computational Linguistics (2007, v33, 1)

He writes that the web is attractive to linguists because it is “enormous, free, immediately available, and largely linguistic”. But, he points out four major flaws:

1. search engines do not lemmatise or part-of-speech tag
2. search syntax is limited
3. there are constraints on numbers of queries and numbers of hits per query
4. search hits are for pages, not for instances.

Kilgarriff offers this alternative: “work like the search engines, downloading and indexing substantial proportions of the web, but to do so transparently, giving reliable figures, and supporting language researchers’ queries”

The squib goes on to detail how we might go about doing that in a principled way. It’s well worth the read.

TEORI DAN KONSEP PEMEROLEHAN BAHASA (Teori Behaviorisme, Teori Nativisme, dan Teori Kognitivisme)

Teori Behaviorisme

1. Teori Behaviorisme mulanya adalah teori belajar dalam psikologi yang telah muncul sejak 1940-an s/d awal 1950-an dan John B. Watson dianggap sebagai pelopor utama dalam teori ini.

2. Otak bayi waktu dilahirkan sama sekali seperti kertas kosong/piring kosong(tabularasa/blank slate), yang nanti akan diisi dengan pengalaman-pengalaman.

3. Bagi mereka istilah bahasa menyiratkan suatu wujud, sesuatu yang dimiliki dan digunakan, dan bukan sesuatu yang dilakukan. Itulah sebabnya mereka menyebutnya dengan Verbal Behavior (perilaku verbal) yang kemudian konsep-konsep tersebut tertuang dalam bukunya B.F. Skinner yang berjudul Verbal Behavior (1957)

4. pengetahuan dalam bahasa manusia yang tampak dalam perilaku berbahasa adalah merupakan hasil dari integrasi peristiwa-peristiwa linguistik yang diamati dan dialami manusia

5. Kemampuan berbicara dan memahami bahasa oleh anak diperoleh melalui rangsangan dari lingkungannya dan anak dianggap sebagai penerima pasif dari tekanan lingkungannya, tidak memiliki peranan yang aktif didalam proses perkembangan perilaku verbalnya.

6. Mereka juga tidak mengakui penguasaan anak terhadap kaidah bahasa dan kemampuannya untuk mengabsrakkan ciri-ciri penting dari bahasa di lingkungannya. Namun adapun ketika anak berbicara itu disebabkan oleh keberhasilan lingkungan yang membentuk anak itu

7. Mereka juga tidak mengakuai kematangan si anak dalam perkembangan pemerolehan bahasa, tetapi proses perkembangan sama sekali ditentukan oleh lamanya latihan yang diberikan oleh lingkungannya. Adapun perkembangan bahasa dipandang sebagai kemajuan dari penerapan prinsip stimulus-respon dan proses imitasi (peniruan)

8. Kekurangannya: teori ini tidak mampu menjelaskan proses pemerolehan bahasa itu sendiri dan faktor kreatifitas dalam penggunaan bahasa serta bagaimana kompetensi bahasa digunakan untuk membuat dan memahami kalimat-kalimat baru yang belum pernah didengarnya.

9. Dalam kaitannya dengan belajar B2, Lado (1964), mengatakan bahwa seseorang yang memulai belajar B2 cendrung akan menggunakan kebiasaan-kebiasaan (kaidah) yang dibentuk pada B1-nya, sehingga kebiasaan itulah yang terbawa ketika belajar B2.

10. Itulah sebabnya teori Behaviorisme sering dikaitkan dengan hipotesis analisis kontrastif (suatu metode sinkronis dalam analisis bahasa untuk melihat/mencari persamaan dan perbedaan antara kedua bahasa atau lebih). Jadi, jika ada kemiripan B1 dan BT/B2, maka anak akan memperoleh struktur BT/B2 dengan mudah, tetapi jika sebaliknya maka anak akan menemui kesulitan.

11. Jadi bagi kaum behaviorism bahwa belajar bahasa dan perkembangannya hanyalah persoalan bagaimana mengkondisikan anak dengan cara “imitation, practice, reinforcement, and habituation”, yang merupakan langkah pemerolehan bahasa.

12. Dalam pengajaran bahasa, behaviorisme mengembangkan metode drill atau memperbanyak latihan baik dalam bentuk lisan atau tulisan.

Teori Nativisme

1. Rounded Rectangle: Diskusi teori Pemerolehan Bahasa oleh Asbah dan Roni Amrullah Kamis, 12 November 2008 Teori ini dipelopori oleh Noam Chomsky pada awal tahun 1960-an sebagai bantahan terhadap teori belajar bahasa yang dilontarkan oleh kaum behaviorisme tersebut, yang kemudian menulis buku berjudul “(Review of B. F. Skinner’s Verbal Behavior, 1959) sebagai bantahan terhadap konsep skinner tentang belajar bahasa yang ada dalam buku Verbal Behavior (1957).

2. Nativisme berpendapat bahwa selama proses pemerolehan bahasa pertama, anak sedikit demi sedikit membuka kemampuan lingualnya yang secara genetis telah diprogramkan. Jadi lingkungan sama sekali lingkungan tidak punya pengaruh dalam proses pemerolehan (acquisition).

3. Chomsky mengatakan bahwa Bahasa terlalu kompleks untuk dipelajari dalam waktu dekat melalui metode imitation seperti anggapan kaum behaviorisme. Dan juga bahasa pertama itu penuh dengan kesalahan dan penyimpangan kaidah ketika pengucapan atau pelaksanaan bahasa (performance). Manusia tidak mungkin belajar bahasa pertama dari orang lain seperti klaim Skinner

4. Menurut Chomsky bahasa hanya dapat dikuasai oleh manusia, karena: 1) perilaku bahasa adalah sesuatu yang diturunkan (genetik), pola perkembangan bahasa berlaku universal, dan lingkungan hanya memiliki peran kecil dalam proses pematangan bahasa. 2) bahasa dapat dikuasai dalam waktu singkat , tidak bergantung pada lamanya latihan seperti pendapat kaum behaviorism. Lihar proses perkembangan bahasa anak.

5. Chomsky menganggap Skinner keliru dalam memahami kodrat bahasa. Bahasa bukan suatu kebiasaan tetapi suatu sistem yang diatur oleh seperangkat peraturan (rule-governed). Bahasa juga bersifat kreatif dan memiliki ketergantungan struktur.

6. Jadi, pemerolehan bahasa bukan didasarkan pada nurture (pemerolehan itu ditentukan oleh alam lingkungan) tetapi pada nature. Artinya anak memperoleh bahasa seperti dia memperoleh kemampuan untuk berdiri dan berjalan. Anak tidak dilahirkan sebagai tabularasa, tetapi telah dibekali dengan Innate Properties (bekal kodrati) yaitu Faculties of the Mind (kapling minda) yang salah satu bagiannya khusus untk memperoleh bahasa, yaitu “Language Acquisition Device.

7. LAD ini dianggap sebagai bagian fisiologis dari otak yang khusus untuk mengolah masukan (input) dan menentukan apa yang dikuasai lebih dahulu seperti bunyi, kata, frasa, kalimat, dan seterusnya. Meskipun kita tidak tahu persis tepatnya dimana LAD itu berada karena sifatnya yang abstrak (invisible).

8. Dalam bahasa juga terdapat konsep universal sehingga secara mental telah mengetahui kodrat-kodrat yang universal ini. Chomsky mengibaratkan anak sebagai entitas yang seluruh tubuhnya telah dipasang tombol serta kabel listrik: mana yang dipencet itulah yang akan menyebanbkan bola lampu tertentu menyala. Jadi, bahasa mana dan wujudnya seperti apa ditentukan oleh input dari sekitarnya,

9. Antara Nurture dan Nature sama-sama saling mendukung. Nature diperlukan karena tampa bekal kodrati makhluk tidak mungkin anak dapat berbahasa dan nurture diperlukan karena tanpa input dari alam sekitar bekal yang kodrati itu tidak akan terwujud (Dardjowidjojo,2003). (Contoh kasus lihat Soenjono,2003; hal, 236-237)

Teori Kognitivisme


1. Munculnya teori ini dipelopori oleh Jean Piaget (1954) yang mengatakan bahwa bahasa itu salah satu di antara beberapa kemampuan yang berasal dari kematangan kognitif. Jadi perkembangan bahasa itu ditentukan oleh urutan-urutan perkembangan kognitif.

2. Menurut Piaget struktur yang kompleks itu bukan pemberian alam dan bukan sesuatu yang dipelajari dari lingkungan melainkan struktur itu timbul secara tak terelakkan sebagai akibat dari interaksi yang terus menerus antara tingkat fungsi kognisi anak dengan lingkungan kebahasaannya

3. Menurut kaum kognitivisme bahwa kemampaun pembelajar sudah terprogram secara biologis untuk memiliki kemampuan kognitif dan proses belajar terjadi dengan cara memetakan kategori linguistik kedalam kategori kognitif, serta apa yang dipelajari adalah tatabahasa sebuah bahasa.

4. Jadi, sebetulnya kaum kognitivisme berusaha menggabungkan peran lingkungan dan faktor bawaan, namun lebih besar ditekankan pada aspek berpikir logis (the power of logical thinking)

5. Rounded Rectangle: Diskusi Teori Pemerolehan Bahasa oleh Asbah dan Roni Amrullah Kamis, 12 November 2008 Urutan pemerolehan bahasa: menuranikan struktur aksi – representasi kecerdasan – membentuk struktur linguistik. (Lebih jelas lihat Chaer, 2003; hal, 178-179).

On Linguistic Fingerprint

Can an author's writing style be defined by the frequency of unique words in their writings? According to physicist Sebastian Bernhardsson, the answer is yes. He found a couple of interesting facts: 1) the more we write, the more we repeat words and 2) the rate of repetition (or rate of change) seems to be unique to individual authors (creating a "linguistic fingerprint"... literally his words, not mine). Let me walk through his claims and findings, just a bit.

Bernhardsson et al. are in press with a corpus linguistics study which compared rates of unique words between short and long form writing (short stories vs. novels vs. corpora). I stumbled on to this research earlier this week when a BBC News title caught my eye: Rare words 'author's fingerprint': Analyses of classic authors' works provide a way to "linguistically fingerprint" them, researchers say.

The idea of linguistically fingerprinting authors has been around for a while. In some ways it acted as a lost leader decades ago, piquing interest in the use of corpora and statistical methods to study language and now there is even a whole journal called Literary and Linguistic Computing. Plus, there is an established practice of forensic linguistics where linguistic methods are used to establish authorship of critical legal documents.

However, Bernhardsson makes a bold claim. He claims that the process of writing (a cognitively complex process) can be described as the process of pulling chunks out of a large meta-book which shows the same statistical regularities of an authors real work (he hedges on this a bit, of course). I always shiver when I run across a non-linguist jumping head first into linguistics making bold claims like this, but I also recognize that Bernhardsson and and his co-authors are pretty smart folks so I gave them the benefit of the doubt and skimmed one of their two available papers (freely available here).

* The meta book and size-dependent properties of written language. Authors: Sebastian Bernhardsson, Luis Enrique Correa da Rocha, Petter Minnhagen. New Journal of Physics (2009), accepted.

First, I concentrated on the first section because the paper goes into a different direction that was not necessary for me to cover (and had lots of scary algorithms; it is Sunday and I do want to watch football, hehe). What they did was count the number of words in a text, then count the number of unique words (this is a classic type/token distinction). Here's what they found:

When the length of a text is increased, the number of different words is also increased. However, the average usage of a specific word is not constant, but increases as well. That is, we tend to repeat the words more when writing a longer text. One might argue that this is because we have a limited vocabulary and when writing more words the probability to repeat an old word increases. But, at the same time, a contradictory argument could be that the scenery and plot, described for example in a novel, are often broader in a longer text, leading to a wider use of ones vocabulary. There is probably some truth in both statements but the empirical data seem to suggest that the dependence of N (types) on M (tokens) reflects a more general property of an authors language. (my emphasis and additions).

First, let's make sure we get what the author's did. We have to use words more than once, right? I've already repeated the word "we" in just the last two sentences. And we repeat words like "the" and "of" all the time. We have to. So there are types of words, like "the" but there are also the number of times those words get repeated (tokens). It's pretty straight forward to simply count the total number of words in a story, then count the total number of types of words. Thus giving us a ratio. For example, let's say we have a short story by Author X with 1000 words it (= tokens). Then we count how many times each word is repeated and we find that there are only 250 unique words (= types), this means there is a ratio of 1000/250, or 100/25 (for comparison's sake I'm using this ratio). This means that only 25% of the words are unique, which also means that, on average, a word is repeated 4 times in this story.

Now let's take a novel by Author X with 100,000 words (= tokens). After counting repetitions we find it has 11000 unique words. Our token/type ration = 100,000/11000, or 100/11. This means that only 11% of the words are unique, which means, on average, a word gets repeated about 9 times. That's higher than in the short story. Words are being repeated more in the novel. Now let's imagine we take all of Author X's written work, put it together into a single corpus and repeat the process and discover that the ratio is 100/7 (on average, a word gets repeated about 14 times).

UPDATE: whoa, my maths was off a bit the first time I did this. That'll teach me to write a blog post while watching Indie crush Denver. Sorry, eh,

This is what the author's found: "The curve shows a decreasing rate of adding new words which means that N grows slower than linear (α less than 1)."

They discovered something potentially even more interesting. there is a rate of change between these ratios is unique to each author: Here's is their graph from the article (H = Thomas Hardy, M = Herman Melville, and L = D.H. Lawrence):
FIG. 1: The number of different words, N, as a function of the total number of words, M, for the authors Hardy, Melville and Lawrence. The data represents a collection of books by each author. The inset shows the exponent = lnN/ lnM as a function of M for each author.

Their conclusions about the meta-book and linguistic fingerprint:

These findings lead us towards the meta book concept : The writing of a text can be described by a process where the author pulls a piece of text out of a large mother book (the meta book) and puts it down on paper. This meta book is an imaginary infinite book which gives a representation of the word frequency characteristics of everything that a certain author could ever think of writing. This has nothing to do with semantics and the actual meaning of what is written, but rather to the extent of the vocabulary, the level and type of education and the personal preferences of an author. The fact that people have such different backgrounds, together with the seemingly different behavior of the function N(M) for the different authors, opens up for the speculation that every person has its own and unique meta book, in which case it can be seen as a fingerprint of an author. (my emphasis)

They are quick to point out that this finding says nothing about the semantic content of the writings. So what does it say? I admit I was having a hard time seeing any conclusion about cognition or the writing process, even while finding this methodology interesting, I'm just not at all sure what it really says about the human brain and language, if anything at all. The speculation that "every person has their own unique meta book" is bold. Unfortunately, it is also almost entirely untestable. Keep in mind that this research had zero psycholinguistic component. They were just counting words on pages. I'd caution against drawing any conclusion about the human language system based solely on this work. (I should note that I skipped one of the most interesting findings, that the section of work doesn't matter, simply the size. meaning, they took random chunks from their corpora and found the same patterns, if I understood that part correctly.) Which begs the question: why is this being published in a physics journal? It's being published in The New Journal of Physics and a quick perusal of the articles from previous editions doesn't show anything remotely similar to this work (no surprise).

I'm a fan of corpus linguistics, but I'm also a fan of caution. I'm not convinced any conclusions about the psycholinguistics of the complex writing process can be drawn from this work. Not as yet. But interesting, nonetheless.

FYI: it's easy enough to fact check some of these results using freely available tools, namely KWIC Concordance. This tool will take any text and count the total tokens and number of repeats for us. I did this for Melville's Bartleby, the Scrivener and Moby Dick. I got text versions of each from Project Gutenberg, then ran the wordlist function within KWIC and here are my results:

Bartleby
Total Tokens: 18111
Total Types: 3462
Type-Token Ratio: 0.191155

Moby Dick
Total Tokens: 221912
Total Types: 17354
Type-Token Ratio: 0.078202

Bartleby = 0.191155
Moby Dick = 0.078202

Yep, the short story Bartleby has more unique words than the longer Moby Dick. FYI, this is a weak test simply because the tokens are not stemmed, meaning morphological variants are treated as different words. I don't know if this is consistent with Bernhardsson's methodology or not.

December 11, 2009

A Troop of One

Today is Veterans Day in the United States, and linguist Neal Whitman has been thinking about a question of military usage: if "50,000 troops" refers to 50,000 people, then does "one troop" refer to one person?

The horrifying events at Fort Hood last week have shown us that even in friendly territory, members of the armed forces may lose their lives in service to the rest of us. Whether overseas in active war zones, or within the relative safety of our own borders, our troops willingly accept that doing their job may require the ultimate sacrifice, and their acceptance humbles and amazes me. Regardless of our views on where the United States should send our troops or for what purposes, we can agree that the men and women who serve have undertaken a dangerous but necessary job that would scare most of us, and for that we owe them our thanks. If you see a troop today, or if you know one, please express your appreciation to him or her. If you are a troop, thank you.

Many speakers might have had a problem with those last two sentences — not for the sentiment (I hope), but for the way I used the word troop. How can one person be a troop? Shouldn't I have said soldier, or some suitable word, perhaps servicemember, that would cover members of any branch of the armed forces?

To tell you the truth, I have a problem with that usage of troop myself. I remember reading in my U.S. history class in high school about how President Kennedy at some point had sent 11,300 troops to Vietnam. I wondered how many people that actually translated to. Were there 20 soldiers in a troop? Fifty? It wasn't until I was in college that I finally began to figure out that for sufficiently large values of X, X troops simply meant X people.

In his Political Dictionary published last year, William Safire had this to say on the issue:

Troops is a word in semantic trouble. In one sense, it means "soldiers"; does this exclude sailors and airmen (now grouped as "service personnel")? Troops means "a group of," but so does a troop; the extent of the number is fuzzy. ... A troop means both "one soldier" and "a group of soldiers," which is not what a word is supposed to do.

Before I go any further, I need to lay out some terminology. Troop in the sense of "a group of soldiers" is an example of a collective noun, like group, family, or collection. If you use the plural form troops to mean "more than one group of soldiers" — the meaning I insisted upon in high school — it's a plural collective noun. I will refer to the use of troops to mean "a group of soldiers" and troop to mean "one soldier" as noncollective troop(s).

Concern about noncollective troops seems to be a 21st-century phenomenon. The earliest complaint I've found is in a letter from one of columnist Barbara Wallraff's readers, quoted in her book Word Court (2000). The reader describes an experience much like mine: When history teachers mentioned statistics like 50,000 troops in Vietnam, they wondered if that translated to half a million soldiers. The reader's position is uncompromising: if 50,000 troops means 50,000 people, then one troop is one person, and that's wrong.

A similarly hard line comes from Susan Jacoby, in her book The Age of American Unreason (2008):

As every dictionary makes plain, the word "troop" is always a collective noun; the "s" is added when referring to a particularly large military force.

So for exactly how long has troops been used noncollectively? The Oxford English Dictionary's earliest citation of troop is from 1545, and Google Books turns up plenty of citations from the 1600s onward with troops following large, round numbers. A few of the hits contain enough math clues to make it clear that noncollective troops is intended. For example, in The History and Proceedings of the Second Session of the Third Parliament of King George II, 1742-1743, we find:

[I]f we take 16,000 into our Pay, fresh Troops must be raised for that Purpose, and, I hope, I may say, without any Derogation, that 16,000 Hanoverians newly raised, are not so good as 16,000 of the Veteran Troops of any other Potentate in Europe.

Clearly, the number 16,000 that keeps coming up refers to 16,000 people.

Moving down an order of magnitude, there are references to hundreds of troops from the same time period. In 1757, during the French and Indian War, the minutes of a colonial governors' meeting contain a breakdown of one army to be assembled, including "200 Provincial Troops from Pensilvania" and "200 Troops from North Carolina", which, together with other groups totaling 1600 soldiers, make a grand total of "2000 men."

So the use of noncollective troops with large, round numbers is nothing recent; it has been going on for at least 250 years, and probably longer. Wallraff's response to her reader recognizes this. After pointing out the need for a term to cover any member of the armed services, she writes:

We are to ignore the special qualities of troops when it appears in a context like "5,000 troops were sent overseas," in which it means a body of soldiers and the number is indicating the size of that body.

Patricia O'Conner, author of Woe is I, wrote something similar on her Grammarphobia blog in 2006:

"Troops" (plural), in the military sense, properly refers to a LARGE number of individuals (as in, "Five thousand troops were deployed.") When "troops" is used with a SMALL number, it properly refers not to individuals but to collections of people. "Three troops were attacked" would mean three units were attacked.

So the singular "troop" in reference to an individual soldier ("one troop was slightly injured") is considered a misuse, as is the plural "troops" to refer to a small number of individuals ("four troops were captured"). We do hear and read such irregular usages, though.

The earliest complaint I've been able to find about troops with these smaller numbers starts appears only a year after Wallraff's reader's complaint about troops with any number, but during that interval came the attacks of 9/11 and the beginning of the war on terror(ism). The news reports from Iraq and Afghanistan have told stories involving roadside bombs, rocket-propelled grenades, suicide bombers, and frequent, small numbers of deaths, both of civilians and military personnel. The complaint appeared on the alt.usage.english online forum in December 2001, and took issue with the New York Times's phrasing "four Arab al Qaeda troops."

Complaints for various numbers of troops less than 100 have appeared in one forum or other nearly every year since. In 2005, James Kilpatrick responded to some of his readers, who were complaining about 19 and 31 troops. In 2003, northern California columnist Debra DeAngelo ranted about "twelve troops." Patricia O'Conner's blog entry from 2006 was in response to a reader's complaint about "2 or 3 troops." In fact, "two troops" attracts quite a bit of attention. In April 2004, I myself blogged about a headline mentioning "two troops," and a year later, Geoff Pullum wrote on Language Log about hearing "two troops" on NPR. In November 2007, Ralph Harrington at the Grey Cat Blog complained about a BBC story that referred to "two Danish troops." "If one individual had been reported dead," he asked, "would the headline have referred to the killing of 'one Danish troop'?"

That's the question, isn't it? What happens when we take noncollective troops to its mathematical conclusion, and find ourselves with one troop corresponding to one person, in the most overt possible ambiguity with the collective noun troop? In her blog post, O'Conner rejects such a usage for the same reason as for rejecting troops with small numbers. (Just last week, O'Conner updated the post to allow for small numbers with troops, but is still silent on the singular.) Other grammarians, while grudgingly allowing troops with small numbers, draw the line at one troop. A post from November 2001 on a grammar blog associated with Capital Community College in Hartford, Connecticut states, "It seems awfully clumsy to refer to an individual as a 'troop,' though, and I don't think that happens (even though it would seem logical enough)."

In a 2007 essay on NPR, linguist John McWhorter states, "One cannot refer to a single soldier as a troop." Bryan Garner concurs in the third edition of his Modern American Usage (2009), which finally takes on the question of noncollective troops. Even while admitting as standard phrases such as three troops (for which he provides an 1853 citation), Garner writes that "a single soldier, sailor, or pilot would never be termed a troop."

But it had to happen sometime. On June 9, 2008, Ralph Harrington returned with another post, informing us that CBS News had broken a "stupidity barrier" when they put out the headline "Afghan violence claims 100th British troop."

So the "one troop" stupidity barrier was crossed just last year. Or was it? O'Conner mentioned having read or heard "such forms" prior to August 2006, and the Corpus of Contemporary American English contains this example from 2002: "We do know that there was one troop killed over the weekend." On November 24, 1990, during the Persian Gulf crisis, President George Bush stated, "As long as I have one American troop — one man, one woman — out there, I will work closely with all those who stand up against this aggression."

When members of the American Dialect Society's listserv talked about noncollective troop in April 2003, one recalled being addressed individually as "troop" while in the Army during the late 1960s. The Oxford English Dictionary has "Can you spare a bite for a front-line troop?" from 1947. The June 20, 1874 Sydney Mail reports, "One of the troops was severely handled, being beaten with clubs, and two others received shot wounds." And finally, the OED provides the earliest citation yet, from 1823: "The monkey stowed himself away..till the same marine passed.., and laid hold of him by the calf of the leg... As the wounded 'troop' was not much hurt, a sort of truce was proclaimed."

In short, although noncollective troop(s) with small numbers (including one) has been in existence for almost 200 years, language watchers noticed it only recently. Some reject it with any number; some allow it only with large numbers; some allow it with any number greater than one. And one writer has gone so far as to allow it with any number, including one. It was none other than the curmudgeonly James Kilpatrick who wrote in his 2005 column: "In today's nomenclature, a troop is also an individual soldier."

source: http://www.visualthesaurus.com/cm/dictionary/2062/

Linguistic Predictions

Introduction

A word of comfort to those of you who are unhappy about an important and so far unique and unprecedented linguistic development underway in the world and the present shift in English identity to a world property or world identity.

The rise of global English doesn't mean the loss of a geographical identity. People in Britain remain British, those in the US stay American. A model is Switzerland. You can be German, French or Italian but you are ultimately Swiss. Even now the term Spanish is misleading because it doesn't consider the Catalans and others.

Although I still believe the loss of a linguistic identity has far more advantages than drawbacks, other languages in the world won't be spared. A much more disastrous fate is awaiting them. It will also free us from the shackles of nationalism and arrogance. We will save time and energy which we still put into translation. A global English might overrun and replace lots of languages (compare French and la grande nation). So everybody in the world will acquire a new linguistic identity. I wonder, what's wrong with that. After all languages come and go. This is a rule of nature. Just think of what happened to Latin. It gave birth before it died to a number of daughters Like: Italian, Spanish, and Portuguese....

On the other hand a lot of people are worried that the English invasion will bring with it American and English culture. They believe it's a kind of cultural colonialism. The whole world will be Anglo-Saxonized or Americanized. Even in China people have already started eating MacDonald's hamburgers (symbolic). There is no reason to worry that Chinese (Mandarin) one day will be a more powerful language than English. Chinese sounds are more difficult to the people of the world. In addition its writing system is cumbersome and not suitable for international communication. Human beings have always lamented the demise of the present. Shelly's "West Wind" makes way for a new life.

Language is the basis for all human interactions. No thinking is possible without language. It is indeed surprising why linguistics and language have not got the attention and focus they deserve till now. After all, human knowledge and science is only possible through the medium of language. Acquiring a new language opens new perspectives on life and broadens our minds. Our identity is still based on language. However, in case of English and because of its international role the idea of identity might get lost or is in deed on its way to be lost (it is becoming a world or a human identity). Yet, English culture and literature can only be seen in English. That's why although translation is a brilliant discipline it cannot overcome its deficiencies. Just imagine translating Shakespeare or Coleridge into Dutch.In short, language is superior to all other disciplines and is always first in sequence. It is knowledge, pleasure, emotions (mind and body or body and soul in one). After all we are not only heads but bodies too.
Density and Speed

One question which doesn't leave my mind is: Can the human language we know cope with the information density and the speed we are experiencing. As you know, information is growing and coding or packing information into language is necessary. Informal language uses more verbs. It is verbal or verbose (i.e. information density is low). Academic language tries to avoid verbs as much as possible apart from some basic ones like: be, have and a couple of other high frequency ones. So nominalization is a feature of Academic English because of the problem of information density. You can do more operations on nouns than on verbs. For example you can count nouns; use adjectives with them (describe them) etc... But nominalization means using nouns and as you knows most of them, at least the academic ones, are of Romance origin i.e. very long words (multi-syllabic). What will be if this information density grows to such an extent that the present human language is no more capable of packing information. On the other hand human language is slow for communicating messages. You need more time to pronounce long Romance words. Will there be a new communication medium to cope with the problems of density and speed.
Ambiguity, Density, Identity and Speed

Now I would like to extend the ideas of density and speed which I touched upon to two additional important phenomena which are part and parcel of human language.
Ambiguity

Some people complain about the imprecise use of language. If human language won't suffice to code information in the future due to information density and speed and of course due to misunderstandings (ambiguity) then the analogue language we have now might make way for a digital language Nowadays, everything is becoming digital. Why not human language? Once human language is digitized or replaced by a digital (computer) language (whether prescribed or agreed on), not only ambiguity ends but also beauty and mysticism and culture of human heritage. This means there will definitely be advantages of density, speed and clarity but a lot of disadvantages as I already mentioned will ensue foremost among those is reduction. Digital data is compressed or zipped. Compression means losing part of the information which is beyond human perception. Thus! , Digitalization means reducing human language to two modes, there is current or no current, a duality of yes and no like vending machines or computers. It is always a win/lose situation. This is an economic principle. We have to make a decision and set priorities.
Identity

Another problem of human language is identity. Identity doesn't only help us to belong to a nation and provide a profile but also create big human conflicts. Just take nationalism which is not only based on skin colour and facial features but also on language.

There are different peoples (nations) in the world: In the Middle East there are Arabs, Turks, Kurds or Persians. In East Asia there are: Chinese, Koreans, and Vietnamese. In Germanic Europe there are: the English, the Germans, and the Dutch etc. In Romance Europe: the French, the Italians, the Spanish. All these People look alike in their own parts of the world and it is difficult to tell them a part like one cell twins but they have lots of conflicts mostly based on nationalism (language identity). Using a digital language might solve some national conflicts. Human relations will probably then be based on economy and not languages.
Growth

Perhaps one day we will cease to be bodies or at least some parts of our bodies will be left as remnants based on a different anatomy (giving way to big heads) in the process of human evolution. The problem is not one of technology as much as that of growth .The pace of information growth is scaring. This is in deed a gloomy picture (at least to us now). Our biggest problem and enemy is ultimately growth not only that of information. Every thing is growing: world population is growing, pollution is growing, and economy is growing. People usually think it is positive but any growth means more consumption and more damage. This means we have to set priorities. I personally find it difficult to cope with the information overload. Sometimes I develop interest in a variety of issues and I find myself lost. It has already become difficult to make a choice or a decision.

Maths or translating language into a digital or formal language can only substitute natural language by means of reduction. Reduction means parts of our analogue language (like intonation and other features) are lost because digital or mathematical language stops the language flow, creates boundaries and compresses data as I already mentioned. So because of information density, speed and the possible need for a language that doesn’t allow ambiguity. The language we know now might change or be replaced. I mean human beings have already thought of a language like Esperanto void of identity based on linguistic differences which has caused a lot of human suffering. I am not saying this is what I personally prefer. I know the price we have to pay but our present language has to cope with the big challenges it is going to face in the future.

Suppose extraterrestrials landed on our Earth I am sure they would be greatly surprised to find that people on a small planet like Earth can neither communicate freely and direct (without the help of an intermediary i.e. an interpreter or a translator) nor can get in touch easily. Their astonishment would grow further when they find out how much suffering this linguistic diversity has caused so far. People's nationalities have been defined foremost on linguistic grounds. Languages create different identities and cultures. This on the one hand makes the world more interesting but on the other hand prepares ground for tragic conflicts. The loss of a linguistic identity doesn't necessarily mean the loss of a geographic or ethnic identity but it will mean one obstacle being removed.

I would like to elaborate on this issue by drawing a comparison from economy. In Europe, we used to have different currencies and in a sense it was a nice feeling to see foreign currencies when we were on holiday but no one can deny that the introduction of a single currency has also solved a lot of problems. The advantages of a single currency certainly outweigh the disadvantages by far if any. The EURO has set an example and paved the way for a single, at least, official European language. People can go on using their languages but we need a common European or world language to make us strong in unity. People are afraid of the loss of cultures and languages. However, cultures and languages have never been static. They are destined to change. In the age of globalization, information and communication technology, satellite TV and fast travelling the gap has become narrower. In addition, we can save more time and energy when everybody can communicate without any linguistic barriers. I hope one day the present state becomes history and we can say: A long time ago people on our Earth used to have different languages and we wonder how they could cope without a single world language.

English has already become global and no more the property of a certain community. Nobody feels at a disadvantage when speaking English. It's no more Germanic in quality because it has at least incorporated vocabulary from nearly all languages in the world. This makes English indeed global. Every nation can find a bit of its linguistic heritage integrated into English. Moreover, its writing system has no diacritical points as in some languages which make them difficult to use, pronounce and communicate in writing. English has become a powerful tool and rich; it has become simple and complex at the same time. Gender is nearly non-existent because the article remains "the" whatever the gender and the position in a sentence. There are no difficult case endings and sounds as in lots of other languages which sometimes constitute insurmountable barriers. We can express any idea most powerfully and precisely. It has reached the level of maturity and deserves the label of a global language. In addition, it sounds beautiful and appealing to the ear or at least acceptable and learnable by the majority of people. Finally, global English is experiencing a simplification of its grammar and phonetics.
Density, Memory and Speed

Information is increasing on a daily basis. Our human knowledge has grown, is still growing and will continue to grow due to advances in nearly all disciplines. We are already experiencing information overload. The question is what will be in 50or so years? Even computers are facing difficulty with memory challenges and new search engines like Google are adapted to more effective ways of information storage and retrieval. Academic language takes refuge in nominalization. Our present languages are not prepared to keep pace with such density and speed not experienced before and it is not clear whether our memories and brains can accommodate and cope with these developments.
Global English

English is experiencing a development unprecedented in human history. Every day there are new speakers of English. Again what will be in 50 years or so? The impact of this new situation will be three-fold:

* First, the dominance of English will be to the "disadvantage" of other languages and cultures.
* Second, loss of linguistic identity (to the "disadvantage" of the so-called "native speakers") which might have advantages since linguistic variety despite its beauty has been a source of tragic conflicts in the world. No more BE or AmE but a global English.
* Third, English will change in its new role to accommodate other cultures and languages. People have already started mixing BE and AmE and don't care about the pedantic view not to mix the two varieties.

Of course, in every change and development lie advantages and disadvantages but perhaps the advantages outweigh the advantages. In addition to what is mentioned, we will save a lot of time and energy spent on translations and thus communications barriers are removed. This doesn't mean that a Global English in turn won't change or split but it is a fact that people in our global village and a 24-hour society are not kept apart as they used to by geographical barriers. We can communicate freely and quickly through the medium of Global English. We have already reached the age of more direct and instant contact. Future linguistic changes consequently will be of a different nature.
The principle of convention and democracy in change

This is not plea for an artificial replacement of the present world languages. On the contrary natural languages are a means of social cuddling and cannot be changed by the dictatorship of minorities. It's always the dictatorship of the majority. There will be no revolution but changes are already underway. It's a fact that a Global English is already underway to overrun many a language.
The versatile and analogue character of Natural Languages

Natural (analogue) languages are certainly superior to artificial or digital languages because they accommodate nearly all of our present needs and can be used for all disciplines. However, we have already developed mathematical and computer languages to satisfy specific needs. In addition, natural languages leave room for ambiguity which still might be very useful to satisfy certain human needs (like literature, playing on words, implications, jokes ....) but can also be a source of misunderstanding. On the other hand digitization is simple, boring and poor cannot satisfy our present social needs although it is precise, mathematical, can be reduced, stored and manipulated.

As already mentioned these are only predictions made due to a variety of changes and developments in the last 20 years or so. Perhaps the most threatening force is growth. This word might sound positive but is in fact behind a lot of evil. Just imagine everything is growing, Earth population, economy (which means more consumption and Pollution and more...). Human knowledge has grown exponentially. This is the reality and not fiction even though a lot of trash is being produced daily but you can’t make an omelette without breaking eggs. In order to cope with this growth we need resources. For example we need food for the growing population, but producing and consuming food means in turn more pollution, more damage...

As far as human knowledge is concerned we also need resources to store and retrieve information. There are big advances in science and technology and our knowledge is growing on a daily basis. Just take the number of books, websites published everyday in comparison what was some years ago.

The computer networks worldwide, the phone, TV (satellite, cable, terrestrial), internet (email and the web), modern airlines have made it possible to contact each other just in-time, interact with each other, discuss issues online, share work, brainstorm ideas, pool resources and so on much faster and more productively. There are practically no boundaries left. All sciences are linked and have become inter-disciplinary. In Europe the EU and the single currency have also removed borders. Thus growth or density of information necessitates a tool to communicate and interact faster. Human language might not be capable of keeping pace with this growth and speed. Academic language is compact uses more nouns than verbs (nominalization: independent of tense, aspect and mood) because you can pack more information into nouns than verbs, use many of them (cram your text but still not wordy). They are quieter, more objective and do a lot of other operations on them. For example, you can count them, modify them.... Verbs in comparison are verbal or verbose (more talk than matter). They i.e. verbs are more subjective, dynamic (the majority of verbs are dynamic not stative and even among the limited number of stative verbs which exist some behave dynamically. Verbs of high frequency belong to informal register, show change and are conjugated which you don't have in nouns. Nouns are static (almost lifeless), neutral to change and emotions and more objective. Counting the number of nouns and verbs in a page of an academic paper will show this tendency. In addition, verbs are subordinate to nouns because they relate nouns to each other like prepositions and we can do only with a few of them. Some are transitive with one object or some have two open connections. So we need something beyond English either as an adapted natural language or an artificial functioning next to our natural one. It can be any tool.

However, human language is beautiful, encodes more than linguistic information, and allows room for ambiguity. There are a lot of implications and layers in human language next to the basic linguistic layer. It is analogue, has no boundaries and is far much superior to mathematical or digital languages.

source: http://www.usingenglish.com/articles/linguistic-predictions.html

Linguistics

Linguistics is the scientific[1][2] study of natural language.[3][4] Linguistics encompasses a number of sub-fields. An important topical division is between the study of language structure (grammar) and the study of meaning (semantics and pragmatics). Grammar encompasses morphology (the formation and composition of words), syntax (the rules that determine how words combine into phrases and sentences) and phonology (the study of sound systems and abstract sound units). Phonetics is a related branch of linguistics concerned with the actual properties of speech sounds (phones), non-speech sounds, and how they are produced and perceived. Other sub-disciplines of linguistics include the following: evolutionary linguistics, which considers the origins of language; historical linguistics, which explores language change; sociolinguistics, which looks at the relation between linguistic variation and social structures; psycholinguistics, which explores the representation and functioning of language in the mind; neurolinguistics, which looks at the representation of language in the brain; language acquisition, which considers how children acquire their first language and how children and adults acquire and learn their second and subsequent languages; and discourse analysis, which is concerned with the structure of texts and conversations, and pragmatics with how meaning is transmitted based on a combination of linguistic competence, non-linguistic knowledge, and the context of the speech act.

Linguistics is narrowly defined as the scientific approach to the study of language, but language can, of course, be approached from a variety of directions, and a number of other intellectual disciplines are relevant to it and influence its study. Semiotics, for example, is a related field concerned with the general study of signs and symbols both in language and outside of it. Literary theorists study the use of language in artistic literature. Linguistics additionally draws on work from such diverse fields as psychology, speech-language pathology, informatics, computer science, philosophy, biology, human anatomy, neuroscience, sociology, anthropology, and acoustics.

Within the field, linguist is used to describe someone who either studies the field or uses linguistic methodologies to study groups of languages or particular languages. Outside the field, this term is commonly used to refer to people who speak many languages or have a great vocabulary.
Contents
[hide]

* 1 Names for the discipline
* 2 Fundamental concerns and divisions
* 3 Variation and universality
* 4 Structures
* 5 Selected sub-fields
o 5.1 Historical linguistics
o 5.2 Semiotics
o 5.3 Descriptive linguistics and language documentation
o 5.4 Applied linguistics
* 6 Description and prescription
* 7 Speech and writing
* 8 History
* 9 Schools of study
o 9.1 Generative grammar
o 9.2 Cognitive linguistics
* 10 See also
* 11 References
* 12 External links

Names for the discipline

Before the twentieth century, the term "philology", first attested in 1716,[5] was commonly used to refer to the science of language, which was then predominantly historical in focus.[6] Since Ferdinand de Saussure's insistence on the importance of synchronic analysis, however, this focus has shifted[7] and the term "philology" is now generally used for the "study of a language's grammar, history and literary tradition," especially in the United States,[8] where it was never as popular as it was elsewhere (in the sense of the "science of language").[5]

Although the term "linguist" in the sense of "a student of language" dates from 1641,[9] the term "linguistics" is first attested in 1847.[9] It is now the usual academic term in English for the scientific study of language.
Fundamental concerns and divisions

Linguistics concerns itself with describing and explaining the nature of human language. Relevant to this are the questions of what is universal to language, how language can vary, and how human beings come to know languages. All humans (setting aside extremely pathological cases) achieve competence in whatever language is spoken (or signed, in the case of signed languages) around them when growing up, with apparently little need for explicit conscious instruction. While non-humans acquire their own communication systems, they do not acquire human language in this way (although many non-human animals can learn to respond to language, or can even be trained to use it to a degree).[10] Therefore, linguists assume, the ability to acquire and use language is an innate, biologically-based potential of modern human beings, similar to the ability to walk. There is no consensus, however, as to the extent of this innate potential, or its domain-specificity (the degree to which such innate abilities are specific to language), with some theorists claiming that there is a very large set of highly abstract and specific binary settings coded into the human brain, while others claim that the ability to learn language is a product of general human cognition. It is, however, generally agreed that there are no strong genetic differences underlying the differences between languages: an individual will acquire whatever language(s) he or she is exposed to as a child, regardless of parentage or ethnic origin.[11]

Linguistic structures are pairings of meaning and form; such pairings are known as Saussurean signs. In this sense, form may consist of sound patterns, movements of the hands, written symbols, and so on. There are many sub-fields concerned with particular aspects of linguistic structure, ranging from those focused primarily on form to those focused primarily on meaning:

* Phonetics, the study of the physical properties of speech (or signed) production and perception
* Phonology, the study of sounds (or signs) as discrete, abstract elements in the speaker's mind that distinguish meaning
* Morphology, the study of internal structures of words and how they can be modified
* Syntax, the study of how words combine to form grammatical sentences
* Semantics, the study of the meaning of words (lexical semantics) and fixed word combinations (phraseology), and how these combine to form the meanings of sentences
* Pragmatics, the study of how utterances are used in communicative acts, and the role played by context and non-linguistic knowledge in the transmission of meaning
* Discourse analysis, the analysis of language use in texts (spoken, written, or signed)

Many linguists would agree that these divisions overlap considerably, and the independent significance of each of these areas is not universally acknowledged. Regardless of any particular linguist's position, each area has core concepts that foster significant scholarly inquiry and research.

Alongside these structurally-motivated domains of study are other fields of linguistics, distinguished by the kinds of non-linguistic factors that they consider:

* Applied linguistics, the study of language-related issues applied in everyday life, notably language policies, planning, and education. (Constructed language fits under Applied linguistics.)
* Biolinguistics, the study of natural as well as human-taught communication systems in animals, compared to human language.
* Clinical linguistics, the application of linguistic theory to the field of Speech-Language Pathology.
* Computational linguistics, the study of computational implementations of linguistic structures.
* Developmental linguistics, the study of the development of linguistic ability in individuals, particularly the acquisition of language in childhood.
* Evolutionary linguistics, the study of the origin and subsequent development of language by the human species.
* Historical linguistics or diachronic linguistics, the study of language change over time.
* Language geography, the study of the geographical distribution of languages and linguistic features.
* Linguistic typology, the study of the common properties of diverse unrelated languages, properties that may, given sufficient attestation, be assumed to be innate to human language capacity.
* Neurolinguistics, the study of the structures in the human brain that underlie grammar and communication.
* Psycholinguistics, the study of the cognitive processes and representations underlying language use.
* Sociolinguistics, the study of variation in language and its relationship with social factors.
* Stylistics, the study of linguistic factors that place a discourse in context.

The related discipline of semiotics investigates the relationship between signs and what they signify. From the perspective of semiotics, language can be seen as a sign or symbol, with the world as its representation.[citation needed]
Variation and universality

Much modern linguistic research, particularly within the paradigm of generative grammar, has concerned itself with trying to account for differences between languages of the world. This has worked on the assumption that if human linguistic ability is narrowly constrained by human biology, then all languages must share certain fundamental properties.

In generativist theory, the collection of fundamental properties all languages share are referred to as universal grammar (UG). The specific characteristics of this universal grammar are a much debated topic. Typologists and non-generativist linguists usually refer simply to language universals, or universals of language.

Similarities between languages can have a number of different origins. In the simplest case, universal properties may be due to universal aspects of human experience. For example, all humans experience water, and all human languages have a word for water. Other similarities may be due to common descent: the Latin language spoken by the Ancient Romans developed into Spanish in Spain and Italian in Italy; similarities between Spanish and Italian are thus in many cases due to both being descended from Latin. In other cases, contact between languages — particularly where many speakers are bilingual — can lead to much borrowing of structures, as well as words. Similarity may also, of course, be due to coincidence. English much and Spanish mucho are not descended from the same form or borrowed from one language to the other;[12] nor is the similarity due to innate linguistic knowledge (see False cognate).

Arguments in favor of language universals have also come from documented cases of sign languages (such as Al-Sayyid Bedouin Sign Language) developing in communities of congenitally deaf people, independently of spoken language. The properties of these sign languages conform generally to many of the properties of spoken languages. Other known and suspected sign language isolates include Kata Kolok, Nicaraguan Sign Language, and Providence Island Sign Language.
Structures
Ferdinand de Saussure

It has been perceived that languages tend to be organized around grammatical categories such as noun and verb, nominative and accusative, or present and past, though, importantly, not exclusively so. The grammar of a language is organized around such fundamental categories, though many languages express the relationships between words and syntax in other discrete ways (cf. some Bantu languages for noun/verb relations, ergative-absolutive systems for case relations, several Native American languages for tense/aspect relations).

In addition to making substantial use of discrete categories, language has the important property that it organizes elements into recursive structures; this allows, for example, a noun phrase to contain another noun phrase (as in "the chimpanzee's lips") or a clause to contain a clause (as in "I think that it's raining"). Though recursion in grammar was implicitly recognized much earlier (for example by Jespersen), the importance of this aspect of language became more popular after the 1957 publication of Noam Chomsky's book Syntactic Structures,[13] which presented a formal grammar of a fragment of English. Prior to this, the most detailed descriptions of linguistic systems were of phonological or morphological systems.

Chomsky used a context-free grammar augmented with transformations. Since then, following the trend of Chomskyan linguistics, context-free grammars have been written for substantial fragments of various languages (for example GPSG, for English). It has been demonstrated, however, that human languages (most notably Dutch and Swiss German) include cross-serial dependencies, which cannot be handled adequately by context-free grammars.[14]
Selected sub-fields
Historical linguistics
Main article: Historical Linguistics

Historical linguistics studies the history and evolution of languages through the comparative method. Often the aim of historical linguistics is to classify languages in language families descending from a common ancestor. This evolves comparison of elements in different languages to detect possible cognates in order to be able to reconstruct how different languages have changed over time. This also involves the study of etymology, the study of the history of single words. Historical linguistics is also called "diachronic linguistics" and is opposed to "synchronic linguistics" that study languages in a given moment in time without regarding its previous stages.In universities in the United States, the historic perspective is often out of fashion. Historical linguistics was among the first linguistic disciplines to emerge and was the most widely practiced form of linguistics in the late 19th century. The shift in focus to a synchronic perspective started with Saussure and became predominant in western linguistics with Noam Chomsky's emphasis on the study of the synchronic and universal aspects of language.
Semiotics
Main article: Semiotics

Semiotics is the study of sign processes (semiosis), or signification and communication, signs and symbols, both individually and grouped into sign systems, including the study of how meaning is constructed and understood. Semioticians often do not restrict themselves to linguistic communication when studying the use of signs but extend the meaning of "sign" to cover all kinds of cultural symbols. Nonetheless semiotic disciplines closely related to linguistics are literary studies, discourse analysis, text linguistics, and philosophy of language.
Descriptive linguistics and language documentation
Main article: Descriptive linguistics

Since the inception of the discipline of linguistics linguists have been concerned with describing and documenting languages previously unknown to science. Starting with Franz Boas in the early 1900s descriptive linguistics became the main strand within American linguistics until the rise of formal structural linguistics in the mid 20th century. The rise of American descriptive linguistics was caused by the concern with describing the languages of indigenous peoples that were (and are) rapidly moving towards extinction. The ethnographic focus of the original Boasian type of descriptive linguistics occasioned the development of disciplines such as Sociolinguistics, anthropological linguistics, and linguistic anthropology, disciplines that investigate the relations between language, culture and society.

The emphasis on linguistic description and documentation has since become more important outside of North America as well, as the documentation of rapidly dying indigenous languages has become a primary focus in many of the worlds' linguistics programs. Language description is a work intensive endeavour usually requiring years of field work for the linguist to learn a language sufficiently well to write a reference grammar of it. The further task of language documentation requires the linguist to collect a preferably large corpus of texts and recordings of sound and video in the language, and to arrange for its storage in accessible formats in open repositories where it may be of the best use for further research by other researchers.[15]
Applied linguistics
Main article: Applied linguistics

Linguists are largely concerned with finding and describing the generalities and varieties both within particular languages and among all language. Applied linguistics takes the result of those findings and "applies" them to other areas. The term "applied linguistics" is often used to refer to the use of linguistic research in language teaching only[citation needed], but results of linguistic research are used in many other areas as well, such as lexicography and translation. "Applied linguistics" has been argued to be something of a misnomer[who?], since applied linguists focus on making sense of and engineering solutions for real-world linguistic problems, not simply "applying" existing technical knowledge from linguistics; moreover, they commonly apply technical knowledge from multiple sources, such as sociology (e.g. conversation analysis) and anthropology.

Today, computers are widely used in many areas of applied linguistics. Speech synthesis and speech recognition use phonetic and phonemic knowledge to provide voice interfaces to computers. Applications of computational linguistics in machine translation, computer-assisted translation, and natural language processing are areas of applied linguistics which have come to the forefront. Their influence has had an effect on theories of syntax and semantics, as modeling syntactic and semantic theories on computers constraints.

Linguistic analysis is a subdiscipline of applied linguistics used by many governments to verify the claimed nationality of people seeking asylum who do not hold the necessary documentation to prove their claim.[16] This often takes the form of an interview by personnel in an immigration department. Depending on the country, this interview is conducted in either the asylum seeker's native language through an interpreter, or in an international lingua franca like English.[16] Australia uses the former method, while Germany employs the latter; the Netherlands uses either method depending on the languages involved.[16] Tape recordings of the interview then undergo language analysis, which can be done by either private contractors or within a department of the government. In this analysis, linguistic features of the asylum seeker are used by analysts to make a determination about the speaker's nationality. The reported findings of the linguistic analysis can play a critical role in the government's decision on the refugee status of the asylum seeker.[16]
Description and prescription

Main articles: Descriptive linguistics, Linguistic prescription

Linguistics is descriptive; linguists describe and explain features of language without making subjective judgments on whether a particular feature is "right" or "wrong". This is analogous to practice in other sciences: a zoologist studies the animal kingdom without making subjective judgments on whether a particular animal is better or worse than another.

Prescription, on the other hand, is an attempt to promote particular linguistic usages over others, often favouring a particular dialect or "acrolect". This may have the aim of establishing a linguistic standard, which can aid communication over large geographical areas. It may also, however, be an attempt by speakers of one language or dialect to exert influence over speakers of other languages or dialects (see Linguistic imperialism). An extreme version of prescriptivism can be found among censors, who attempt to eradicate words and structures which they consider to be destructive to society.
Speech and writing

Most contemporary linguists work under the assumption that spoken (or signed) language is more fundamental than written language. This is because:

* Speech appears to be universal to all human beings capable of producing and hearing it, while there have been many cultures and speech communities that lack written communication;
* Speech evolved before human beings invented writing;
* People learn to speak and process spoken languages more easily and much earlier than writing;

Linguists nonetheless agree that the study of written language can be worthwhile and valuable. For research that relies on corpus linguistics and computational linguistics, written language is often much more convenient for processing large amounts of linguistic data. Large corpora of spoken language are difficult to create and hard to find, and are typically transcribed and written. Additionally, linguists have turned to text-based discourse occurring in various formats of computer-mediated communication as a viable site for linguistic inquiry.

The study of writing systems themselves is in any case considered a branch of linguistics.
History
Main article: History of linguistics

Some of the earliest linguistic activities can be recalled from Iron Age India with the analysis of Sanskrit. The Pratishakhyas (from ca. the 8th century BC) constitute as it were a proto-linguistic ad hoc collection of observations about mutations to a given corpus particular to a given Vedic school. Systematic study of these texts gives rise to the Vedanga discipline of Vyakarana, the earliest surviving account of which is the work of Pānini (c. 520 – 460 BC), who, however, looks back on what are probably several generations of grammarians, whose opinions he occasionally refers to. Pānini formulates close to 4,000 rules which together form a compact generative grammar of Sanskrit. Inherent in his analytic approach are the concepts of the phoneme, the morpheme and the root. Due to its focus on brevity, his grammar has a highly unintuitive structure, reminiscent of contemporary "machine language" (as opposed to "human readable" programming languages).

Indian linguistics maintained a high level for several centuries; Patanjali in the 2nd century BC still actively criticizes Panini. In the later centuries BC, however, Panini's grammar came to be seen as prescriptive, and commentators came to be fully dependent on it. Bhartrihari (c. 450 – 510) theorized the act of speech as being made up of four stages: first, conceptualization of an idea, second, its verbalization and sequencing (articulation) and third, delivery of speech into atmospheric air, the interpretation of speech by the listener, the interpreter.

Western linguistics begins in Classical Antiquity with grammatical speculation such as Plato's Cratylus. The first important advancement of the Greeks was the creation of the alphabet. As a result of the introduction of writing, poetry such as the Homeric poems became written and several editions were created and commented, forming the basis of philology and critic. The sophists and Socrates introduced dialectics as a new text genre. Aristotle defined the logic of speech and the argument. Furthermore Aristotle works on rhetoric and poetics were of utmost importance for the understating of tragedy, poetry, public discussions etc. as text genres.

One of the greatest of the Greek grammarians was Apollonius Dyscolus.[17] Apollonius wrote more than thirty treatises on questions of syntax, semantics, morphology, prosody, orthography, dialectology, and more. In the 4th c., Aelius Donatus compiled the Latin grammar Ars Grammatica that was to be the defining school text through the Middle Ages.[18] In De vulgari eloquentia ("On the Eloquence of Vernacular"), Dante Alighieri expanded the scope of linguistic enquiry from the traditional languages of antiquity to include the language of the day.[citation needed]

In the Middle East, the Persian linguist Sibawayh made a detailed and professional description of Arabic in 760, in his monumental work, Al-kitab fi al-nahw (الكتاب في النحو, The Book on Grammar), bringing many linguistic aspects of language to light. In his book he distinguished phonetics from phonology.[citation needed]

Sir William Jones noted that Sanskrit shared many common features with classical Latin and Greek, notably verb roots and grammatical structures, such as the case system. This led to the theory that all languages sprung from a common source and to the discovery of the Indo-European language family. He began the study of comparative linguistics, which would uncover more language families and branches.

In 19th century Europe the study of linguistics was largely from the perspective of philology (or historical linguistics). Some early-19th-century linguists were Jakob Grimm, who devised a principle of consonantal shifts in pronunciation – known as Grimm's Law – in 1822; Karl Verner, who formulated Verner's Law; August Schleicher, who created the "Stammbaumtheorie" ("family tree"); and Johannes Schmidt, who developed the "Wellentheorie" ("wave model") in 1872.

Ferdinand de Saussure was the founder of modern structural linguistics, with an emphasis on synchronic (i.e. non-historical) explanations for language form.

In North America, the structuralist tradition grew out of a combination of missionary linguistics (whose goal was to translate the bible) and Anthropology. While originally regarded as a sub-field of anthropology in the United States,[19][20] linguistics is now considered a separate scientific discipline in the US, Australia and much of Europe.

Edward Sapir, a leader in American structural linguistics, was one of the first who explored the relations between language studies and anthropology. His methodology had strong influence on all his successors. Noam Chomsky's formal model of language, transformational-generative grammar, developed under the influence of his teacher Zellig Harris, who was in turn strongly influenced by Leonard Bloomfield, has been the dominant model since the 1960s.

The structural linguistics period was largely superseded in North America by generative grammar in the 1950s and 60s. This paradigm views language as a mental object, and emphasizes the role of the formal modeling of universal and language specific rules. Noam Chomsky remains an important but controversial linguistic figure. Generative grammar gave rise to such frameworks such as Transformational grammar, Generative Semantics, Relational Grammar, Generalized Phrase-structure Grammar, Head-Driven Phrase Structure Grammar (HPSG) and Lexical Functional Grammar (LFG). Other linguists working in Optimality Theory state generalizations in terms of violable constraints that interact with each other, and abandon the traditional rule-based formalism first pioneered by early work in generativist linguistics.

Functionalist linguists working in functional grammar and Cognitive Linguistics tend to stress the non-autonomy of linguistic knowledge and the non-universality of linguistic structures, thus differing significantly from the formal approaches.
Schools of study

There are a wide variety of approaches to linguistic study. These can be loosely divided (although not without controversy) into formalist and functionalist approaches. Formalist approaches stress the importance of linguistic forms, and seek explanations for the structure of language from within the linguistic system itself. For example, the fact that language shows recursion might be attributed to recursive rules. Functionalist linguists by contrast view the structure of language as being driven by its function. For example, the fact that languages often put topical information first in the sentence, may be due to a communicative need to pair old information with new information in discourse.
Generative grammar
Main article: Generative grammar

During the last half of the twentieth century, following the work of Noam Chomsky, linguistics was dominated by the generativist school. While formulated by Chomsky in part as a way to explain how human beings acquire language and the biological constraints on this acquisition, in practice it has largely been concerned with giving formal accounts of specific phenomena in natural languages. Generative theory is modularist and formalist in character. Formal linguistics remains the dominant paradigm for studying linguistics,[21] though Chomsky's writings have also gathered much criticism.
Cognitive linguistics
Main article: Cognitive linguistics

In the 1970s and 1980s, a new school of thought known as cognitive linguistics emerged as a reaction to generativist theory. Led by theorists such as Ronald Langacker and George Lakoff, linguists working within the realm of cognitive linguistics posit that language is an emergent property of basic, general-purpose cognitive processes, though cognitive linguistics has also been the subject of much criticism.[22] In contrast to the generativist school of linguistics, cognitive linguistics is non-modularist and functionalist in character. Important developments in cognitive linguistics include cognitive grammar, frame semantics, and conceptual metaphor, all of which are based on the idea that form-function correspondences based on representations derived from embodied experience constitute the basic units of language.
See also
Main article: Outline of linguistics

* Cognitive science
* Speech-Language Pathology
* History of linguistics
* International Linguistics Olympiad
* Linguistics Departments at Universities
* Summer schools for linguistics
* List of linguists

Branches and fields

Anthropological linguistics, Semiotics, Philology, Discourse, Structuralism, Post-structuralism, Cognitive linguistics, Cognitive science, Comparative linguistics, Sociolinguistics, Varieties, Developmental linguistics, Discourse Analysis, Descriptive linguistics, Ecolinguistics, Embodied cognition, Endangered languages.

History of linguistics, Historical linguistics, Intercultural competence, Lexicography/Lexicology, Linguistic typology, Evolutionary linguistics.

Articulatory phonology, Biolinguistics, Computational linguistics, Biosemiotics, Articulatory synthesis, Machine translation, Natural language processing, Speaker recognition (authentication), Speech processing, Speech recognition, Speech synthesis, Concept Mining, Corpus linguistics, Critical discourse analysis, Cryptanalysis, Decipherment, Asemic writing, Grammar Writing.

Forensic linguistics, Global language system, Glottometrics, Integrational linguistics, International Linguistic Olympiad, Language acquisition, Language attrition, Language engineering, Language geography, Metacommunicative competence, Microlinguistics, Natural Language Processing, Neurolinguistics, Orthography, Reading, Second language acquisition, Sociocultural linguistics, Stratificational linguistics, Text linguistics, Writing systems, Xenolinguistics.
References

1. ^ Fromkin, Victoria; Bruce Hayes; Susan Curtiss, Anna Szabolcsi, Tim Stowell, Donca Steriade (2000). Linguistics: An Introduction to Linguistic Theory. Oxford: Blackwell. p. 3. ISBN 0631197117.
2. ^ Martinet, André (1960). Elements of General Linguistics. Tr. Elisabeth Palmer (Studies in General Linguistics, vol. i.). London: Faber. p. 15.
3. ^ Halliday, Michael A. K.; Jonathan Webster (2006). On Language and Linguistics. Continuum International Publishing Group. p. vii. ISBN 0826488242.
4. ^ Greenberg, Joseph (1948). "Linguistics and ethnology". Southwestern Journal of Anthropology 4: 140–47.
5. ^ a b Online Etymological Dictionary: philology
6. ^ McMahon, A. M. S. (1994), Understanding Language Change, Cambridge University Press, p. 19, ISBN 0-521-44665-1
7. ^ McMahon, A. M. S. (1994), Understanding Language Change, Cambridge University Press, p. 9, ISBN 0-521-44665-1
8. ^ A. Morpurgo Davies Hist. Linguistics (1998) 4 I. 22.
9. ^ a b Online Etymological Dictionary: linguist
10. ^ "Animal Language Article"
11. ^ Nevertheless, recent research suggests that even weak genetic biases in speakers may, over a number of generations, influence the evolution of particular languages, leading to a non-random distribution of certain linguistic features across the world. (Dediu, D. & Ladd, D.R. (2007). Linguistic tone is related to the population frequency of the adaptive haplogroups of two brain size genes, ASPM and Microcephalin, PNAS 104:10944-10949; summary available here)
12. ^ Much is from Middle English muchel, which is from Proto-Germanic *mekilaz[1], while mucho is from Latin multus[2].
13. ^ Chomsky, Noam. 1957. "Syntactic Structures". Mouton, The Hague
14. ^ Carl Vogel, Ulrike Hahn, Holly Branigan 1996, "Cross serial dependencies are not hard to process", Proceedings of the 16th conference on Computational linguistics - Volume 1
15. ^ Himmelman, Nikolaus Language documentation: What is it and what is it good for? in P. Gippert, Jost, Nikolaus P Himmelmann & Ulrike Mosel. (2006) Essentials of Language documentation. Mouton de Gruyter, Berlin & New York.
16. ^ a b c d Eades, Diana (2005). "Applied Linguistics and Language Analysis in Asylum Seeker Cases". Applied Linguistics 26 (4): 503–526. doi:10.1093/applin/ami021. http://songchau.googlepages.com/503.pdf.
17. ^ Apollonius Dyscolus
18. ^ linguistics : Greek and Roman antiquity -- Britannica Online Encyclopedia
19. ^ The "four fields" in American anthropology are cultural anthropology, physical anthropology, archeology and linguistics.
20. ^ Kemmer, Suzanne (2008). Biographical sketch of Franz Boas. Houston: Rice University. http://www.ruf.rice.edu/~kemmer/Found/boasbio.html.
21. ^ McMahon, A. M. S. (1994), Understanding Language Change, Cambridge University Press, p. 32, ISBN 0-521-44665-1
22. ^ See Newmeyer 1998, Language Form and Language Function (Cambridge, MA: MIT Press), and Culicover and Jackendoff 2005, Simpler Syntax (OUP)[3]

Fans