Corpus_linguistics

Corpus linguistics

Branch of linguistics that studies language through examples contained in real texts

Corpus linguistics is the study of a language as that language is expressed in its text corpus (plural corpora), its body of "real world" text. Corpus linguistics proposes that a reliable analysis of a language is more feasible with corpora collected in the field—the natural context ("realia") of that language—with minimal experimental interference. The large collections of text allow linguists to run quantitative analyses on linguistic concepts, otherwise harder to quantify.^[1]

The text-corpus method uses the body of texts written in any natural language to derive the set of abstract rules which govern that language. Those results can be used to explore the relationships between that subject language and other languages which have undergone a similar analysis. The first such corpora were manually derived from source texts, but now that work is automated.

Corpora have not only been used for linguistics research, they have also been used to compile dictionaries (starting with The American Heritage Dictionary of the English Language in 1969) and grammar guides, such as A Comprehensive Grammar of the English Language, published in 1985.

Experts in the field have differing views about the annotation of a corpus. These views range from John McHardy Sinclair, who advocates minimal annotation so texts speak for themselves,^[2] to the Survey of English Usage team (University College, London), who advocate annotation as allowing greater linguistic understanding through rigorous recording.^[3]

History

Some of the earliest efforts at grammatical description were based at least in part on corpora of particular religious or cultural significance. For example, Prātiśākhya literature described the sound patterns of Sanskrit as found in the Vedas, and Pāṇini's grammar of classical Sanskrit was based at least in part on analysis of that same corpus. Similarly, the early Arabic grammarians paid particular attention to the language of the Quran. In the Western European tradition, scholars prepared concordances to allow detailed study of the language of the Bible and other canonical texts.

English corpora

A landmark in modern corpus linguistics was the publication of Computational Analysis of Present-Day American English in 1967. Written by Henry Kučera and W. Nelson Francis, the work was based on an analysis of the Brown Corpus, which was a contemporary compilation of about a million American English words, carefully selected from a wide variety of sources.^[4] Brown's corpus was the first computerized corpus designed for linguistic research.^[5] Kučera and Francis subjected the Brown Corpus to a variety of computational analyses and then combined elements of linguistics, language teaching, psychology, statistics, and sociology to create a rich and variegated opus. A further key publication was Randolph Quirk's "Towards a description of English Usage" in 1960^[6] in which he introduced the Survey of English Usage. Quirk's corpus was the first modern corpus to be built with the purpose of representing the whole language.^[7]

Shortly thereafter, Boston publisher Houghton-Mifflin approached Kučera to supply a million-word, three-line citation base for its new American Heritage Dictionary, the first dictionary compiled using corpus linguistics. The AHD took the innovative step of combining prescriptive elements (how language should be used) with descriptive information (how it actually is used).

Other publishers followed suit. The British publisher Collins' COBUILD monolingual learner's dictionary, designed for users learning English as a foreign language, was compiled using the Bank of English. The Survey of English Usage Corpus was used in the development of one of the most important Corpus-based Grammars, which was written by Quirk et al. and published in 1985 as A Comprehensive Grammar of the English Language.^[8]

The Brown Corpus has also spawned a number of similarly structured corpora: the LOB Corpus (1960s British English), Kolhapur (Indian English), Wellington (New Zealand English), Australian Corpus of English (Australian English), the Frown Corpus (early 1990s American English), and the FLOB Corpus (1990s British English). Other corpora represent many languages, varieties and modes, and include the International Corpus of English, and the British National Corpus, a 100 million word collection of a range of spoken and written texts, created in the 1990s by a consortium of publishers, universities (Oxford and Lancaster) and the British Library. For contemporary American English, work has stalled on the American National Corpus, but the 400+ million word Corpus of Contemporary American English (1990–present) is now available through a web interface.

The first computerized corpus of transcribed spoken language was constructed in 1971 by the Montreal French Project,^[9] containing one million words, which inspired Shana Poplack's much larger corpus of spoken French in the Ottawa-Hull area.^[10]

Notes and references

[1]
Hunston, S. (1 January 2006), Brown, Keith (ed.), "Corpus Linguistics", Encyclopedia of Language & Linguistics (Second Edition), Oxford: Elsevier, pp. 234–248, doi:10.1016/b0-08-044854-2/00944-5, ISBN 978-0-08-044854-1, retrieved 31 October 2023
[2]
Sinclair, J. 'The automatic analysis of corpora', in Svartvik, J. (ed.) Directions in Corpus Linguistics (Proceedings of Nobel Symposium 82). Berlin: Mouton de Gruyter. 1992.
[3]
Wallis, S. 'Annotation, Retrieval and Experimentation', in Meurman-Solin, A. & Nurmi, A.A. (ed.) Annotating Variation and Change. Helsinki: Varieng, [University of Helsinki]. 2007. e-Published
[4]
Francis, W. Nelson; Kučera, Henry (1 June 1967). Computational Analysis of Present-Day American English. Providence: Brown University Press. ISBN 978-0870571053.
[5]
Kennedy, G. (1 January 2001), Smelser, Neil J.; Baltes, Paul B. (eds.), "Corpus Linguistics", International Encyclopedia of the Social & Behavioral Sciences, Oxford: Pergamon, pp. 2816–2820, ISBN 978-0-08-043076-8, retrieved 31 October 2023
[6]
Quirk, Randolph (November 1960). "Towards a description of English Usage". Transactions of the Philological Society. 59 (1): 40–61. doi:10.1111/j.1467-968X.1960.tb00308.x.
[7]
Kennedy, G. (1 January 2001), Smelser, Neil J.; Baltes, Paul B. (eds.), "Corpus Linguistics", International Encyclopedia of the Social & Behavioral Sciences, Oxford: Pergamon, pp. 2816–2820, doi:10.1016/b0-08-043076-7/03056-4, ISBN 978-0-08-043076-8, retrieved 31 October 2023
[8]
Quirk, Randolph; Greenbaum, Sidney; Leech, Geoffrey; Svartvik, Jan (1985). A Comprehensive Grammar of the English Language. London: Longman. ISBN 978-0582517349.
[9]
Sankoff, David; Sankoff, Gillian (1973). Darnell, R. (ed.). "Sample survey methods and computer-assisted analysis in the study of grammatical variation". Canadian Languages in Their Social Context. Edmonton: Linguistic Research Incorporated: 7–63.
[10]
Poplack, Shana (1989). Fasold, R.; Schiffrin, D. (eds.). "The care and handling of a mega-corpus". Language Change and Variation. Current Issues in Linguistic Theory. 52. Amsterdam: Benjamins: 411–451. doi:10.1075/cilt.52.25pop. ISBN 978-90-272-3546-6.
[11]
"National Center for Sign Language and Gesture Resources at B.U." www.bu.edu. Retrieved 31 October 2023.
[12]
Andersen, Francis I.; Forbes, A. Dean (2003), "Hebrew Grammar Visualized: I. Syntax", Ancient Near Eastern Studies, vol. 40, pp. 43–61 [45]
[13]
Eyland, E. Ann (1987), "Revelations from Word Counts", in Newing, Edward G.; Conrad, Edgar W. (eds.), Perspectives on Language and Text: Essays and Poems in Honor of Francis I. Andersen's Sixtieth Birthday, July 28, 1985, Winona Lake, IN: Eisenbrauns, p. 51, ISBN 0-931464-26-9
[14]
Dukes, K., Atwell, E. and Habash, N. 'Supervised Collaboration for Syntactic Annotation of Quranic Arabic'. Language Resources and Evaluation Journal. 2011.
[15]
"Digital Corpus of Sanskrit (DCS)". Retrieved 28 June 2022.
[16]
Wahle, Jan Philip; Ruas, Terry; Mohammad, Saif; Gipp, Bela (2022). "D3: A Massive Dataset of Scholarly Metadata for Analyzing the State of Computer Science Research". Proceedings of the Thirteenth Language Resources and Evaluation Conference. Marseille, France: European Language Resources Association: 2642–2651. arXiv:2204.13384.
[17]
Mohammad, Saif M. (2020). "NLP Scholar: A Dataset for Examining the State of NLP Research". Proceedings of the Twelfth Language Resources and Evaluation Conference. Marseille, France: European Language Resources Association: 868–877. ISBN 979-10-95546-34-4.
[18]
Bernardini, S. (1 January 2006), Brown, Keith (ed.), "Machine Readable Corpora", Encyclopedia of Language & Linguistics (Second Edition), Oxford: Elsevier, pp. 358–375, doi:10.1016/b0-08-044854-2/00476-4, ISBN 978-0-08-044854-1, retrieved 31 October 2023
[19]
Mainz, Johannes Gutenberg-Universität. "Corpus Linguistics | ENGLISH LINGUISTICS". Johannes Gutenberg-Universität Mainz (in German). Retrieved 31 October 2023.
[20]
Wallis, S. and Nelson G. Knowledge discovery in grammatically analysed corpora. Data Mining and Knowledge Discovery, 5: 307–340. 2001.
[21]
Baker, Paul; Egbert, Jesse, eds. (2016). Triangulating Methodological Approaches in Corpus-Linguistic Research. New York: Routledge.

Share this article:

This article uses material from the Wikipedia article Corpus_linguistics, and is written by contributors. Text is available under a CC BY-SA 4.0 International License; additional terms may apply. Images, videos and audio are available under their respective licenses.

[1] [1]
Hunston, S. (1 January 2006), Brown, Keith (ed.), "Corpus Linguistics", Encyclopedia of Language & Linguistics (Second Edition), Oxford: Elsevier, pp. 234–248, doi:10.1016/b0-08-044854-2/00944-5, ISBN 978-0-08-044854-1, retrieved 31 October 2023

[2] [2]
Sinclair, J. 'The automatic analysis of corpora', in Svartvik, J. (ed.) Directions in Corpus Linguistics (Proceedings of Nobel Symposium 82). Berlin: Mouton de Gruyter. 1992.

[3] [3]
Wallis, S. 'Annotation, Retrieval and Experimentation', in Meurman-Solin, A. & Nurmi, A.A. (ed.) Annotating Variation and Change. Helsinki: Varieng, [University of Helsinki]. 2007. e-Published

[4] [4]
Francis, W. Nelson; Kučera, Henry (1 June 1967). Computational Analysis of Present-Day American English. Providence: Brown University Press. ISBN 978-0870571053.

[5] [5]
Kennedy, G. (1 January 2001), Smelser, Neil J.; Baltes, Paul B. (eds.), "Corpus Linguistics", International Encyclopedia of the Social & Behavioral Sciences, Oxford: Pergamon, pp. 2816–2820, ISBN 978-0-08-043076-8, retrieved 31 October 2023

[6] [6]
Quirk, Randolph (November 1960). "Towards a description of English Usage". Transactions of the Philological Society. 59 (1): 40–61. doi:10.1111/j.1467-968X.1960.tb00308.x.

[7] [7]
Kennedy, G. (1 January 2001), Smelser, Neil J.; Baltes, Paul B. (eds.), "Corpus Linguistics", International Encyclopedia of the Social & Behavioral Sciences, Oxford: Pergamon, pp. 2816–2820, doi:10.1016/b0-08-043076-7/03056-4, ISBN 978-0-08-043076-8, retrieved 31 October 2023

[8] [8]
Quirk, Randolph; Greenbaum, Sidney; Leech, Geoffrey; Svartvik, Jan (1985). A Comprehensive Grammar of the English Language. London: Longman. ISBN 978-0582517349.

[9] [9]
Sankoff, David; Sankoff, Gillian (1973). Darnell, R. (ed.). "Sample survey methods and computer-assisted analysis in the study of grammatical variation". Canadian Languages in Their Social Context. Edmonton: Linguistic Research Incorporated: 7–63.

[10] [10]
Poplack, Shana (1989). Fasold, R.; Schiffrin, D. (eds.). "The care and handling of a mega-corpus". Language Change and Variation. Current Issues in Linguistic Theory. 52. Amsterdam: Benjamins: 411–451. doi:10.1075/cilt.52.25pop. ISBN 978-90-272-3546-6.

[11] [11]
"National Center for Sign Language and Gesture Resources at B.U." www.bu.edu. Retrieved 31 October 2023.

[12] [12]
Andersen, Francis I.; Forbes, A. Dean (2003), "Hebrew Grammar Visualized: I. Syntax", Ancient Near Eastern Studies, vol. 40, pp. 43–61 [45]

[13] [13]
Eyland, E. Ann (1987), "Revelations from Word Counts", in Newing, Edward G.; Conrad, Edgar W. (eds.), Perspectives on Language and Text: Essays and Poems in Honor of Francis I. Andersen's Sixtieth Birthday, July 28, 1985, Winona Lake, IN: Eisenbrauns, p. 51, ISBN 0-931464-26-9

[14] [14]
Dukes, K., Atwell, E. and Habash, N. 'Supervised Collaboration for Syntactic Annotation of Quranic Arabic'. Language Resources and Evaluation Journal. 2011.

[15] [15]
"Digital Corpus of Sanskrit (DCS)". Retrieved 28 June 2022.

[16] [16]
Wahle, Jan Philip; Ruas, Terry; Mohammad, Saif; Gipp, Bela (2022). "D3: A Massive Dataset of Scholarly Metadata for Analyzing the State of Computer Science Research". Proceedings of the Thirteenth Language Resources and Evaluation Conference. Marseille, France: European Language Resources Association: 2642–2651. arXiv:2204.13384.

[17] [17]
Mohammad, Saif M. (2020). "NLP Scholar: A Dataset for Examining the State of NLP Research". Proceedings of the Twelfth Language Resources and Evaluation Conference. Marseille, France: European Language Resources Association: 868–877. ISBN 979-10-95546-34-4.

[18] [18]
Bernardini, S. (1 January 2006), Brown, Keith (ed.), "Machine Readable Corpora", Encyclopedia of Language & Linguistics (Second Edition), Oxford: Elsevier, pp. 358–375, doi:10.1016/b0-08-044854-2/00476-4, ISBN 978-0-08-044854-1, retrieved 31 October 2023

[19] [19]
Mainz, Johannes Gutenberg-Universität. "Corpus Linguistics | ENGLISH LINGUISTICS". Johannes Gutenberg-Universität Mainz (in German). Retrieved 31 October 2023.

[20] [20]
Wallis, S. and Nelson G. Knowledge discovery in grammatically analysed corpora. Data Mining and Knowledge Discovery, 5: 307–340. 2001.

[21] [21]
Baker, Paul; Egbert, Jesse, eds. (2016). Triangulating Methodological Approaches in Corpus-Linguistic Research. New York: Routledge.

[1]

[2]

[3]

[4]

[5]

[6]

[7]

[8]

[9]

[10]

[11]

[12]

[13]

[14]

[15]

[16]

[17]

[18]

[19]

[20]

[21]

Corpus_linguistics

Corpus linguistics

History

English corpora

Multilingual Corpora

Ancient languages corpora

Corpora from specific fields

Methods

See also

Notes and references

Further reading

Books

Book series

Journals

External links

Share this article: