Abstract
Arabic is a widely-spoken language with a rich and long history spanning more than fourteen centuries. Yet existing Arabic corpora largely focus on the modern period or lack sufficient diachronic information. We develop a large-scale, historical corpus of Arabic of about 1 billion words from diverse periods of time. We clean this corpus, process it with a morphological analyzer, and enhance it by detecting parallel passages and automatically dating undated texts. We demonstrate its utility with selected case-studies in which we show its application to the digital humanities.
| Original language | Undefined/Unknown |
|---|---|
| Title of host publication | Proceedings of the Workshop on Language Technology Resources and Tools for Digital Humanities (LT4DH) |
| Place of Publication | Osaka, Japan |
| Pages | 45-53 |
| Number of pages | 9 |
| State | Published - 1 Dec 2016 |