Constructing Integrated Corpus and Lexicon Models for Multi-Layer Annotation in OWL DL
Keywords:corpus linguistics, lexicon
We present a general approach to formally modelling corpora with multi-layered annotation in a typed logical representation language, OWL DL. By defining abstractions over the corpus data, we can generalise from a large set of individual corpus annotations, thereby inducing a lexicon model. The resulting combined corpus and lexicon model can be interpreted as a graph structure that offers flexible querying functionality beyond current XML-based query languages. Its powerful methods for characterising and checking consistency can be used for incremental model refinement. In addition, the formalisation in a graph-based structure offers the means of defining flexible lexicon views over the corpus data. These views can be tailored for linguistic inspection or to define clean interfaces with other linguistic resources. We illustrate our approach by applying it to the syntactically and semantically annotated SALSA/TIGER corpus, a collection of German newspaper text.
How to Cite
This work is licensed under CC BY 4.0, which permits you to use, share, adapt, distribute, and reproduce it in any medium or format, provided you credit the original author(s) and source.