Better tags give better trees – or do they?

Ines Rehbein; Hagen Hirschmann; Anke Lüdeling; Marc Reznicek

doi:10.33011/lilt.v7i.1279

Authors

Ines Rehbein Universität Potsdam
Hagen Hirschmann Humboldt-Universität zu Berlin
Anke Lüdeling Humboldt-Universität zu Berlin
Marc Reznicek Humboldt-Universität zu Berlin

DOI:

https://doi.org/10.33011/lilt.v7i.1279

Keywords:

treebank, annotation, POS tags, German

Abstract

Parsing learner data poses a great challenge for standard tools, since non-canonical and unusual structures may lead to wrong interpretations on the part of the taggers and parsers. It is well known that providing a statistical parser with perfect part-of-speech (POS) tags is of great benefit for parsing accuracy, and that parsing results can decrease considerably when the parser has to predict its own POS tags. Therefore one might expect that even small improvements in POS accuracy have a positive effect on parsing performance. In this paper we test this assumption and assess the impact of POS tag accuracy on constituency parsing for German learner language. We compare different strategies to manual correction of the learner text and specific POS tags, and we measure the time requirements for each strategy. We show that tagging a canonical equivalent of the non-canonical learner text substantially improves POS tag accuracy. Correcting selected POS tags can only lead to parsing results comparable to a setting where all POS tags are corrected, while reducing annotation time substantially. However, the manual corrections of the POS tags do not result in a statistically significant improvement for parsing, giving evidence for the high quality of the automatically predicted parts-of-speech for the corrected learner data.

Better tags give better trees – or do they?

Authors

DOI:

Keywords:

Abstract

Downloads

Published

How to Cite

Issue

Section

License

Information