Building a Chinese AMR Bank with Concept and Relation Alignments

Bin Li; Yuan Wen; Li Song; Weiguang Qu; Nianwen Xue

doi:10.33011/lilt.v18i.1429

Authors

Bin Li Nanjing Normal University
Yuan Wen Nanjing Normal University
Li Song Nanjing Normal University
Weiguang Qu Nanjing Normal University
Nianwen Xue Brandeis University

DOI:

https://doi.org/10.33011/lilt.v18i.1429

Abstract

Abstract Meaning Representation (AMR) is a meaning representation framework in which the meaning of a full sentence is represented as a single-rooted, acyclic, directed graph. In this article, we describe an on-going project to build a Chinese AMR (CAMR) corpus, which currently includes 10,149 sentences from the newsgroup and weblog portion of the Chinese TreeBank (CTB). We describe the annotation specifications for the CAMR corpus, which follow the annotation principles of English AMR but make adaptations where needed to accommodate the linguistic facts of Chinese. The CAMR specifications also include a systematic treatment of sentence-internal discourse relations. One significant change we have made to the AMR annotation methodology is the inclusion of the alignment between word tokens in the sentence and the concepts/relations in the CAMR annotation to make it easier for automatic parsers to model the correspondence between a sentence and its meaning representation. We develop an annotation tool for CAMR, and the inter-agreement as measured by the Smatch score between the two annotators is 0.83, indicating reliable annotation. We also present some quantitative analysis of the CAMR corpus. 46.71% of the AMRs of the sentences are non-tree graphs. Moreover, the AMR of 88.95% of the sentences has concepts inferred from the context of the sentence but do not correspond to a specific word.

Building a Chinese AMR Bank with Concept and Relation Alignments

Authors

DOI:

Abstract

Downloads

Published

How to Cite

Issue

Section

License

Information