Abstract
Clustering is crucial for many NLP tasks and applications. However, evaluating the results of a clustering algorithm is hard. In this paper we focus on the evaluation setting in which a gold standard solution is available. We discuss two existing information theory based measures, V and VI, and show that they are both hard to use when comparing the performance of different algorithms and different datasets. The V measure favors solutions having a large number of clusters, while the range of scores given by VI depends on the size of the dataset. We present a new measure, NVI, which normalizes VI to address the latter problem. We demonstrate the superiority of NVI in a large experiment involving an important NLP application, grammar induction, using real corpus data in English, German and Chinese.
| Original language | English |
|---|---|
| Title of host publication | CoNLL 2009 - Proceedings of the 13th Conference on Computational Natural Language Learning |
| Editors | Suzanne Stevenson, Xavier Carreras |
| Publisher | Association for Computational Linguistics (ACL) |
| Pages | 165-173 |
| Number of pages | 9 |
| ISBN (Electronic) | 9781932432299 |
| DOIs | |
| State | Published - 2009 |
| Event | 13th Conference on Computational Natural Language Learning, CoNLL 2009 in conjunction with NAACL HLT - Boulder, United States Duration: 4 Jun 2009 → 5 Jun 2009 |
Publication series
| Name | CoNLL 2009 - Proceedings of the 13th Conference on Computational Natural Language Learning |
|---|
Conference
| Conference | 13th Conference on Computational Natural Language Learning, CoNLL 2009 in conjunction with NAACL HLT |
|---|---|
| Country/Territory | United States |
| City | Boulder |
| Period | 4/06/09 → 5/06/09 |
Bibliographical note
Publisher Copyright:© 2009 Association for Computational Linguistics.
Fingerprint
Dive into the research topics of 'The NVI Clustering Evaluation Measure'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver