TY - GEN
T1 - Unsupervised graph-based topic labelling using DBpedia
AU - Hulpus, Ioana
AU - Hayes, Conor
AU - Karnstedt, Marcel
AU - Greene, Derek
A2 - Stefano Leonardi, Alessandro Panconesi
PY - 2013
Y1 - 2013
N2 - Automated topic labelling brings benefits for users aiming at analysing and understanding document collections, as well as for search engines targetting at the linkage between groups of words and their inherent topics. Current approaches to achieve this suffer in quality, but we argue their performances might be improved by setting the focus on the structure in the data. Building upon research for concept disambiguation and linking to DBpedia, we are taking a novel approach to topic labelling by making use of structured data exposed by DBpedia. We start from the hypothesis that words co-occuring in text likely refer to concepts that belong closely together in the DBpedia graph. Using graph centrality measures, we show that we are able to identify the concepts that best represent the topics. We comparatively evaluate our graph-based approach and the standard text-based approach, on topics extracted from three corpora, based on results gathered in a crowd-sourcing experiment. Our research shows that graph-based analysis of DBpedia can achieve better results for topic labelling in terms of both precision and topic coverage.
AB - Automated topic labelling brings benefits for users aiming at analysing and understanding document collections, as well as for search engines targetting at the linkage between groups of words and their inherent topics. Current approaches to achieve this suffer in quality, but we argue their performances might be improved by setting the focus on the structure in the data. Building upon research for concept disambiguation and linking to DBpedia, we are taking a novel approach to topic labelling by making use of structured data exposed by DBpedia. We start from the hypothesis that words co-occuring in text likely refer to concepts that belong closely together in the DBpedia graph. Using graph centrality measures, we show that we are able to identify the concepts that best represent the topics. We comparatively evaluate our graph-based approach and the standard text-based approach, on topics extracted from three corpora, based on results gathered in a crowd-sourcing experiment. Our research shows that graph-based analysis of DBpedia can achieve better results for topic labelling in terms of both precision and topic coverage.
KW - dbpedia
KW - graph centrality measures
KW - latent dirichlet allocation
KW - topic labelling
UR - http://hdl.handle.net/10379/4528
UR - https://www.scopus.com/pages/publications/84874233631
U2 - 10.13025/21045
DO - 10.13025/21045
M3 - Conference Publication
SN - 9781450318693
T3 - WSDM 2013 - Proceedings of the 6th ACM International Conference on Web Search and Data Mining
SP - 465
EP - 474
BT - WSDM 2013 - Proceedings of the 6th ACM International Conference on Web Search and Data Mining
T2 - 6th ACM International Conference on Web Search and Data Mining, WSDM 2013
Y2 - 4 February 2013 through 8 February 2013
ER -