Experimental data for computing semantic similarity between concepts using multiple inheritances in Wikipedia category graph

Muhammad Jawad Hussain; Shahbaz Hassan Wasti; Guangjian Huang; Yuncheng Jiang

doi:10.1016/j.dib.2020.105377

Experimental data for computing semantic similarity between concepts using multiple inheritances in Wikipedia category graph

Data Brief. 2020 Mar 10:30:105377. doi: 10.1016/j.dib.2020.105377. eCollection 2020 Jun.

Authors

Muhammad Jawad Hussain¹, Shahbaz Hassan Wasti^{1

2}, Guangjian Huang¹, Yuncheng Jiang¹

Affiliations

¹ School of Computer Science, South China Normal University, Guangzhou 510631, China.
² Division of Science and Technology, University of Education, Lahore, Pakistan.

Abstract

This data article compiles the detailed and descriptive experimental data of Wikipedia-based semantic similarity approach called as Neighbourhood Aggregated Semantic Contribution (NASC), presented in Husain, et al. [1]. The JWPL (Java Wikipedia Library)-DataMachine and JWPL WikipediaAPI are used to extract the required Wikipedia features from Wikipedia dump. The dataset presents the disambiguated Wikipedia concepts of the gold standard word similarity benchmarks MC30 (English), RG65_es (Spanish) and RG65_fr (French) and their associated set of categories in the corresponding Wikipedia category graph (WCG). The dataset also contains the number of ancestors, common ancestors, pages, and common pages in the k-neighbourhood of the associated categories for different levels of parameter k in the English, Spanish, and French WCGs. The presented dataset can be used to assess the semantic similarity between Wikipedia concepts in English (MC30), Spanish (RG65_es), and French (RG65_fr) languages benchmarks. Moreover, the dataset will be useful for the further analysis and comparison of the taxonomic structures of the English, Spanish, and French WCGs.

Keywords: Information content; Multiple inheritances; Semantic similarity; Wikipedia category graph.