TY - JOUR
T1 - Integrating lexical semantics for deep conceptual similarity
T2 - A WordNet-based vector framework
AU - Hussain, Muhammad Jawad
AU - Bai, Heming
AU - Seong, Myeongsu
AU - Wasti, Shahbaz Hassan
N1 - Publisher Copyright:
© 2026 Elsevier Inc.
PY - 2026/11/5
Y1 - 2026/11/5
N2 - Semantic similarity (SS) and semantic relatedness (SR) computation between concepts play a pivotal role in computational linguistics, supporting key applications such as information retrieval, text classification, and machine translation. Traditional knowledge-based approaches often rely solely on the taxonomic hierarchy of lexical resources like WordNet, overlooking non-taxonomic features and thereby limiting semantic richness. To address this, we propose the WordNet Relations-Based Model (WRBM), a novel vector-space framework that integrates both taxonomic relations (e.g., hypernyms and hyponyms) and non-taxonomic attributes (e.g., glosses, synonyms, examples, sister-terms, derivations, holonyms, and meronyms) to generate more expressive conceptual representations. We introduce an information-content (IC)-based mechanism to quantify the semantic contribution of each feature dimension, enabling the construction of low-dimensional, dense concept vectors that preserve semantic integrity while enhancing computational efficiency. Cosine similarity is employed to compute SS and SR between concepts. WRBM is evaluated on eight benchmark datasets, demonstrating consistent improvements over baseline models: 22.5% on MC30, 17.1% on RG65, 20.3% on WS203, 24.9% on SimLex, 43.8% on 353ALL, and 52.7% on MTurk287. These results highlight the robustness, scalability, and effectiveness of WRBM in advancing semantic computations and offer promising directions for future research in knowledge-based semantic modeling.
AB - Semantic similarity (SS) and semantic relatedness (SR) computation between concepts play a pivotal role in computational linguistics, supporting key applications such as information retrieval, text classification, and machine translation. Traditional knowledge-based approaches often rely solely on the taxonomic hierarchy of lexical resources like WordNet, overlooking non-taxonomic features and thereby limiting semantic richness. To address this, we propose the WordNet Relations-Based Model (WRBM), a novel vector-space framework that integrates both taxonomic relations (e.g., hypernyms and hyponyms) and non-taxonomic attributes (e.g., glosses, synonyms, examples, sister-terms, derivations, holonyms, and meronyms) to generate more expressive conceptual representations. We introduce an information-content (IC)-based mechanism to quantify the semantic contribution of each feature dimension, enabling the construction of low-dimensional, dense concept vectors that preserve semantic integrity while enhancing computational efficiency. Cosine similarity is employed to compute SS and SR between concepts. WRBM is evaluated on eight benchmark datasets, demonstrating consistent improvements over baseline models: 22.5% on MC30, 17.1% on RG65, 20.3% on WS203, 24.9% on SimLex, 43.8% on 353ALL, and 52.7% on MTurk287. These results highlight the robustness, scalability, and effectiveness of WRBM in advancing semantic computations and offer promising directions for future research in knowledge-based semantic modeling.
KW - Information content
KW - Semantic relatedness
KW - Semantic similarity
KW - Vector space
KW - WordNet
UR - https://www.scopus.com/pages/publications/105041520632
U2 - 10.1016/j.ins.2026.123757
DO - 10.1016/j.ins.2026.123757
M3 - Article
AN - SCOPUS:105041520632
SN - 0020-0255
VL - 755
JO - Information Sciences
JF - Information Sciences
M1 - 123757
ER -