Document Set Expansion with Positive-Unlabelled Learning Using Intractable Density Estimation

Haiyang Zhang; Qiuyi Chen; Yuanjie Zou; Yushan Pan; Jia Wang; Mark Stevenson

Document Set Expansion with Positive-Unlabelled Learning Using Intractable Density Estimation

Haiyang Zhang^*, Qiuyi Chen, Yuanjie Zou, Yushan Pan, Jia Wang, Mark Stevenson

^*Corresponding author for this work

Research output: Chapter in Book or Report/Conference proceeding › Conference Proceeding › peer-review

Abstract

The Document Set Expansion (DSE) task involves identifying relevant documents from large collections based on a limited set of example documents. Previous research has highlighted Positive and Unlabeled (PU) learning as a promising approach for this task. However, most PU methods rely on the unrealistic assumption of knowing the class prior for positive samples in the collection. To address this limitation, this paper introduces a novel PU learning framework that utilizes intractable density estimation models. Experiments conducted on PubMed and Covid datasets in a transductive setting showcase the effectiveness of the proposed method for DSE. Code is available from https://github.com/Beautifuldog01/Document-set-expansion-puDE.

Original language	English
Title of host publication	2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings
Editors	Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, Nianwen Xue
Publisher	European Language Resources Association (ELRA)
Pages	5167-5173
Number of pages	7
ISBN (Electronic)	9782493814104
Publication status	Published - 2024
Event	Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024 - Hybrid, Torino, Italy Duration: 20 May 2024 → 25 May 2024

Publication series

Name	2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings

Conference

Conference	Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024
Country/Territory	Italy
City	Hybrid, Torino
Period	20/05/24 → 25/05/24

Keywords

Density estimation
Document set expansion
Information retrieval
PU learning

Cite this

Zhang, H., Chen, Q., Zou, Y., Pan, Y., Wang, J., & Stevenson, M. (2024). Document Set Expansion with Positive-Unlabelled Learning Using Intractable Density Estimation. In N. Calzolari, M.-Y. Kan, V. Hoste, A. Lenci, S. Sakti, & N. Xue (Eds.), 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings (pp. 5167-5173). (2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings). European Language Resources Association (ELRA).

Zhang, Haiyang ; Chen, Qiuyi ; Zou, Yuanjie et al. / Document Set Expansion with Positive-Unlabelled Learning Using Intractable Density Estimation. 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings. editor / Nicoletta Calzolari ; Min-Yen Kan ; Veronique Hoste ; Alessandro Lenci ; Sakriani Sakti ; Nianwen Xue. European Language Resources Association (ELRA), 2024. pp. 5167-5173 (2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings).

@inproceedings{f0282c53eefe4c18881c4ae972909507,

title = "Document Set Expansion with Positive-Unlabelled Learning Using Intractable Density Estimation",

abstract = "The Document Set Expansion (DSE) task involves identifying relevant documents from large collections based on a limited set of example documents. Previous research has highlighted Positive and Unlabeled (PU) learning as a promising approach for this task. However, most PU methods rely on the unrealistic assumption of knowing the class prior for positive samples in the collection. To address this limitation, this paper introduces a novel PU learning framework that utilizes intractable density estimation models. Experiments conducted on PubMed and Covid datasets in a transductive setting showcase the effectiveness of the proposed method for DSE. Code is available from https://github.com/Beautifuldog01/Document-set-expansion-puDE.",

keywords = "Density estimation, Document set expansion, Information retrieval, PU learning",

author = "Haiyang Zhang and Qiuyi Chen and Yuanjie Zou and Yushan Pan and Jia Wang and Mark Stevenson",

note = "Publisher Copyright: {\textcopyright} 2024 ELRA Language Resource Association: CC BY-NC 4.0.; Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024 ; Conference date: 20-05-2024 Through 25-05-2024",

year = "2024",

language = "English",

series = "2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings",

publisher = "European Language Resources Association (ELRA)",

pages = "5167--5173",

editor = "Nicoletta Calzolari and Min-Yen Kan and Veronique Hoste and Alessandro Lenci and Sakriani Sakti and Nianwen Xue",

booktitle = "2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings",

}

Zhang, H, Chen, Q, Zou, Y, Pan, Y , Wang, J & Stevenson, M 2024, Document Set Expansion with Positive-Unlabelled Learning Using Intractable Density Estimation. in N Calzolari, M-Y Kan, V Hoste, A Lenci, S Sakti & N Xue (eds), 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings. 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings, European Language Resources Association (ELRA), pp. 5167-5173, Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024, Hybrid, Torino, Italy, 20/05/24.

Document Set Expansion with Positive-Unlabelled Learning Using Intractable Density Estimation. / Zhang, Haiyang; Chen, Qiuyi; Zou, Yuanjie et al.
2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings. ed. / Nicoletta Calzolari; Min-Yen Kan; Veronique Hoste; Alessandro Lenci; Sakriani Sakti; Nianwen Xue. European Language Resources Association (ELRA), 2024. p. 5167-5173 (2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings).

Research output: Chapter in Book or Report/Conference proceeding › Conference Proceeding › peer-review

TY - GEN

T1 - Document Set Expansion with Positive-Unlabelled Learning Using Intractable Density Estimation

AU - Zhang, Haiyang

AU - Chen, Qiuyi

AU - Zou, Yuanjie

AU - Pan, Yushan

AU - Wang, Jia

AU - Stevenson, Mark

PY - 2024

Y1 - 2024

N2 - The Document Set Expansion (DSE) task involves identifying relevant documents from large collections based on a limited set of example documents. Previous research has highlighted Positive and Unlabeled (PU) learning as a promising approach for this task. However, most PU methods rely on the unrealistic assumption of knowing the class prior for positive samples in the collection. To address this limitation, this paper introduces a novel PU learning framework that utilizes intractable density estimation models. Experiments conducted on PubMed and Covid datasets in a transductive setting showcase the effectiveness of the proposed method for DSE. Code is available from https://github.com/Beautifuldog01/Document-set-expansion-puDE.

AB - The Document Set Expansion (DSE) task involves identifying relevant documents from large collections based on a limited set of example documents. Previous research has highlighted Positive and Unlabeled (PU) learning as a promising approach for this task. However, most PU methods rely on the unrealistic assumption of knowing the class prior for positive samples in the collection. To address this limitation, this paper introduces a novel PU learning framework that utilizes intractable density estimation models. Experiments conducted on PubMed and Covid datasets in a transductive setting showcase the effectiveness of the proposed method for DSE. Code is available from https://github.com/Beautifuldog01/Document-set-expansion-puDE.

KW - Density estimation

KW - Document set expansion

KW - Information retrieval

KW - PU learning

UR - http://www.scopus.com/inward/record.url?scp=85195956973&partnerID=8YFLogxK

M3 - Conference Proceeding

AN - SCOPUS:85195956973

T3 - 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings

SP - 5167

EP - 5173

BT - 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings

A2 - Calzolari, Nicoletta

A2 - Kan, Min-Yen

A2 - Hoste, Veronique

A2 - Lenci, Alessandro

A2 - Sakti, Sakriani

A2 - Xue, Nianwen

PB - European Language Resources Association (ELRA)

T2 - Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024

Y2 - 20 May 2024 through 25 May 2024

ER -

Zhang H, Chen Q, Zou Y, Pan Y , Wang J, Stevenson M. Document Set Expansion with Positive-Unlabelled Learning Using Intractable Density Estimation. In Calzolari N, Kan MY, Hoste V, Lenci A, Sakti S, Xue N, editors, 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings. European Language Resources Association (ELRA). 2024. p. 5167-5173. (2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings).

Document Set Expansion with Positive-Unlabelled Learning Using Intractable Density Estimation

Abstract

Publication series

Conference

Keywords

Other files and links

Fingerprint

Cite this