Retrieving arXiv, SocArXiv, and SSRN metadata for initial review screening

Rubia Fatima; Affan Yasin; Lin Liu; Jianmin Wang; Wasif Afzal

doi:10.1016/j.infsof.2023.107251

Retrieving arXiv, SocArXiv, and SSRN metadata for initial review screening

Rubia Fatima, Affan Yasin, Lin Liu^*, Jianmin Wang, Wasif Afzal

^*Corresponding author for this work

Research output: Contribution to journal › Article › peer-review

1 Citation (Scopus)

Abstract

Context: Researchers around the globe invest a lot of time searching the literature for performing reviews (Systematic Literature Review (SLR), Multivocal Literature Review (MLR)). The steps to performing the review includes inclusion of the grey literature, preprints, and quality assessed non-peer reviewed literature (the purpose is to minimize the publication bias). The initial screening of the papers takes time and bibliographic information is only available online for the researcher(s). Objective: Objective of our study is to propose, design, and develop a method that will help the research community to download the basic information of the papers (title, abstract, author) for the searched query from arxiv, SSRN, and SocArxiv (Social Science ArXiv). Method: We used Web scraping to extract data from the servers and save it in excel file. To retrieve the desired query from the databases, a Python code is used. Two methods have been discussed in the study to download the metadata of the searched query. Results: We have used different queries (such as “grey literature”, “testing software”, and “python” etc.) to see the results of our proposed method. Furthermore, we cross-verified the results with the online search results of the databases. Conclusion: Initial results from the preliminary pilot evaluations show that it is a viable method to search, download, and shortlist the research articles information (title, abstract etc.) from arXiv, SSRN, and SocArXiv. For external validity more evaluations are needed.

Original language	English
Article number	107251
Journal	Information and Software Technology
Volume	161
DOIs	https://doi.org/10.1016/j.infsof.2023.107251
Publication status	Published - Sept 2023
Externally published	Yes

Keywords

Bibliographic
Information retrieval
Initial screening
Metadata
Software engineering

Access to Document

10.1016/j.infsof.2023.107251

Cite this

@article{d9540faf96ee42b1a22d14e67cf3887c,

title = "Retrieving arXiv, SocArXiv, and SSRN metadata for initial review screening",

abstract = "Context: Researchers around the globe invest a lot of time searching the literature for performing reviews (Systematic Literature Review (SLR), Multivocal Literature Review (MLR)). The steps to performing the review includes inclusion of the grey literature, preprints, and quality assessed non-peer reviewed literature (the purpose is to minimize the publication bias). The initial screening of the papers takes time and bibliographic information is only available online for the researcher(s). Objective: Objective of our study is to propose, design, and develop a method that will help the research community to download the basic information of the papers (title, abstract, author) for the searched query from arxiv, SSRN, and SocArxiv (Social Science ArXiv). Method: We used Web scraping to extract data from the servers and save it in excel file. To retrieve the desired query from the databases, a Python code is used. Two methods have been discussed in the study to download the metadata of the searched query. Results: We have used different queries (such as “grey literature”, “testing software”, and “python” etc.) to see the results of our proposed method. Furthermore, we cross-verified the results with the online search results of the databases. Conclusion: Initial results from the preliminary pilot evaluations show that it is a viable method to search, download, and shortlist the research articles information (title, abstract etc.) from arXiv, SSRN, and SocArXiv. For external validity more evaluations are needed.",

keywords = "Bibliographic, Information retrieval, Initial screening, Metadata, Software engineering",

author = "Rubia Fatima and Affan Yasin and Lin Liu and Jianmin Wang and Wasif Afzal",

note = "Publisher Copyright: {\textcopyright} 2023 Elsevier B.V.",

year = "2023",

month = sep,

doi = "10.1016/j.infsof.2023.107251",

language = "English",

volume = "161",

journal = "Information and Software Technology",

issn = "0950-5849",

}

TY - JOUR

T1 - Retrieving arXiv, SocArXiv, and SSRN metadata for initial review screening

AU - Fatima, Rubia

AU - Yasin, Affan

AU - Liu, Lin

AU - Wang, Jianmin

AU - Afzal, Wasif

PY - 2023/9

Y1 - 2023/9

N2 - Context: Researchers around the globe invest a lot of time searching the literature for performing reviews (Systematic Literature Review (SLR), Multivocal Literature Review (MLR)). The steps to performing the review includes inclusion of the grey literature, preprints, and quality assessed non-peer reviewed literature (the purpose is to minimize the publication bias). The initial screening of the papers takes time and bibliographic information is only available online for the researcher(s). Objective: Objective of our study is to propose, design, and develop a method that will help the research community to download the basic information of the papers (title, abstract, author) for the searched query from arxiv, SSRN, and SocArxiv (Social Science ArXiv). Method: We used Web scraping to extract data from the servers and save it in excel file. To retrieve the desired query from the databases, a Python code is used. Two methods have been discussed in the study to download the metadata of the searched query. Results: We have used different queries (such as “grey literature”, “testing software”, and “python” etc.) to see the results of our proposed method. Furthermore, we cross-verified the results with the online search results of the databases. Conclusion: Initial results from the preliminary pilot evaluations show that it is a viable method to search, download, and shortlist the research articles information (title, abstract etc.) from arXiv, SSRN, and SocArXiv. For external validity more evaluations are needed.

AB - Context: Researchers around the globe invest a lot of time searching the literature for performing reviews (Systematic Literature Review (SLR), Multivocal Literature Review (MLR)). The steps to performing the review includes inclusion of the grey literature, preprints, and quality assessed non-peer reviewed literature (the purpose is to minimize the publication bias). The initial screening of the papers takes time and bibliographic information is only available online for the researcher(s). Objective: Objective of our study is to propose, design, and develop a method that will help the research community to download the basic information of the papers (title, abstract, author) for the searched query from arxiv, SSRN, and SocArxiv (Social Science ArXiv). Method: We used Web scraping to extract data from the servers and save it in excel file. To retrieve the desired query from the databases, a Python code is used. Two methods have been discussed in the study to download the metadata of the searched query. Results: We have used different queries (such as “grey literature”, “testing software”, and “python” etc.) to see the results of our proposed method. Furthermore, we cross-verified the results with the online search results of the databases. Conclusion: Initial results from the preliminary pilot evaluations show that it is a viable method to search, download, and shortlist the research articles information (title, abstract etc.) from arXiv, SSRN, and SocArXiv. For external validity more evaluations are needed.

KW - Bibliographic

KW - Information retrieval

KW - Initial screening

KW - Metadata

KW - Software engineering

UR - http://www.scopus.com/inward/record.url?scp=85162768520&partnerID=8YFLogxK

U2 - 10.1016/j.infsof.2023.107251

DO - 10.1016/j.infsof.2023.107251

M3 - Article

AN - SCOPUS:85162768520

SN - 0950-5849

VL - 161

JO - Information and Software Technology

JF - Information and Software Technology

M1 - 107251

ER -

Retrieving arXiv, SocArXiv, and SSRN metadata for initial review screening

Abstract

Keywords

Access to Document

Other files and links

Fingerprint

Cite this