TY - JOUR
T1 - Retrieving arXiv, SocArXiv, and SSRN metadata for initial review screening
AU - Fatima, Rubia
AU - Yasin, Affan
AU - Liu, Lin
AU - Wang, Jianmin
AU - Afzal, Wasif
N1 - Publisher Copyright:
© 2023 Elsevier B.V.
PY - 2023/9
Y1 - 2023/9
N2 - Context: Researchers around the globe invest a lot of time searching the literature for performing reviews (Systematic Literature Review (SLR), Multivocal Literature Review (MLR)). The steps to performing the review includes inclusion of the grey literature, preprints, and quality assessed non-peer reviewed literature (the purpose is to minimize the publication bias). The initial screening of the papers takes time and bibliographic information is only available online for the researcher(s). Objective: Objective of our study is to propose, design, and develop a method that will help the research community to download the basic information of the papers (title, abstract, author) for the searched query from arxiv, SSRN, and SocArxiv (Social Science ArXiv). Method: We used Web scraping to extract data from the servers and save it in excel file. To retrieve the desired query from the databases, a Python code is used. Two methods have been discussed in the study to download the metadata of the searched query. Results: We have used different queries (such as “grey literature”, “testing software”, and “python” etc.) to see the results of our proposed method. Furthermore, we cross-verified the results with the online search results of the databases. Conclusion: Initial results from the preliminary pilot evaluations show that it is a viable method to search, download, and shortlist the research articles information (title, abstract etc.) from arXiv, SSRN, and SocArXiv. For external validity more evaluations are needed.
AB - Context: Researchers around the globe invest a lot of time searching the literature for performing reviews (Systematic Literature Review (SLR), Multivocal Literature Review (MLR)). The steps to performing the review includes inclusion of the grey literature, preprints, and quality assessed non-peer reviewed literature (the purpose is to minimize the publication bias). The initial screening of the papers takes time and bibliographic information is only available online for the researcher(s). Objective: Objective of our study is to propose, design, and develop a method that will help the research community to download the basic information of the papers (title, abstract, author) for the searched query from arxiv, SSRN, and SocArxiv (Social Science ArXiv). Method: We used Web scraping to extract data from the servers and save it in excel file. To retrieve the desired query from the databases, a Python code is used. Two methods have been discussed in the study to download the metadata of the searched query. Results: We have used different queries (such as “grey literature”, “testing software”, and “python” etc.) to see the results of our proposed method. Furthermore, we cross-verified the results with the online search results of the databases. Conclusion: Initial results from the preliminary pilot evaluations show that it is a viable method to search, download, and shortlist the research articles information (title, abstract etc.) from arXiv, SSRN, and SocArXiv. For external validity more evaluations are needed.
KW - Bibliographic
KW - Information retrieval
KW - Initial screening
KW - Metadata
KW - Software engineering
UR - http://www.scopus.com/inward/record.url?scp=85162768520&partnerID=8YFLogxK
U2 - 10.1016/j.infsof.2023.107251
DO - 10.1016/j.infsof.2023.107251
M3 - Article
AN - SCOPUS:85162768520
SN - 0950-5849
VL - 161
JO - Information and Software Technology
JF - Information and Software Technology
M1 - 107251
ER -