A SCRAPY–SPLASH PIPELINE FOR AUTOMATED COLLECTION AND RANKING OF SCIENTIFIC LITERATURE FROM JAVASCRIPT-RENDERED SEARCH ENGINES

Authors

  • Asfandyar Ahmed
  • Obaid Ur Rehman
  • Muhammad Touseef Irshad

Keywords:

Web scraping; Scrapy; Splash; Google Scholar; Locality Sensitive Hashing; Cosine similarity; Text mining; Literature review automation.

Abstract

More than 1.8 million scientific articles are published each year across roughly 28,000 peer-reviewed journals, and searching academic engines such as Google Scholar page by page while copying metadata by hand has become a slow, error-prone bottleneck in literature reviews. This study presents an integrated, reproducible pipeline that automates this process for JavaScript-rendered search engines. Scrapy performs asynchronous crawling, Splash renders dynamic result pages before XPath-based extraction, and automated pagination follows the results to the last available page. A request-management layer combines a randomised download delay of 2.5–7.5 seconds, Scrapy’s AutoThrottle extension, user-agent rotation and proxy-based IP rotation. Extracted records are cleaned in Scrapy’s item pipeline, stored in CSV and MySQL, and processed in RapidMiner through tokenisation, stop-word removal, case normalisation and n-gram generation. Each record is then ranked by its cosine similarity to the research query, with Locality Sensitive Hashing (LSH) serving as the scaling layer for larger collections. Evaluated with the query “web scraping”, the pipeline collected 200 metadata records in a single session across all available result pages. HTTP 429 (Too Many Requests) responses, frequent in the baseline configuration, were eliminated once request management was enabled, and HTTP 504 timeouts fell from about 160 per session to zero after the resources of the Splash container were increased. In manual inspection, the highest-ranked records were on-topic for the query. The study also documents these failure modes and their resolution, and it examines the legal and ethical limits of scraping academic search engines. The pipeline substantially reduces the manual effort of small-scale academic literature collection, and for large-scale use it can be adapted to open scholarly APIs such as OpenAlex and Semantic Scholar.

Downloads

Published

2026-03-23

How to Cite

Asfandyar Ahmed, Obaid Ur Rehman, & Muhammad Touseef Irshad. (2026). A SCRAPY–SPLASH PIPELINE FOR AUTOMATED COLLECTION AND RANKING OF SCIENTIFIC LITERATURE FROM JAVASCRIPT-RENDERED SEARCH ENGINES. Spectrum of Engineering Sciences, 4(3), 6510–6522. Retrieved from https://www.thesesjournal.com.medicalsciencereview.com/index.php/1/article/view/3916