CRISPR-CAS9 OFF-TARGET PREDICTION USING DEEP LEARNING AND ITS BIOINFORMATICS TOOLS: A SYSTEMATIC REVIEW

Authors

  • Akhtar Sayed
  • Kifayat Ullah
  • Naeem Uddin

Keywords:

CRISPR-Cas9; off-target prediction; deep learning; bioinformatics; machine learning; gene-editing safety

Abstract

CRISPR-Cas9 has become the dominant platform for programmable genome editing, but its clinical, agricultural, and research applications are constrained by unintended cleavage at genomic loci that are imperfectly complementary to the guide RNA (gRNA). Accurate prediction of these off-target sites is essential for experimental design, regulatory review, and the safety of gene-editing therapeutics. Since approximately 2018–2020, deep learning (DL) architectures (convolutional neural networks (CNNs), recurrent networks (RNNs/LSTMs/GRUs), attention mechanisms and transformers, and hybrid or ensemble models) have progressively displaced earlier alignment-based and rule-based scoring systems such as Cas-OFFinder, the MIT score, and the Cutting Frequency Determination (CFD) score, as well as classical machine-learning predictors such as CRISTA and Elevation. This systematic review synthesizes the literature on DL-based off-target prediction tools for CRISPR-Cas9 published between 1 January 2020 and the final search date in 2026. Searches were conducted across PubMed, Google Scholar, IEEE Xplore, and bioRxiv, supplemented by citation tracking. Seventeen studies describing distinct deep-learning-based tools or frameworks were included, spanning CNN-only architectures (e.g., CnnCrispr, DL-CRISPR), recurrent/hybrid CNN–RNN architectures (e.g., CRISPR-Net, R-CRISPR, Crispr-SGRU, CRISPR-DIPOFF), attention- and transformer-based models (e.g., CRISPRIP, CRISPR-M, CRISPR-HW, transformer anti-noise models), and emerging language-model, molecular-simulation, hybrid-fusion, and transfer-learning approaches (e.g., CRISOT, CCLMoff, CRISPR-MCA, similarity-based transfer learning, crispAI, and the recently published EHNN). Across the included studies, DL models consistently outperformed classical scoring methods on held-out or leave-one-guide-out validation, but the field is characterized by heterogeneous and overlapping benchmark datasets (chiefly derived from GUIDE-seq, CIRCLE-seq, CHANGE-seq, and the DeepCRISPR HEK293T/K562 corpora), inconsistent negative-sampling strategies, near-total reliance on retrospective rather than prospective validation, and limited interpretability of "black-box" predictions. We identify a marked shift toward attention-based, multi-view, and pretrained-language-model architectures in 2023–2026, alongside a growing but still nascent emphasis on model interpretability (e.g., integrated gradients, SHAP, attention-map visualization) and, most recently, an emphasis on prediction calibration (well-behaved probability outputs, not only discrimination) exemplified by the 2026 EHNN model. No formal meta-analysis was performed given the heterogeneity of reported metrics and test sets; findings are synthesized narratively. To our knowledge, this is the first review focused specifically on deep-learning-based CRISPR-Cas9 off-target prediction tools, and it highlights an urgent need for standardized, independent benchmark suites and interpretable, clinically validated platforms. No external funding was received for this review, and the protocol was not prospectively registered.

Downloads

Published

2026-03-31

How to Cite

Akhtar Sayed, Kifayat Ullah, & Naeem Uddin. (2026). CRISPR-CAS9 OFF-TARGET PREDICTION USING DEEP LEARNING AND ITS BIOINFORMATICS TOOLS: A SYSTEMATIC REVIEW. Spectrum of Engineering Sciences, 4(3), 3252–3272. Retrieved from https://www.thesesjournal.com.medicalsciencereview.com/index.php/1/article/view/3532