Clustered Regularly Interspersed Short Palindromic Repeats and CRISPR-associated genes (CRISPR-Cas) is a bacterial immune system also famous for its use in genome editing. The diversity of known systems could be significantly increased by metagenomic data. Here we present the Metagenomic CRISPR Array Analysis Tool (MCAAT), a highly sensitive algorithm for finding CRISPR arrays in unassembled metagenomic data. It takes advantage of the properties of CRISPR arrays that form multicycles in de Bruijn graphs. We show that MCAAT reliably predicts CRISPR arrays in bacterial genome sequences and that its assembly-free graph-based strategy outperforms assembly-based workflows and other assembly-free methods on synthetic and real metagenomes. Our new approach will help to increase the diversity of known CRISPR-Cas systems and enable studies of spacer evolution within metagenomic data sets.
Keywords: CRISPR array; CRISPR-Cas; bioinformatics; cycles; de Bruijn graph; metagenomic.
© The Author(s) 2025. Published by Oxford University Press on behalf of FEMS.