Malva: a rapid, reference-free search engine for single-cell sequence data

1 min read
Source: Nature
Malva: a rapid, reference-free search engine for single-cell sequence data
Photo: Nature
TL;DR Summary

A new platform called Malva enables ultrafast, reference-free querying of raw single-cell and spatial transcriptomics data at atlas scale. It builds a scalable index of non-overlapping k-mers linked to cell barcodes (Malva Index), allowing exact, protein- and genome-agnostic searches for any sequence, mutations, splice junctions, circRNAs, pathogens, or nonreference transcripts. Malva supports sequence-based cell-type clustering, de novo marker sequence assembly, and integration with neural networks via an API and natural-language query translator. TheIndex aggregates tens of millions of cells from public datasets (e.g., Human Cell Atlas) in a compressed, on-disk structure, delivering real-time results (milliseconds to minutes) and enabling analyses beyond traditional gene-centric pipelines. Demonstrations include germline SNP allele frequencies, isoform and 3′ UTR usage, circRNA detection, viral/contaminant screening, and pan-cancer somatic mutation detection, all without read alignment. Limitations include a minimum 24-nt query, reliance on exact k-mer matching, and pseudocounts rather than absolute molecule counts. Malva provides a new discovery layer for sequence-defined questions and AI-assisted biology across huge public atlases.

Share this article

Reading Insights

Total Reads

1

Unique Readers

17

Time Saved

82 min

vs 84 min read

Condensed

99%

16,622162 words

Want the full story? Read the original article

Read on Nature