The database SILVA provides comprehensive, quality checked and regularly updated datasets of aligned small (16S/18S, SSU) and large subunit (23S/28S, LSU) ribosomal RNA (rRNA) sequences for all three domains of life (Bacteria, Archaea and Eukarya). It is part of the DSMZ Digital Diversity and listed as Global Core Biodata Resource and ELIXIR Core Data Resource.
With the current release, SILVA enhances its functionality:
- Release of the SILVA SSU 144 dataset focused on prokaryotic small-subunit (SSU) rRNA sequences, following a complete redesign of the curation workflow for improved robustness and sustainability.
- Introduction of a new curation workflow developed by a new team of curators, ensuring a more reliable and scalable process for future releases.
- Inclusion of all newly available eukaryotic sequences in the dataset, although they have not undergone manual curation and are currently classified based on the existing SILVA guide tree.
- New versioning scheme: SILVA 144 is the first release using a new versioning system, with version numbers increasing by one for each full release (e.g., 145, 146), independent of ENA release numbers.
- Integration of multiple classifiers: The release includes pre-trained classifiers for DADA2, Kraken2, MEGAN, and Qiime2.
No large-subunit (LSU) dataset released with SILVA 144, as efforts were prioritized on finalizing the new curation workflow and ensuring the timely release of the prokaryotic SSU dataset.
Future roadmap
- Habitat-specific Qiime2 classifiers are still under development and will be released when ready.
- Enhanced metadata availability planned for future releases, including additional data such as BioSamples.
- The eukaryotic dataset will be curated in subgroups (e.g., Fungi, Metazoa) and released incrementally starting in 2027, with the goal of a full, high-quality release of all domains (Archaea, Bacteria, Eukaryota) in SILVA 145.
- Ongoing software improvements: The development team will focus on enhancing pipeline scalability and modernizing outdated software components to handle current data volumes.
