Skip to content

Drug Discovery

Dereplication at screening throughput.

Hundreds of extracts a week, with library generation, interactive molecular networking and database search in the same environment your chemists already use.

The bottleneck

Rediscovering a known compound is the expensive failure

In natural product discovery the cost is not in the screen, it is in chasing a hit that turns out to be a known metabolite. Dereplication has to happen at the speed of acquisition, not weeks later.

  • Known compounds re-isolated because dereplication lagged the screen
  • In-house libraries trapped in spreadsheets and personal folders
  • Networking run in one tool, annotation in another, quantitation in a third

The workflow

From raw files to something you can publish.

Every step ships with default settings. You may customise each one, but you don’t have to.

  • Import

    Hundreds of extracts as one batch, from any major vendor format.

  • Detect and resolve

    Feature detection and deconvolution across the whole screen at once.

  • Group the ions

    Adducts and in-source fragments collapsed so a cluster counts compounds, not signals.

  • Network

    MS2 similarity explored interactively, without exporting to a separate tool.

  • Dereplicate

    Public and in-house libraries queried in one pass to find the knowns early.

  • Build the library

    Confirmed identifications versioned into a reference set the next screen starts from.

Capabilities

What you get out of the box.

Adduct series annotated across a spectrum, linked into a molecular network graph

Library generation

Build and version an in-house spectral library from confirmed identifications.

Interactive molecular networking

Explore the network without exporting to a separate tool.

Database search

Public and proprietary libraries queried in one pass.

Screening-scale batches

Hundreds of complex extracts processed as a single reproducible job.

In the software

The molecular network.

Twenty-four rows left, out of 9,526. The pie on each node shows which sample groups the feature turned up in, so anything that is only medium background is gone before a single spectrum is opened. Ion identity edges collapse the adducts of one molecule onto one node first, so a cluster counts compounds rather than the same compound three times over, and each annotation names the library entry behind it, down to the accession and the collision energy it was measured at.

Drug Discovery
The mzmine molecular networking view: feature and neutral-molecule nodes, each drawn as a pie of the sample groups it appears in, joined by MS2 cosine, ion identity and MS1 shape correlation edges, with named annotations including paracetamol, tryptophan and tyrosine; beside it a mirror plot of the measured spectrum against the MassBank library entry with the compound and instrument details listed, and below it the aligned feature table filtered to 24 rows

In production

Labs already running this.

“As part of the R&D scent team at IFF, working with complex and diverse natural product extracts, efficient data processing, deconvolution, annotation, chemometrics, and interactive visualization are crucial for high-throughput understanding. mzmine Pro has provided us with not only a powerful and efficient GC-MS workflow to investigate volatile compounds but also a highly advanced LC-MS workflow, perfectly aligned with our aspirations. These tools are essential for enhancing our work, fostering innovation, and deepening our understanding of complex natural matrices.”
Dr. Melissa Nothias-Esposito
Dr. Melissa Nothias-Esposito
Senior Scientist · LMR by IFF, Grasse, France
“As part of our drug discovery workflow, we analyze hundreds of complex fungal samples every week. mzmine PRO provides the necessary processing power and innovative compound discovery tools, such as library generation, interactive molecular networking, and database search to swiftly accelerate our daily routine. Data processing is no longer a bottleneck.”
Dr. Annika Jagels
Dr. Annika Jagels
Senior Scientist · LifeMine Therapeutics, Cambridge, USA

The science

The methods behind this workflow, peer-reviewed.

2023
Integrative analysis of multimodal mass spectrometry data in MZmine 3 (opens in a new tab)
Schmid et al. · Nature Biotechnology 41, 447–449
The methods paper for the platform itself. Cite this one if you cite only one.
2026
Self-supervised learning of molecular representations from millions of tandem mass spectra using DreaMS (opens in a new tab)
Bushuiev et al. · Nature Biotechnology 44, 630–640
A foundation model for MS2 spectra, trained on millions of unannotated spectra.
2021
Ion identity molecular networking for mass spectrometry-based metabolomics in the GNPS environment (opens in a new tab)
Schmid et al. · Nature Communications 12, 3832
Resolves adducts and in-source fragments into single molecular identities.
2020
Feature-based molecular networking in the GNPS analysis environment (opens in a new tab)
Nothias, Petras, Schmid et al. · Nature Methods 17, 905–908
The networking method the drug discovery and dereplication workflows rest on.

Resources

Go deeper on Drug Discovery

We have written this up in more detail. There is one paper. One short form and all of it unlocks.

Get started

See it on your own data.

Send us a few representative files. We will build the drug discovery workflow and show you the result before you commit to anything.