Documentation
What DreaMS Search does, how to run a search, and how to read what comes back.
Overview
DreaMS Search compares tandem mass spectra you upload against large reference libraries of public MS/MS data, and ranks what looks most like them. The DreaMS model measures the likeness from the whole spectrum, so it surfaces related compounds as well as exact repeats of your measurement.
The name covers three things. This page documents the second, and the two surfaces beside it.
| Name | What it is | Use it when |
|---|---|---|
| DreaMS | The model and its Python package (source, paper). | You want embeddings in your own code, or to train on your own data. |
| DreaMS Search | This application. It runs the model over libraries that are already indexed. | You have a file and want matches. It works in the browser. |
| Spectrum viewer | A surface of this application that draws and scores spectra straight from your files. | You want to look at two spectra and compare them. |
| DreaMS Atlas | A map of public spectra and how they relate, built once from the reference data. | You are exploring what is out there, with a file or without one. |
DreaMS Search
Give the search page a file of tandem mass spectra and choose which reference libraries to search. It embeds every spectrum with the DreaMS model, ranks each library against them, and hands back a table per library plus a CSV of everything.
Running a search
1. Give it a file
Drop one file of tandem mass spectra, or click to browse. Your browser reads it first, so the page tells you how many spectra it holds before anything is sent.


Short of a file, pick an example. They run as an upload does, against the same libraries.


2. Choose what to search against
Pick at least one reference library. Each one is ranked on its own and gets its own table, so you can always see which library a hit came from. Pick several and you get several tables.


Loading the reference libraries…
Your own library
Attach a reference library of your own alongside the file you are searching. It is parsed and embedded for that one search, held in memory while the search runs, and matched exactly.


Attach one and it stands as a valid search on its own, with or without a deployed library beside it.
3. Run it
The search runs on the server. Progress arrives live, and the page names the stage each chunk of your file is in.


Submitting starts a queue wait and a parse, which the server reports once the first chunk is under way and the counters begin to move.


What your file needs
Formats
One file per search. The extension decides whether the upload is accepted, so name a file for what it holds.
The upload form lists the file formats this deployment accepts, and rejects anything else before the search starts.
Centroided MS/MS
Each peak should be one m/z and one intensity. Everything here reads a spectrum that way, so centroided input is what the scores are built for. A profile-mode file is still searched, and the results report how many spectra declared it, which tells you when the representation explains a disappointing ranking.
Positive ion mode
The model was trained on positive-mode spectra, so it is at its strongest there. Negative-mode spectra are accepted and searched, and the results state the count and mark the rows, so you can weigh those rankings accordingly. Ion mode comes from whatever your file declares, and stays unknown when the file is silent.
Current limits
What this deployment is running with. The upload form enforces them and names the one you hit.
Loading the current limits…
Reading results
Each library gets its own table, ranked by DreaMS similarity. Compare a score with others from the same library, where the size of the collection and the way it was built are held constant.
Pick a query spectrum
You read results one query spectrum at a time. The panel beside the table lists the spectra that matched, with each one's precursor mass and best score, so you can see which are worth opening.


The hit table
One row per matched reference spectrum. A column appears once some row fills it, so richly annotated hits bring more columns with them.


The five scores
Every hit carries five. They answer different questions, and a hit scoring well on all of them is far stronger than one carried by a single column.
| Score | What it answers | What it needs | When it is filled |
|---|---|---|---|
| DreaMS similarity | Does the model think these are related compounds? | The two spectra alone | Always filled, for a deployed library |
| Cosine similarity | Do the two agree on actual fragment masses? | Reference peaks | Filled where the library stores peaks |
| Modified cosine | Do they agree once a precursor-mass shift is allowed? | Reference peaks | Filled where the library stores peaks |
| Neutral loss | Do they lose the same things from their precursors? | Reference peaks and a precursor on both | Filled where the library stores peaks and both precursors are known |
| Entropy similarity | Do they agree on the peaks that carry the most information? | Reference peaks | Filled where the library stores peaks |
Cosine, modified cosine and neutral loss each publish the number of peak pairs they matched. That count is what tells you how much evidence a score rests on: read a high score beside a pair count of two with that in mind.
The mirror plot
Every row draws your spectrum against the one it matched: yours up, the reference down. Peaks that line up are fragments the two agree on, and a peak standing alone belongs to that spectrum only.


Below is the same plot at reading size, same code, on two standards this app ships. Caffeine and theophylline differ by one methyl group, so their precursors sit 14.0157 apart and several fragments are shared. That is what modified cosine is for: it holds up across a precursor shift where plain cosine drops.
Narrow and export
The toolbar filters rows, chooses columns, and downloads the CSV. Filters narrow the hits you already have, so a library that carries the column you filtered on answers and the rest stay quiet.


Limits and caveats
Quality filtering
The form offers to drop your low-quality spectra before they reach the model, by the criterion the DreaMS authors defined. It acts on your own spectra, deciding which of them are searched, and leaves the libraries and the results as they are.


It starts off. A spectrum unlike the model's training data is worth reading carefully, which keeps the measurement in play. The results say how many were dropped, and a file whose every spectrum is dropped fails outright, so an empty result always has a reason attached.
Input the model was trained for
| Input | What happens | How you are told |
|---|---|---|
| Negative mode | Searched. The model is positive-mode-trained, so read these rankings with more caution. | A count above the table, a pill on each affected spectrum. |
| Profile mode | Searched. Centroided peaks are what the scores are built for. | A count of the spectra declaring it. |
| Unknown mode | Searched, and recorded as unknown, which is what a silent file leaves it as. | The spectrum carries no mode. |
One library carries negative-mode references, and those hits are labelled on the row.
Fixed result depth
Every search returns the same number of hits per query spectrum per library, and the CSV carries all of them. The depth is a property of the deployment, listed with the other limits above.
Approximate matching
Searching tens of millions of spectra uses an approximate index, so a ranking comes very close to the exact answer at a fraction of the cost. A library you upload yourself is small enough to match exactly.
Results expire
A finished search and its download are kept for a bounded time, listed with the limits above. Recent searches lives in your browser and outlasts that, so it can list a search whose result the server has already swept.
Spectrum viewer
Spectrum viewer opens one or two files and draws their spectra straight away. Your browser does the parsing, so the file stays on your machine and a spectrum appears as soon as you drop it in.


Comparing two spectra
Tick one scan, then tick a second. The viewer draws them head to head and scores that pair on the same five columns the results table carries, so you get the numbers straight from your own files.


The two scans can come from one file or from two, which makes this the quickest way to check a replicate against a standard.
DreaMS Atlas
The Atlas is a map of public spectra and how they relate to each other, built once from the reference data. You browse it, and it answers about the reference collection itself.


Browsing the graph
A spectrum is drawn with its nearest neighbours, so you read it in the company of what it resembles. Follow an edge and the neighbourhood moves with you, which is how a class of related compounds becomes visible as a shape.
Reading the graph
- Amber: sample metadata is recorded
- At least one of its spectra has a sample record saying what it was measured in, from ReDU or a MASST table.
- Grey: no sample metadata
- No sample record is known for any of its spectra.
- A ring: a known structure
- From the annotated library, so it can be drawn. A ring also marks a structure that is withheld for licensing.
- The number: how many spectra it stands for
- The size of the cluster the node represents, written in it. A node with no cluster recorded carries no number rather than a zero.
- Bigger: the spectrum the panel is describing
- Size marks the one the detail card is about, and nothing else changes size, so it never competes with what the node is.
- A soft band: where the walk is rooted
- The spectrum this neighbourhood was drawn around. Drawn outside the circle, so it never covers what the node is.
- Thicker, darker, shorter: more similar
- Three channels carry it at once, so none of them has to be seen alone. Relative to what is on screen, not an absolute score: every link in this network is a near neighbour, so the widths compare these links to each other and nothing else.
Further reading
- DreaMS model documentationThe model, its Python package, and tutorials on computing embeddings in your own code.
- Bushuiev et al., Nature Biotechnology 2025How the model was trained and evaluated, and the quality criterion the form offers.
- DreaMS source repositoryModel code and released checkpoints.
Write to dreams@fi.muni.cz about anything here: a question this page does not answer, a result that looks wrong, or a reference library you would like served.