Showing posts with label PEAKS Studio. Show all posts
Showing posts with label PEAKS Studio. Show all posts

Monday, December 2, 2013

Boost your analysis speed with PEAKS 7

New mass spectrometry instruments come out every year with higher resolution and/or scan rate. This causes an increasing amount of data generated per hour. While the accuracy and sensitivity of the software tools are critical, many researchers have the need to complete the analysis sooner without compromising the accuracy or sensitivity.

In PEAKS 7, we have re-engineered many parts of the algorithms to make them more efficient.
"The increase in speed [of PEAKS 7] is quite remarkable."
Paul Taylor, Sick Kids Hospital
In this post, we will use an internal test to illustrate the dramatic speed improvements our users love.

Test data and search parameters

12 files generated from Thermo Q-Exactive instrument was used. Both MS and MS/MS were acquired on high resolution. In total, there are more than 100,000 MS/MS spectrum.
Test hardware and operating system

Mac mini late 2012 version. Intel i7-3720QM processor at 2.6GHz. 16GB RAM. Windows 7 Professional Edition 64 bit. The Windows OS is installed using Boot Camp.

The test and the result

PEAKS 6 and PEAKS 7 were installed on the same Mac mini computer. The computer was rebooted between analysis.

First off, we want to test our much loved de novo sequencing speed improvement and here is the result. PEAKS 7 de novo is more than five time faster!

The complete analysis in PEAKS can identify database peptides, novel peptides, PTMs and mutated peptides in one go. We use the total running time in version 6 as the baseline and see how much shorter researchers can finish the same analysis.

The same search completed in just one quarter of what it takes in the previous version, with the same level of accuracy and sensitivity.




Friday, August 9, 2013

de novo only peptides in PEAKS

One of the unique features in PEAKS is that it provides a list of de novo only peptides. What does it mean?

de novo only peptides are the de novo sequences derived from the spectra that do not a have confident database match.

In PEAKS, de novo only peptides are listed in the de novo only tab in every PEAKS DB, PEAKS PTM, SPIDER and inChorus results. The actual de novo only peptides displayed in the list will be affected by the filters in the summary tab. Let's use PEAKS DB result as an example.

The de novo only peptides are defined in the second row of the filters. The first part, TLC and ALC filters are the same as it is in the de novo result. It tells PEAKS what should be considered as a good de novo sequence. The second part, peptide -10lgP filter tells PEAKS what should not be considered as a confident database match. After the filters are applied, PEAKS will go through all the de novo sequences that passed the TLC and ALC filters. For each of such de novo sequence, PEAKS will look at the corresponding spectrum. If the spectrum does not produce any database matches that have a score higher than the de novo only peptide -10lgP filter, the de novo sequence of the spectrum will be considered as a de novo only peptide.

PEAKS does not stop at just providing a list of de novo only peptides, it also tries to associate them with the proteins. In the protein coverage view below, the gray bars represents the de novo only peptides that share at least 10 consecutive AAs with the protein sequence. This is particularly useful when looking for unexpected PTMs or glycosylation site.

Monday, June 24, 2013

Carbamidomethyl @ C, D, E, H, K, N-term

Recently I worked on an ETD dataset generated from Orbitrap Velos. The user mentioned that Carbamidomethyl on Cysteine and some phosphopeptides are expected but he was not able to get any good identification results using Mascot.

My first attempt is to use the information provided by the user and run the PEAKS DB search. The result shocked me as under 1% FDR, PEAKS DB only reported 170 PSMs. What could possibly go wrong?

Looking at the PEAKS DB result, there are many "de novo only" peptides, which means many spectrum can produce confident de novo sequences but they do not have a confident database hit. This could be a result of unsuspected PTMs and mutations.

So I decided to do a PEAKS PTM search on the db result as my second attempt. PEAKS PTM reported 582 PSMs under 1% FDR. In the summary view, PTM profile section, there are many PSMs found with the PTM Carbamidomethyl on D, E, H, K and N-term. This could be an indication of excess iodoacetamide during alkylation procedure.

Now I know more PTM information of this data. In the last attempt, I added Carbamidomethyl @ DEHK,N-term and Dehydration @ DST into the variable PTM list along with Phosphorylation @ STY. PEAKS DB search was then performed with PEAKS PTM option enabled. This time, under 1% FDR, 736 PSMs were reported, more than four times the number of the initial search.

PEAKS PTM is a great tool to find unsuspected PTMs thus help explaining more spectrum.

Monday, June 3, 2013

Stop Guessing, Start Knowing - Score threshold and FDR control in PEAKS

In an MCP guideline published in 2004, "a significant but undefined number of proteins being reported as identified in proteomics articles are likely to be false positives".

To tackle this problem and control the "quality" of the identification results, over the past decade, false discovery rate (FDR) becomes the most accepted result validation method.

Although many search engines have the FDR estimation integrated into their package, in order to get the result under a particular FDR, e.g. 1%, the researchers usually have to do several round of "trial and error" on the score threshold used for filtering the result. Even worse, this process may have to be performed on every results.

PEAKS introduced a score selection tool to make this process very simple. Normally, it only takes three mouse clicks. Here is how to do it. On the top of the result summary pane, there is a FDR button.

Simply click that button will give you the score selection dialog. In this dialog, you can quickly select the commonly used FDR values on the right by a single click of the corresponding button. Or you can hover the mouse cursor on the chart to find the desired FDR value. When the value is found, right-click and then select "copy score threshold", the score to be used for the desired FDR will be filled into the score filter automatically. Click the "Apply Filters" button, you will get the result under the desired FDR.




Wednesday, May 22, 2013

More than a decade of PEAKS

Staring at the calendar, I can not believe it is almost half way into the year of 2013. The PEAKS team is hard at work and strive to better serve the mass spec proteomics community by adding tons of new features to every releases.

Ever since the first release, PEAKS has gone through many iterations and/or dramatic changes to become what it is today.


What new features will be in this year's release? Stay tuned!

Monday, May 13, 2013

Multiple enzymes support in PEAKS - Full Protein Coverage

PEAKS 6 introduced a new feature specifically targeting the experiments that use multiple enzyme digestions to increase protein coverage.

In the past, users have to search each sample separately and combine all the results manually afterwards or using none enzyme option to analyze all samples in one go which may cause higher false positives. Now in PEAKS, users can specify enzyme for each sample when creating a multi-sample project. Then in de novo sequencing and PEAKS DB, user can choose 'sample enzyme' in the enzyme list as the search option. PEAKS will use the correct enzyme when analyzing each sample.

From our users feedback, this feature is extremely useful when you want to fully characterize a single protein.The following example shows how big a difference this feature may make.

ALBU_BOVIN protein ordered from a reputable vendor was digested with Trypsin, LysC, GluC. The dataset is generated from Thermo Orbitrap instrument. Three searches were performed. The first one uses inChorus function to launch Mascot search (version 1.4) on the trypsin sample only. The second search uses standard PEAKS DB search on the trypsin sample. The third search uses the complete analysis workflow, including PEAKS PTM and SPIDER, on all three samples and uses "sample enzyme" as the enzyme option. The results are all filtered to only keep the very confident PSMs at 0.1% FDR level.

Mascot and PEAKS DB are able to achieve 73% and 86% protein coverage using only the trypsin sample respectively. In the protein coverage view below, the blue bars are the PSMs that matched the protein sequence at that position.

PEAKS complete analysis on all three samples reported 96% coverage on the protein. The uncovered 4% is in the protein N-terminal region, which is most likely cleaved-off and not in the purchased sample1.
1specific binding site (Asp-Thr-His-Lys) for Cu(II) ions. T. Peters Jr., F.A. Blumenstock. J. Biol. Chem., 242 (1967), p. 1574

Monday, May 6, 2013

Configure FASTA database in PEAKS

Configuring FASTA databases in PEAKS is fairly easy especially if the FASTA file has the same header format as one of the public databases (e.g. NR, Swiss-Prot, IPI). It is just a matter of selecting the pre-defined format and the parsing rules will be automatically filled in.

There are also a large number of users use PEAKS to search on their in-house, customized FASTA databases. In this situation, the header format is very hard to predict and it varies case by case.

In PEAKS, the parsing rule is defined using regular expression. While regular expression is very powerful, it will take people quite a bit of time to master it. Since we got tons of searches to run every week, against FASTA files with so many different header formats, I created this lazy, generic parsing rule for internal use and in most cases, it worked good enough.

Accession. The regular expression tries to use everything before the first white space as the accession. If no white space were found within the first 30 characters, the first 30 characters will be used as accession.
>\([^\s|]{1,30}\)
Description. The whole line after ">" will be used as the description.
>\(.*\)



Wednesday, May 1, 2013

100% vs 50% CPU usage, twice as fast? Not really!

Some user observed that when performing a search, the CPU usage for PEAKS would only go up to 50%. Why PEAKS does not use 100% of the CPU?

The observation is for sure valid, but the CPU usage reported by Windows Task Manager is somewhat misleading. 100% CPU usage does not mean the program is running twice as fast as under 50% CPU usage. The reason for this is, in my opinion, due to the Hyper-Threading technology most Intel CPUs have enabled by default. While the technology can improve the performance and responsiveness of a computer in some situation, it does not help much for computation heavy application, like PEAKS.

I did a performance test on a desktop PC with the following specification, a quad-core CPU with lots of RAM.
Intel i7 3770 3.4GHz CPU (quad core with hyperthreading)
16GB RAM
System drive SSD, data drive 7200RPM HDD
Windows 8 Pro 64bit
The dataset contains about 9000 MS/MS spectrum from Thermo Orbitrap instrument. PEAKS de novo was manually configured to run on 1, 2, 4 and 8 threads configuration respectively. For each configuration, two searches were done, one with hyper-threading enabled, one with hyper-threading disabled. So in total, there were 8 runs, a clean project was created for each run and the PC was rebooted between the runs.
With HT enabled, there is only about 10% performance gain running 8 threads (100% CPU usage) than running 4 threads (50% CPU usage). The search indeed run slightly faster, but the computer is not very responsive for even the simplest tasks like email, excel, etc.




Friday, April 26, 2013

Common ptifalls of FDR estimation part three


The third pitfall is also caused due to the over-emphasis on sensitivity.

There is another trend in database search software to re-score the peptide identification results by using machine learning. The idea is straightforward: After the search, we know what the decoy hits are. The algorithm should take advantage of it, and retrain the parameters of the scoring function to get rid of the decoy hits. With this effort, it will get rid of a lot of the target false hits as well.

The method is valid, except that it may cause FDR underestimation. This is because the target false hits are unknown to the machine learning algorithm. Therefore, there is a risk that the machine learning algorithm removes more decoy hits than the target false hits.
This overfit risk is well known in machine learning. A machine learning expert can reduce the risk but can never get rid of it. 

The solution to this pitfall number 3 is trickier.

The first suggestion: don’t use it. The philosophy here is that judges cannot be players. If we want to use the decoy for result validate, the decoy information should never be released to the search algorithm.

If this re-scoring method must be used due to the low-performance of some database search software, it should only be used for very large dataset to reduce the risk of over-fit.

Perhaps the best solution is the third one. That is, the retraining of the score parameters should be done for each different instrument type, instead of each dataset. This will gain much of the benefit provided by machine learning, but without the problem of over-fitting. Indeed, this third approach is what we do in the PEAKS DB algorithm.


*The content of this post is extracted from "Practical Guide to Significantly Improve Peptide Identification Sensitivity and Accuracy" by Dr. Bin Ma, CTO of Bioinformatics Solutions Inc. You can find the link to the guide on this page.

Monday, April 22, 2013

Common ptifalls of FDR estimation part two


The second pitfall of the traditional target-decoy strategy is caused by another popular technique used to increase the peptide identification sensitivity.  

The idea is clever: if a weakly identified peptide happens to be on a highly-confident protein, then the peptide is likely to be correct regardless of its low score. So, to increase the sensitivity, the software can add a score bonus to each peptide on a multiple-hit protein. Indeed, this protein bonus will save some weak true hits, but it will save some weak false hits at the same time. The bigger problem is that the target database will provide more multiple-hit proteins than the decoy. As a result, more weak false hits will be saved from the target database. This will cause the FDR underestimation.
In PEAKS,
decoy fusion approach can solve this problem effectively.
 

Because the target and decoy sequences are concatenated into a single protein sequence, when a protein bonus is added to the multiple-hit proteins, the same bonus will be added to the target and decoy hits equally. So, weak false hits are saved with approximately equal probabilities in the target and decoy. This recreates the balance and provides accurate FDR estimation.

By using the decoy fusion as the validation method, we can safely apply the protein bonus. We get the sensitivity, but did not compromise  the FDR estimation.

 
*The content of this post is extracted from "Practical Guide to Significantly Improve Peptide Identification Sensitivity and Accuracy" by Dr. Bin Ma, CTO of Bioinformatics Solutions Inc. You can find the link to the guide on this page.

Friday, April 19, 2013

Common ptifalls of FDR estimation part one

Today’s most widely used method for FDR estimation is the target-decoy strategy. This is a well-established method in statistics and started to be used in proteomics around 2007.

In this approach, a decoy database that contains the same number of proteins as the target database are searched together by the database search engine to identify peptides. The blue colors indicate the target hits and the orange colors indicate the decoy hits, the squares are the false hits, and circles are true hits. 
The decoy proteins are randomly generated so that any decoy hit is supposedly a false hit. Since the search engine doesn’t know which sequences are from target and which are from decoy, when it makes a mistake, the mistake falls in the target and decoy databases with equal probability. Thus, the total number of false target hits can be approximated by the number of decoy hits in the final result. And the FDR can be estimated by the ratio between the numbers of decoy hits and the number of target hits. 

The target-decoy strategy is a powerful method for FDR estimation. However, as we will discover in the next little while, such a powerful method must be used with caution to avoid FDR underestimation. 

The first pitfall in the use of target-decoy approach for FDR estimation is due to the so-called multiple round search strategy in today’s database search software. 

This multi-round search was popularized by the X!Tandem program published in 2004, in order to speed up the computation. The first round uses a fast but less sensitive search method to quickly identify a shortlist of proteins from the large database. Then, the second round uses a more sensitive but slower search method to identify peptides, but only from the short list of proteins. This effectively speeds up the search without sacrificing too much sensitivity. Indeed, X!Tandem is one of the fastest search algorithm used today.

However, as pointed out by a paper published in JPR in 2010, this multiple-round search strategy screws up the target-decoy estimation of the FDR. The reason is that after the first round, there will be more target proteins than the decoy in the short list. Thus, if the second round search makes a mistake, the mistake will be more likely in the target proteins. So, we will end up with fewer decoy hits than the actual false target hits. This causes the FDR underestimation.
The JPR paper in 2010 provided a fix to this problem. But a year later, in another JPR paper, Bern and
Kil pointed out that the fix was wrong, and proposed a different fix that required the change of the search engine’s algorithm. This shows that the FDR estimation is very tricky, even the experts can sometimes get it wrong. 


In PEAKS, we used a new approach, called decoy fusion to solve this problem. 

Instead of mixing the target and decoy databases, we append a decoy sequence to each target protein.

So, after the fast search round, the protein shortlist will still contain the same length of target and decoy sequences. And the false hits of the second round will have the equal chance to be from the target and decoy sequences. This recreates the balance and can accurately estimate the FDR in the multiple-round search setting.


*The content of this post is extracted from "Practical Guide to Significantly Improve Peptide Identification Sensitivity and Accuracy" by Dr. Bin Ma, CTO of Bioinformatics Solutions Inc. You can find the link to the guide on this page.

Monday, April 15, 2013

Use new / non-standard amino acids in PEAKS

Some researchers are looking for unknown amino acids in their proteomic studies. This requires the software to be able to take new theoretical AA into account. Although, PEAKS internally only use the 20 common amino acids residues, here is a trick we can use to add new AAs.

The general concept is to add this new amino acid X as a PTM of an existing amino acid Y. The existing amino acid Y should be picked so that the mass difference between X and Y is distinguishable from the mass of all PTMs in the search. In the analysis, on top of the normal PTMs you may set, set this special PTM as a variable modification. In the result, this new amino acid X will appear as a variable modification of Y.

Here is an example. Let's add the amino acid Pyrrolysine to PEAKS. Open the PTM configuration by click the menu "Window"-->"Configuration", then select the PTM tab. Click the "New" button on the lower right corner of the dialog. I am going to configure it as a modification of Alanine. The mass difference is 184.1212Da.
Click "OK" to complete the configuration. When performing a PEAKS search, the new PTM, Pyrrolysine, can then be selected in the "Customized" tab of the "PTM Options" dialog.
Use the same trick, we can substitute or remove an amino acid by setting it as a fixed modification in the search.

Wednesday, April 10, 2013

Why everyone should de novo?

When I was doing some house cleaning on old documents, I found a few slides created 5 years ago about why de novo sequencing should be included in all research workflows. Most of these things still stay true today and it aligns very well with PEAKS development. I'd like to share them here.

Problem:
“The organism I’m studying isn’t in any of the databases.”

Solution:
De novo sequencing is the only way to go.

Problem:
“I need more confidence in the peptides found by MS/MS ion search”

Solution:
Use an orthogonal approach like sequence tag searching or hybrid searching to validate the results

Problem:
“I can only explain <10% of my data. I need to boost my search engine’s performance.”

Solution:
Use two or more unique search engines. Then try de novo sequencing.

Problem:
“I think I’m still missing some peptides, maybe because of with PTM”

Solution:
Use de novo to find a partial sequence; matching this to a database will highlight peptides of interest.

Problem:
“I am still missing some peptides even after trying all PTM. But the data is great!”

Solution:
De novo + peptide homology search like SPIDER or BLAST.

Problem:
“My lab creates too much data for de novo sequencing”

Solution:
Use PEAKS. It’s fast enough.

Tuesday, April 2, 2013

A few quick issues of running PEAKS on OS X and Linux

I have played PEAKS as a viewer on a Mac OS X for a while and do notice some minor issues. That's expected as PEAKS on OS X is not officially supported!

No instrument raw file loading ability. This could probably never be fixed unless the instrument vendors port their libraries to OS X. So in short, PEAKS can only read text formats (mgf, mzxml, mzml, pkl, dta, etc) and of course PEAKS projects on OS X.

Images on the summary view is broken. This seems to be a coding issue related to the path (file path on Windows and OS X are different). I do manage to get a workaround though. On the summary view, click "Notes" button on the top of the view, a text editor will show up. Type in the following and click "OK".
<a href="">go</a>
A "go" link will be displayed in the "Notes" section of the summary view. Click on the link, the summary view will be correctly displayed in your default web browser.

Vertical tabs are too small and the text on them are not visible. Well, not sure how to workaround this. But when you mouse over the tabs, the tooltips do work.

I will keep playing PEAKS on OS X when I have chance. Please comment if you find other interesting issues :D

Wednesday, March 27, 2013

PEAKS de novo sequencing on 3.4 million MS2 spectrum

It is unbelievable how fast the mass spectrometry data can grow in size. A decade ago, when Dr. Bin Ma created PEAKS, a data file containing hundreds of spectrum is already considered large. Now we are seeing dataset over a million spectrum on a regular basis. On this scale, it is quite challenging for proteomics software to analyze the data within a reasonable amount of time under finite computing resources.

I recently received a large dataset from one of our collaborators. The data contains 3.4 million MS2 spectrum in total (about 160 hours LC for all the samples) and it is generated on a Thermo Orbitrap instrument.

I now have the project created and started de novo sequencing on all the spectrum in one go. It will take a while until the job completion.

Stay tuned!

Tuesday, March 26, 2013

How to run PEAKS Studio/Viewer on Mac OS X or Linux?

Occasionally, we are asked by users whether PEAKS can run on Mac OS X or Linux. PEAKS is written in Java, theoretically it should work. This post will show the steps you need to do to make PEAKS run a Mac OS X. The steps needed should be very similar to make PEAKS work on Linux.

Disclaimer: PEAKS does not officially support any OS other than Windows as of the time I am writing the post. The software may not be fully functional. Activating the software on OS X or Linux will consume the license, which means the same license can not be used again. I strongly recommend only following the steps to configure PEAKS Viewer (the unlicensed Studio) on OS X or Linux for PEAKS result sharing and presentation purposes.

Before you start

PEAKS is a Java program. So before porting PEAKS to Mac, we will make sure that Java is installed. Open a terminal window and type in the command:
java -version
If Java is installed, the version information will be displayed. In OS X Mountain Lion, if Java is not installed, this command will also trigger a window for Java installation.

Get the files

Since PEAKS only have the installer for Windows, you will need a Windows computer to install PEAKS and copy the installed files over to Mac.

To proceed, download PEAKS from the website on a Windows PC. Run the installer, follow the on screen instructions to complete the installation. By default, PEAKS 6 will be installed on C:\PeaksStudio6 directory. Copy the directory to a USB drive and copy it to Mac OS X, e.g. /Users/userx/PeaksStudio6.

Configure PEAKS on OS X

Open a Terminal window and change the directory to the PEAKS directory, for example, /Users/userx/PeaksStudio6, by typing the command:
cd /Users/userx/PeaksStudio6
We need to replace the Windows version JRE with the one installed in OS X:
rm -r -f jre
mkdir jre
cd jre
mkdir bin
cd bin
ln -s /usr/bin/java java.exe
cd /Users/userx/PeaksStudio6
We want to make sure PEAKS starts in one JVM (type the following on one line with a white space after ".jar"):
jre/bin/java.exe -cp peaksstudio.jar com.bsi.tools.computenodenumber.PerformanceConfigDialog
In the performance configuration dialog, select "Manually configure PEAKS performance" option and make sure that the "Start Client Separately" and "Start Compute Node Separately" checkboxes are unchecked. Click "Apply" and close the dialog.

Use a text editor, e.g. vim, to create the start up script peaks.sh:
#!/bin/sh
jre/bin/java.exe -Xmx12000m -splash:splash.png -jar peaksstudio.jar
The number 12000 means that PEAKS can use up to 12GB of RAM. You can change this value based on your computer configuration, but a higher amount is always preferred.

We need to make the script executable:
chmod u+x peaks.sh
Now you can start PEAKS by simply run the script:
./peaks.sh
There are one more thing to do. When PEAKS is opened, go to Preferences. In the "General" section, change the default project folder to a correct directory in OS X.

Now you can view your PEAKS results on a Mac!



Tuesday, March 19, 2013

New discovery or an error?

We have seen an interesting search result that the precursor mass error plot forms two clusters. The dataset was generated with Thermo LTQ-Orbitrap instrument. Will this be related to some new science discovery or is it caused by an error? Our scientist took a closer look to find out the reason.
In the result summary view, PTM profile section, we noticed that many PSMs have the Deamidation modification.
We then look at some of those PSMs and found out that the precursor mass reported by the instrument is wrong. Instead of reporting the m/z of the monoisotopic peak, in many cases the instrument reported the m/z of the highest isotopic peak. Thus resulting a mass shift of 1Da on the precursor. Here is an example. For this charge 3 peptide, the instrument reports precursor m/z 1057.17, but the monoisotopic peak is at m/z 1056.83.
At this point, we know why there are two clusters on the plot. The software try to explain the wrong precursor mass by adding a Deamidation.

So how can we fix this? We re-run the search. This time, we checked correct precursor mass option in the data refine stage. PEAKS will try to determine the correct precursor mass by look at the survey scan instead of blindly trust values the instrument reported. Comparing the two search results, we not only removed the false Deamidation but also able to explain 10% more spectrum at 1% FDR.

Monday, March 18, 2013

Decoy Fusion on traditional target + decoy database

We were asked a question today by a PEAKS user about FDR result validation. He used PEAKS DB for peptide identification and enabled the built-in decoy fusion method to estimate the FDR. When examining the result, he realized that the FASTA database used for the search is a concatenation of target and decoy proteins. So his question is that is the FDR control still valid or does he have to re-run the search.

The decoy fusion method concatenate the decoy and target sequences of the same protein together as a "fused" sequence (detail explanation can be found here). This ensures that the target and decoy lengths are always the same. If in the searched database, the decoy length is the same as the target length, then PEAKS DB with decoy fusion searched exactly three times the decoy length.

As long as the decoy protein in the searched database is distinguishable, the user can simply discard those hits. The FDR reported by PEAKS is still safe to be used as it only becomes more conservative. 

Friday, March 15, 2013

'Experiment Control' is a valuable tool

There are so many factors that can make database search engines fail to produce a decent result. Today I saw such an interesting case.

A user send in a dataset generated from Orbitrap instrument. When he used PEAKS to analyze the data, surprisingly, he only got very few PSMs. I examined the figures in the 'Experiment Control' section in the 'Summary View' and noticed that the precursor mass error distribution of the PSMs is strange. So I increased the parent ion error tolerance from 10ppm to 20ppm and got very good result. One third of the spectrum have been identified under 1% FDR.

The figures in the 'Summary View' sometimes can help identify the problem when the search result is not ideal. In this case, the instrument is not well calibrated.