Showing posts with label PEAKS DB. Show all posts
Showing posts with label PEAKS DB. Show all posts

Thursday, January 30, 2014

PEAKS 7 Viewer on Mac OS X Mavericks

I wrote a post in early 2013 here to help users to run PEAKS as a Viewer on Mac OS or Linux. Similar procedure applies to PEAKS 7 as well. Many thanks to user ceinwyn that brought the issue to me that PEAKS DB search results cannot be opened.

I spent some spare time looking into this issue and luckily found a workaround. Just want to re-state my disclaimer here.

Disclaimer: PEAKS does not officially support any OS other than Windows as of the time I am writing the post. The software may not be fully functional. Activating the software on OS X or Linux will consume the license, which means the same license can not be used again. I strongly recommend only following the steps to configure PEAKS Viewer (the unlicensed Studio) on OS X or Linux for PEAKS result sharing and presentation purposes.

Basically what happened here is that PEAKS 7 has some sizing issue with the Auqa look and feel from Apple Mac OS X. The workaround is to change the look and feel to a cross-platform one that Apple also support. To make things easier, I have put a modified version of the jar file here. And here are the steps:
  1. Check the version of PEAKS. Go to Help -> About PEAKS. In the dialog, you should see the build number, e.g. build 20131119. This build number must match the number in the downloaded jar file in the above link.
  2. Close PEAKS.
  3. Download the jar file and replace the peaksstudio.jar in the PEAKS directory with this one.
  4. Start PEAKS and now the PEAKS DB search results can be opened. 

Enjoy PEAKS and happy sharing!

Monday, June 24, 2013

Carbamidomethyl @ C, D, E, H, K, N-term

Recently I worked on an ETD dataset generated from Orbitrap Velos. The user mentioned that Carbamidomethyl on Cysteine and some phosphopeptides are expected but he was not able to get any good identification results using Mascot.

My first attempt is to use the information provided by the user and run the PEAKS DB search. The result shocked me as under 1% FDR, PEAKS DB only reported 170 PSMs. What could possibly go wrong?

Looking at the PEAKS DB result, there are many "de novo only" peptides, which means many spectrum can produce confident de novo sequences but they do not have a confident database hit. This could be a result of unsuspected PTMs and mutations.

So I decided to do a PEAKS PTM search on the db result as my second attempt. PEAKS PTM reported 582 PSMs under 1% FDR. In the summary view, PTM profile section, there are many PSMs found with the PTM Carbamidomethyl on D, E, H, K and N-term. This could be an indication of excess iodoacetamide during alkylation procedure.

Now I know more PTM information of this data. In the last attempt, I added Carbamidomethyl @ DEHK,N-term and Dehydration @ DST into the variable PTM list along with Phosphorylation @ STY. PEAKS DB search was then performed with PEAKS PTM option enabled. This time, under 1% FDR, 736 PSMs were reported, more than four times the number of the initial search.

PEAKS PTM is a great tool to find unsuspected PTMs thus help explaining more spectrum.

Monday, May 13, 2013

Multiple enzymes support in PEAKS - Full Protein Coverage

PEAKS 6 introduced a new feature specifically targeting the experiments that use multiple enzyme digestions to increase protein coverage.

In the past, users have to search each sample separately and combine all the results manually afterwards or using none enzyme option to analyze all samples in one go which may cause higher false positives. Now in PEAKS, users can specify enzyme for each sample when creating a multi-sample project. Then in de novo sequencing and PEAKS DB, user can choose 'sample enzyme' in the enzyme list as the search option. PEAKS will use the correct enzyme when analyzing each sample.

From our users feedback, this feature is extremely useful when you want to fully characterize a single protein.The following example shows how big a difference this feature may make.

ALBU_BOVIN protein ordered from a reputable vendor was digested with Trypsin, LysC, GluC. The dataset is generated from Thermo Orbitrap instrument. Three searches were performed. The first one uses inChorus function to launch Mascot search (version 1.4) on the trypsin sample only. The second search uses standard PEAKS DB search on the trypsin sample. The third search uses the complete analysis workflow, including PEAKS PTM and SPIDER, on all three samples and uses "sample enzyme" as the enzyme option. The results are all filtered to only keep the very confident PSMs at 0.1% FDR level.

Mascot and PEAKS DB are able to achieve 73% and 86% protein coverage using only the trypsin sample respectively. In the protein coverage view below, the blue bars are the PSMs that matched the protein sequence at that position.

PEAKS complete analysis on all three samples reported 96% coverage on the protein. The uncovered 4% is in the protein N-terminal region, which is most likely cleaved-off and not in the purchased sample1.
1specific binding site (Asp-Thr-His-Lys) for Cu(II) ions. T. Peters Jr., F.A. Blumenstock. J. Biol. Chem., 242 (1967), p. 1574

Friday, April 26, 2013

Common ptifalls of FDR estimation part three


The third pitfall is also caused due to the over-emphasis on sensitivity.

There is another trend in database search software to re-score the peptide identification results by using machine learning. The idea is straightforward: After the search, we know what the decoy hits are. The algorithm should take advantage of it, and retrain the parameters of the scoring function to get rid of the decoy hits. With this effort, it will get rid of a lot of the target false hits as well.

The method is valid, except that it may cause FDR underestimation. This is because the target false hits are unknown to the machine learning algorithm. Therefore, there is a risk that the machine learning algorithm removes more decoy hits than the target false hits.
This overfit risk is well known in machine learning. A machine learning expert can reduce the risk but can never get rid of it. 

The solution to this pitfall number 3 is trickier.

The first suggestion: don’t use it. The philosophy here is that judges cannot be players. If we want to use the decoy for result validate, the decoy information should never be released to the search algorithm.

If this re-scoring method must be used due to the low-performance of some database search software, it should only be used for very large dataset to reduce the risk of over-fit.

Perhaps the best solution is the third one. That is, the retraining of the score parameters should be done for each different instrument type, instead of each dataset. This will gain much of the benefit provided by machine learning, but without the problem of over-fitting. Indeed, this third approach is what we do in the PEAKS DB algorithm.


*The content of this post is extracted from "Practical Guide to Significantly Improve Peptide Identification Sensitivity and Accuracy" by Dr. Bin Ma, CTO of Bioinformatics Solutions Inc. You can find the link to the guide on this page.

Monday, April 22, 2013

Common ptifalls of FDR estimation part two


The second pitfall of the traditional target-decoy strategy is caused by another popular technique used to increase the peptide identification sensitivity.  

The idea is clever: if a weakly identified peptide happens to be on a highly-confident protein, then the peptide is likely to be correct regardless of its low score. So, to increase the sensitivity, the software can add a score bonus to each peptide on a multiple-hit protein. Indeed, this protein bonus will save some weak true hits, but it will save some weak false hits at the same time. The bigger problem is that the target database will provide more multiple-hit proteins than the decoy. As a result, more weak false hits will be saved from the target database. This will cause the FDR underestimation.
In PEAKS,
decoy fusion approach can solve this problem effectively.
 

Because the target and decoy sequences are concatenated into a single protein sequence, when a protein bonus is added to the multiple-hit proteins, the same bonus will be added to the target and decoy hits equally. So, weak false hits are saved with approximately equal probabilities in the target and decoy. This recreates the balance and provides accurate FDR estimation.

By using the decoy fusion as the validation method, we can safely apply the protein bonus. We get the sensitivity, but did not compromise  the FDR estimation.

 
*The content of this post is extracted from "Practical Guide to Significantly Improve Peptide Identification Sensitivity and Accuracy" by Dr. Bin Ma, CTO of Bioinformatics Solutions Inc. You can find the link to the guide on this page.

Friday, April 19, 2013

Common ptifalls of FDR estimation part one

Today’s most widely used method for FDR estimation is the target-decoy strategy. This is a well-established method in statistics and started to be used in proteomics around 2007.

In this approach, a decoy database that contains the same number of proteins as the target database are searched together by the database search engine to identify peptides. The blue colors indicate the target hits and the orange colors indicate the decoy hits, the squares are the false hits, and circles are true hits. 
The decoy proteins are randomly generated so that any decoy hit is supposedly a false hit. Since the search engine doesn’t know which sequences are from target and which are from decoy, when it makes a mistake, the mistake falls in the target and decoy databases with equal probability. Thus, the total number of false target hits can be approximated by the number of decoy hits in the final result. And the FDR can be estimated by the ratio between the numbers of decoy hits and the number of target hits. 

The target-decoy strategy is a powerful method for FDR estimation. However, as we will discover in the next little while, such a powerful method must be used with caution to avoid FDR underestimation. 

The first pitfall in the use of target-decoy approach for FDR estimation is due to the so-called multiple round search strategy in today’s database search software. 

This multi-round search was popularized by the X!Tandem program published in 2004, in order to speed up the computation. The first round uses a fast but less sensitive search method to quickly identify a shortlist of proteins from the large database. Then, the second round uses a more sensitive but slower search method to identify peptides, but only from the short list of proteins. This effectively speeds up the search without sacrificing too much sensitivity. Indeed, X!Tandem is one of the fastest search algorithm used today.

However, as pointed out by a paper published in JPR in 2010, this multiple-round search strategy screws up the target-decoy estimation of the FDR. The reason is that after the first round, there will be more target proteins than the decoy in the short list. Thus, if the second round search makes a mistake, the mistake will be more likely in the target proteins. So, we will end up with fewer decoy hits than the actual false target hits. This causes the FDR underestimation.
The JPR paper in 2010 provided a fix to this problem. But a year later, in another JPR paper, Bern and
Kil pointed out that the fix was wrong, and proposed a different fix that required the change of the search engine’s algorithm. This shows that the FDR estimation is very tricky, even the experts can sometimes get it wrong. 


In PEAKS, we used a new approach, called decoy fusion to solve this problem. 

Instead of mixing the target and decoy databases, we append a decoy sequence to each target protein.

So, after the fast search round, the protein shortlist will still contain the same length of target and decoy sequences. And the false hits of the second round will have the equal chance to be from the target and decoy sequences. This recreates the balance and can accurately estimate the FDR in the multiple-round search setting.


*The content of this post is extracted from "Practical Guide to Significantly Improve Peptide Identification Sensitivity and Accuracy" by Dr. Bin Ma, CTO of Bioinformatics Solutions Inc. You can find the link to the guide on this page.

Saturday, April 6, 2013

Update: 3.4 million MS2 spectrum project, PEAKS DB completed!

Both PEAKS de novo and PEAKS DB search were finished successfully on the 3.4 million MS2 dataset. The computer has two Xeon hex-core CPUs and 32GB RAM running Windows 7 Pro 64 bit OS. The search completed in a little over a week.
PEAKS de novo reported over 2 million de novo peptides that have their ALC score greater than or equal to 50%.

PEAKS DB reported over 1 million PSMs at 1% FDR.

Be sure to check out the protein coverage view of one of the top score proteins. It can help demonstrate that a protein is pretty much fully covered.

Monday, March 18, 2013

Decoy Fusion on traditional target + decoy database

We were asked a question today by a PEAKS user about FDR result validation. He used PEAKS DB for peptide identification and enabled the built-in decoy fusion method to estimate the FDR. When examining the result, he realized that the FASTA database used for the search is a concatenation of target and decoy proteins. So his question is that is the FDR control still valid or does he have to re-run the search.

The decoy fusion method concatenate the decoy and target sequences of the same protein together as a "fused" sequence (detail explanation can be found here). This ensures that the target and decoy lengths are always the same. If in the searched database, the decoy length is the same as the target length, then PEAKS DB with decoy fusion searched exactly three times the decoy length.

As long as the decoy protein in the searched database is distinguishable, the user can simply discard those hits. The FDR reported by PEAKS is still safe to be used as it only becomes more conservative.