PEAKS supports FDR estimation on inChorus results. To make this work on Mascot results, there are a few extra steps to follow.
PEAKS uses Decoy-Fusion method for FDR estimation. The first step is to create a decoy-fusion database. Go to PEAKS database configuration dialog. Select the FASTA database you want to search against. Then click the "Export Decoy DB" button. A decoy-fusion FASTA file will be generated.
The second step is to configure the decoy-fusion FASTA file into Mascot. This is very straightforward in Mascot 2.4 as the parsing rule of PEAKS decoy-fusion method can be automatically detected.
After the decoy-fusion database is up and running on Mascot server, the last step is to make sure that the "Search decoy database from PEAKS" option is selected in the search dialog.
PEAKS is a complete software package for proteomics mass spectrometry data analysis. Starting from the raw mass spectrometry data, PEAKS takes care of every step of data conversion. PEAKS effectively performs peptide and protein identification, PTM and mutation characterization, as well as results validation, visualization and reporting.
Showing posts with label FASTA. Show all posts
Showing posts with label FASTA. Show all posts
Thursday, July 4, 2013
Monday, May 6, 2013
Configure FASTA database in PEAKS
Configuring FASTA databases in PEAKS is fairly easy especially if the FASTA file has the same header format as one of the public databases (e.g. NR, Swiss-Prot, IPI). It is just a matter of selecting the pre-defined format and the parsing rules will be automatically filled in.
There are also a large number of users use PEAKS to search on their in-house, customized FASTA databases. In this situation, the header format is very hard to predict and it varies case by case.
In PEAKS, the parsing rule is defined using regular expression. While regular expression is very powerful, it will take people quite a bit of time to master it. Since we got tons of searches to run every week, against FASTA files with so many different header formats, I created this lazy, generic parsing rule for internal use and in most cases, it worked good enough.
Accession. The regular expression tries to use everything before the first white space as the accession. If no white space were found within the first 30 characters, the first 30 characters will be used as accession.
There are also a large number of users use PEAKS to search on their in-house, customized FASTA databases. In this situation, the header format is very hard to predict and it varies case by case.
In PEAKS, the parsing rule is defined using regular expression. While regular expression is very powerful, it will take people quite a bit of time to master it. Since we got tons of searches to run every week, against FASTA files with so many different header formats, I created this lazy, generic parsing rule for internal use and in most cases, it worked good enough.
Accession. The regular expression tries to use everything before the first white space as the accession. If no white space were found within the first 30 characters, the first 30 characters will be used as accession.
>\([^\s|]{1,30}\)Description. The whole line after ">" will be used as the description.
>\(.*\)
Subscribe to:
Posts (Atom)