Showing posts with label mass spectrometry. Show all posts
Showing posts with label mass spectrometry. Show all posts

Monday, 24 February 2020

ThermoRAWFileParser: A small step towards cloud proteomics solutions

Proteomics data analysis is in the middle of a big transition. We are moving from small experiments (e.g. a couple of RAW files, samples) to big large scale experiments. While the average number of RAW files per datasets in PRIDE hasn't grown in the last 6 years (Figure 1), we can see multiple experiments with more than 1000 RAW files (Figure 1 - right).

Figure 1: The boxplot of the number of files per dataset in PRIDE (left - outliers removed; right - outliers included) 

On the other side, File size shows a trend towards large RAW files (Figure 2).

Figure 2: Box plot of file size by datasets in PRIDE (outliers removed)

Then, how proteomics data analysis can be moved towards large scale and elastic compute architectures such as Cloud infrastructures or High-performance computing (HPC) clusters?

Thursday, 9 June 2016

How to estimate and compute the isoelectric point of peptides and proteins?

By +Yasset Perez-Riverol and +Enrique Audain :

Isoelectric point (pI) can be defined as the point of singularity in a titration curve, corresponding to the solution pH value at which the net overall surface charge is equal to zero. Currently, there are available different modern analytical biochemistry and proteomics methods depend on the isoelectric point as a principal feature for protein and peptide characterization. Peptide/Protein fractionation according to their pI is widely used in current proteomics sample preparation procedures previous to the LC-MS/MS analysis. The experimental pI records generated by pI-based fractionation procedures are a valuable information to validate the confidence of the identifications, to remove false positive and and could be used to re-compute peptide/protein posterior error probabilities in MS-based proteomics experiments. 

Theses approaches require an accurate  theoretical prediction of pI. Even thought there are several tools/methods to predict the isoelectric point, it remains hard to define beforehand what methods perform well on a specific dataset.  

We believe that the best way to compute the isoelectric point (pI) is to have a complete package with most of the algorithms and methods in the state of the art that can do the job for you [2]. We recently developed an R package (pIR) to compute isoelectric point using long-standing and novels pI methods that can be grouped in three main categories : a) iterative, b) Bjellvist-based methods and c) machine learning methods. In addition, pIR also offers a statistical and graphical framework to evaluate the performance of each method and its capability to “detect” outliers (those peptides/protein with theoretical pI biased from experimental value) in a high-throughput environment.

First lets install the package:

First, we need to install devtools:
install.packages("devtools")
library(devtools)
Then we just call:
install_github("ypriverol/pIR")
library(pIR)

Friday, 5 September 2014

NEW NIST 2014 mass spectral library

Originally posted in NIST 2014.

Identify your mass spectra with the new NIST 14 Mass Spectral Library and Search Software.

NIST 14 - The successor to NIST 11 (2011) - Is a collection of:


  • Electron ionization (EI) mass spectra
  • Tandem MS/MS spectra (ion trap and collision cell)
  • GC method and retention data
  • Chemical structures and names
  • Software for searching and identifying your mass spectra
  • NIST 14 is integrated with most mass spectral data systems, including Agilent ChemStation/MassHunter, Thermo Xcalibur, and others. The NIST Library is known for its high quality, broad coverage, and accessibility. It is a product of a three decade, comprehensive evaluation and expansion of the world's most widely used and trusted mass spectral reference library compiled by a team of experienced mass spectrometrists in which each spectrum was examined for correctness.


Improvements from 2011 version:


  • Increased coverage in all libraries: 32,355 more EI spectra; 138,875 more MS/MS spectra; 37,706 more GC data sets
  • Retention index usable in spectral match scoring
  • Improved derivative naming, user library features, links to InChIKey, and other metadata.
  • Upgrade discount for any previous version
  • Lowest Agilent format price available

MS/MS and GC libraries may now be optionally purchased separately at very low cost
Learn what`s new http://www.sisweb.com/software/ms/nist.htm#whatsnew

 Pick related PDFs

Sunday, 6 April 2014

SWATH-MS and next-generation targeted proteomics

For proteomics, two main LC-MS/MS strategies have been used thus far. They have in common that the sample proteins are converted by proteolysis into peptides, which are then separated by (capillary) liquid chromatography. They differ in the mass spectrometric method used.

The first and most widely used strategy is known as shotgun proteomics or discovery proteomics. For this method, the MS instrument is operated in data-dependent acquisition (DDA) mode, where fragment ion (MS2) spectra for selected precursor ions detectable in a survey (MS1) scan are generated (Figure 1 - Discovery workflow). The resulting fragment ion spectra are then assigned to their corresponding peptide sequences by sequence database searching (See Open source libraries and frameworks for mass spectrometry based proteomics: A developer's perspective).

The second main strategy is referred to as targeted proteomics. There, the MS instrument is operated in selected reaction monitoring (SRM) (also called multiple reaction monitoring) mode (Figure 1 - Targeted Workflow). With this method, a sample is queried for the presence and quantity of a limited set of peptides that have to be specified prior to data acquisition. SRM does not require the explicit detection of the targeted precursors but proceeds by the acquisition, sequentially across the LC retention time domain, of predefined pairs of precursor and product ion masses, called transitions, several of which constitute a definitive assay for the detection of a peptide in a complex sample (See Targeted proteomics) .

Figure 1 - Discovery and Targeted proteomics workflows

Monday, 3 March 2014

Most read from the Journal of Proteome Research for 2013.

1- Protein Digestion: An Overview of the Available Techniques and Recent
    Developments

    Linda Switzar, Martin Giera, Wilfried M. A. Niessen

    DOI: 10.1021/pr301201x

2-  Andromeda: A Peptide Search Engine Integrated into the MaxQuant
     Environment

     Jürgen Cox, Nadin Neuhauser, Annette Michalski, Richard A. Scheltema, Jesper
     V. Olsen, Matthias Mann

     DOI: 10.1021/pr101065j

2- Evaluation and Optimization of Mass Spectrometric Settings during
     Data-dependent Acquisition Mode: Focus on LTQ-Orbitrap Mass Analyzers
 
     Anastasia Kalli, Geoffrey T. Smith, Michael J. Sweredoski, Sonja Hess

     DOI: 10.1021/pr3011588

3-  An Automated Pipeline for High-Throughput Label-Free Quantitative
     Proteomics

     Hendrik Weisser, Sven Nahnsen, Jonas Grossmann, Lars Nilse, Andreas Quandt,
     Hendrik Brauer, Marc Sturm, Erhan Kenar, Oliver Kohlbacher, Ruedi Aebersold,
     Lars Malmström

     DOI: 10.1021/pr300992u

4-  Proteome Wide Purification and Identification of O-GlcNAc-Modified Proteins
     Using Click Chemistry and Mass Spectrometry

     Hannes Hahne, Nadine Sobotzki, Tamara Nyberg, Dominic Helm, Vladimir S.
     Borodkin, Daan M. F. van Aalten, Brian Agnew, Bernhard Kuster

     DOI: 10.1021/pr300967y

5-  A Proteomics Search Algorithm Specifically Designed for High-Resolution
     Tandem Mass Spectra

     Craig D. Wenger, Joshua J. Coon
   
     DOI: 10.1021/pr301024c

6- Analyzing Protein–Protein Interaction Networks

    Gavin C. K. W. Koh, Pablo Porras, Bruno Aranda, Henning Hermjakob, Sandra E.
    Orchard

    DOI: 10.1021/pr201211w

7-  Combination of FASP and StageTip-Based Fractionation Allows In-Depth
     Analysis of the Hippocampal Membrane Proteome

     Jacek R. Wisniewski, Alexandre Zougman, Matthias Mann

     DOI: 10.1021/pr900748n

8-  The Biology/Disease-driven Human Proteome Project (B/D-HPP): Enabling
     Protein Research for the Life Sciences Community

     Ruedi Aebersold, Gary D. Bader, Aled M. Edwards, Jennifer E. van Eyk, Martin
     Kussmann, Jun Qin, Gilbert S. Omenn

     DOI: 10.1021/pr301151m

 9-  Comparative Study of Targeted and Label-free Mass Spectrometry Methods
      for Protein Quantification

       Linda IJsselstijn, Marcel P. Stoop, Christoph Stingl, Peter A. E. Sillevis Smitt,
       Theo M. Luider, Lennard J. M. Dekker

       DOI: 10.1021/pr301221f

Wednesday, 8 January 2014

News: The first version of PRIDE Inspector 2.0 is now available

PRIDE Inspector 2.0 is an integrated desktop application for MS Proteomics data analysis and visualization. 
- The current version support PRIDE XML, mzIdentML, as well as providing direct access to PRIDE public database. 
- The new version also support of Mass Spectra Formats such as mzxml, mgf, pkl, ms2, dta.
- Some of the new features are: Fragmentation Visualization.
- Protein and Peptide Group Visualization.
- Visualization of Peptide and Protein Properties (Scores, pI, etc)

- New Chart Options.
- Others ...



Links:

 

Wednesday, 9 October 2013

My List of Most Influential Authors in Computational Proteomics (according to Articles References, Google Scholar, twitter, Linkedin, Microsoft Academic Search and ResearchGate)

Young researchers starting their careers will often look for reviews, opinions and research manuscripts from the most influential authors of their chosen field. In science, however, unlike many other topics on the Internet, ranked lists or manuscript repositories of top authors sorted by research topic are hard to come by. For some researchers, the idea of such a task brings the words ‘wasted time’ to their minds; the most critical condemn it as a frivolous pursuit. Maybe so. In my opinion, however, it as an excellent starting point.

ResearchGate Home page
Home Page of ResearchGate with more than 3 millions of users

These days, more people than ever are involved in science and research. Just look at ResearchGate’s homepage.  There are over 3 million persons there –and we’re only counting ResearchGate users. Once simple undertakings, such as finding the right manuscript to cite, the most authoritative group on a topic, or the best software application for a specific task, have become increasingly difficult for graduate students navigating this ocean of data, despite the availability of services such as Google Scholar or Pubmed. The situation will only worsen in the future, as is easy to see by simply tallying the number of  published papers in the fields of Proteomics, Genomics, Bioinformatics and Computational Proteomics since 1997:

Number of published manuscripts in Pubmed per year (1997-2012). the statistics was done using the Medline Trend Service http://dan.corlan.net/medline-trend.html

In 2012 alone, over 6,000 and 17,000 manuscripts were published in the fields of proteomics and bioinformatics, respectively. Our young field, computational proteomics, published more than four hundred papers. Perhaps well-established PI’s or Group Leaders can easily tell apart derivative or me-too contributions from groundbreaking work, but young scientists, who spend most of their time implementing someone else’s ideas, can certainly have a hard time doing so. Although technology has come to the rescue with today’s mixture of search engines and social networking tools (ResearchGate, Google Scholar, twitter and LinkedIn among them), the best way to harness its power is, precisely, by starting from a ranked list of the most authoritative voices within a field of research, whose whereabouts can then be traced in the scientific literature, the blogosphere, and anywhere else.


Wednesday, 6 March 2013

#HavanaBioinfo2012 Hard, but Awesome Experience


From 8th to 11th of last December the "I Bioinformatics for Biotechnology Applications" (#HavanaBioinfo2012) was held in the hotel “Occidental Miramar” located within an elegant area in Havana, Cuba. Putting on a Bioinformatic workshop takes a lot of different pieces (small ones and big ones). You need write lot of mails to invited speakers, you need a nice and comfortable place with lots of tables and chairs and good food and the drinks. 

I started from October, writing the first mails to EBI friends, advisers and professors (www.ebi.ac.uk), and to be complete honest it was fantastic because most if then accepted from the very beginning. Thanks to Henning Hermjakob and Alex Bateman (@Alexbateman1) for the support. At the end the workshop was fully subscribed, with more than 45 attendees and sixteen speakers, participating in two poster sessions and a panel discussion. Speakers from EBI, Belgium University, Mascot (UK), Bioinformatics Solutions (Canada) accepted the invitation to come here and give one or two lectures about proteomics, genomics, bioinformatics.


Saturday, 25 August 2012

Why R for Mass Spectrometrist and Computational Proteomics

Why R:

Actually, It is a common practice the integration of the statistical analysis of the resulted data and in silico predictions of the data generated in your manuscript and your daily research. Mass spectrometrist, biologist and bioinformaticians commonly use programs like excel, calc or other office tools to generate their charts and statistical analysis. In recent years many computational biologists especially those from the Genomics field, regard R and Bioconductor as fundamental tools for their research.

R is a modern, functional programming language that allows for rapid development of ideas; it is a language and environment for statistical computing and graphics.The rich set of inbuilt functions makes it ideal for high-volume analysis or statistical studies.


Monday, 21 May 2012

An "in-house" Tool

One of the small hidden details in publications, even in those with a higher impact, is the use of "in-house programs". What is an "in-house" program or tool: Normally is a piece of software that researchers use to analyze process or visualize the experimental data, but most important the software it-self is not published

The term by itself is inoffensive, but the concept could be extremely dangerous. We can cite hundreds of manuscripts that included in the data analysis "in-house" tools, but never the terms "in-house instruments". The authors always needs to cite the manufacturer, the reagents, even the year and the company. I know, we have a section to describe data processing but mostly we cite some parameters, and the well known software like search engines (Mascot, X!Tandem, Sequest, etc). But at some point of this section several times you can find the term "in-house" tool. It could be a reference to an excel formula or to a complete and complex java program with many tasks like parsing a search engine output, computing the FDR, removing false-positive identifications, computing peptide-spectrum-match redundancy, etc. The are not a real/objective measure to distinguish between a little-simple tool and a complex tool one.