Showing posts with label ms proteomics. Show all posts
Showing posts with label ms proteomics. Show all posts

Monday, 24 February 2020

ThermoRAWFileParser: A small step towards cloud proteomics solutions

Proteomics data analysis is in the middle of a big transition. We are moving from small experiments (e.g. a couple of RAW files, samples) to big large scale experiments. While the average number of RAW files per datasets in PRIDE hasn't grown in the last 6 years (Figure 1), we can see multiple experiments with more than 1000 RAW files (Figure 1 - right).

Figure 1: The boxplot of the number of files per dataset in PRIDE (left - outliers removed; right - outliers included) 

On the other side, File size shows a trend towards large RAW files (Figure 2).

Figure 2: Box plot of file size by datasets in PRIDE (outliers removed)

Then, how proteomics data analysis can be moved towards large scale and elastic compute architectures such as Cloud infrastructures or High-performance computing (HPC) clusters?

Friday, 11 September 2015

An API for all MS-based File formats

We recently released and published our first Java API (Application Programming Interface) for the most common file formats in proteomics, not only ms files but also identification files such as mzIdentML and mztab. 

ms-data-core-api (https://github.com/PRIDE-Utilities/ms-data-core-api)

The library allow the end-users and the developers to use a common data structure for proteomics independently of the file types, and .. But first lets try to understand what is a API.

What is an API?

Imagine you are a builder or civil engineering and your are building your bridge, different components, blocks and different teams needs to be coordinated and plugged for the final results. Wrong communications between the members of the teams, different block sizes or building plans only produced strange results. 

In the simplest terms, APIs are sets of requirements, data structures, objects that govern how applications and software components can talk each other. An API, is a set of routines and protocols that provide building blocks for computer programmers and web developers to build software applications. In the past, APIs were largely associated with computer operating systems and desktop applications. In recent years though, we have seen the emergence of Web APIs (Web Services).


What is ms-data-core-api?

The ms-data-core-api is a free, open-source library for developing computational proteomics tools and pipelines. The Application Programming Interface, written in Java, enables rapid tool creation by providing a robust, pluggable programming interface and common data model. The data model is based on controlled vocabularies/ontologies and captures the whole range of data types included in common proteomics experimental workflows, going from spectra to peptide/protein identifications to quantitative results. 

The library contains readers for three of the most used Proteomics Standards Initiative standard file formats: mzML, mzIdentML, and mzTab. In addition to mzML, it also supports other common mass spectra data formats: dta, ms2, mgf, pkl, apl (text-based), mzXML and mzData (XML-based). Also, it can be used to read PRIDE XML, the original format used by the PRIDE database, one of the world-leading proteomics resources. Finally, we present a set of algorithms and tools whose implementation illustrates the simplicity of developing applications using the library.

Friday, 2 January 2015

Brazil: A place for Science and Friendship


Búzios
Búzios
It's really difficult to break stereotypes, especially for developing countries, like Brazil. If you mention its name around the world they are immediately associated with: sports, music, beaches, rum and "País do Carnaval". If you ask to someone in the streets of Germany or China about personalities from Brazil, they will mention Pelé. Breaking stereotypes is a task for years or centuries but we are going in the right direction.

Hotel Ferradura/ Ferradura Resort
Last December I attended to the 2nd Proteomics Meeting of the Brazilian Proteomics Society jointly with the 2nd Pan American HUPO Meeting in Hotel Ferradura/ Ferradura Resort, Búzios, Rio de Janeiro State, Brazil. The venue was gorgeous, mountains close to a small bay that offers calm, clear waters and the open sea. We arrived after 2 hours by car from Rio international airport. My plans, give a talk about PRIDE and ProteomeXchange but more than that, my talk was about "if we really need to share our proteomics data".  

Tuesday, 25 November 2014

HUPO-PSI Meeting 2014: Rookie’s Notes

Standardisation: the most difficult flower to grow.
The PSI (Proteomics Standard Initiative) 2014 Meeting was held this year in Frankfurt (13-17 of April) and I can say I’m now part of this history. First, I will try to describe with a couple of sentences (for sure I will fai) the incredible venue, the Schloss Reinhartshausen Kempinski. When I saw for the first time the hotel, first thing came to my mind was those films from the 50s. Everything was elegant, classic, sophisticated - from the decoration to a small latch. The food was incredible and the service is first class from the moment you set foot on the front step and throughout the whole stay. 
  
Standardization is the process of developing and implementing technical standards. Standardization can help to maximize compatibility, interoperability, safety, repeatability, or quality. It can also facilitate commoditization of formerly custom processes. In bioinformatics, the standardization of file formats, vocabulary, and resources is a job that all of us appreciate but for several reasons nobody wants to do. First of all, standardization in bioinformatics means that you need to organize and merge different experimental and in-silico pipelines to have a common way to represent the information. In proteomics for example, you can use different sample preparation, combined with different fractionation techniques and different mass spectrometers; and finally using different search engines and post-processing tools. The diversity and possible combinations is needed because allow to explore different solutions for complex problems. (Standarization in Proteomics: From raw data to metadata files).

Thursday, 23 October 2014

Which journals release more public proteomics data!!!

I'm a big fan of data and the -omics family. Also, I like the idea of make more & more our data public available for others, not only for reuse, but also to guarantee the reproducibility and quality assessment of the results (Making proteomics data accessible and reusable: Current state of proteomics databases and repositories). I'm wondering which of these journals (list - http://scholar.google.co.uk/) encourages their submitters and authors to make their data publicly available:



Journal
h5-index
h5-median
Molecular & Cellular Proteomics
74
101
Journal of Proteome Research
70
91
Proteomics
60
76
Biochimica et Biophysica Acta (BBA)-Proteins and Proteomics
52
78
Journal of Proteomics
49
60
Proteomics - Clinical Applications
35
43
Proteome Science
23
32

After a simple statistic, based on PRIDE data:


Number of PRIDE projects by Journal

Saturday, 4 October 2014

Analysis of histone modifications with PEAKS 7: A respond to Search Engines comparison from PEAKs Team

Recently we posted a comparison of different search engines for PTMs studies (Evaluation of Proteomic Search Engines for PTMs Identification). After some discussion of the mentioned results in our post the  PEAKS Team just published a blog post with the reanalysis of the dataset. Here the results:

Originally Posted in Peaks Blog:
The complex nature of histone modification patterns has posed as a challenge for bioinformatics analysis over the years. Yuan et al. [1] conducted a study using two datasets from human HeLa histone samples, to benchmark the performance of current proteomic search engines. This article was published in J Proteome Res. 2014 Aug 28 (PubMed), and the data from the two datasets, HCD_Histone and CID_Histone (PXD001118), was made publically available through ProteomeXchange. With this data, the article uses eight different proteomic search engines to compare and evaluate the performance and capability of each. The evaluated search engines in this study are: pFind, Mascot, SEQUEST, ProteinPilot, PEAKS 6, OMSSA, TPP and MaxQuant. 
In this study, PEAKS 6 was used to compare the performance capabilities between search engines. However, PEAKS 7, which was released November 2013, is the latest version available of the PEAKS Studio software. PEAKS 7 not only includes better performance than PEAKS 6, but a lot of additional and improved features. Our team has reanalyzed the two datasets HCD_Histone and CID_Histone with PEAKS 7 to update the ID results presented in the publication by Yuan et al.  These updated results showed that instead, it is PEAKS, pFind and Mascot that identify the most confident results.

Friday, 5 September 2014

NEW NIST 2014 mass spectral library

Originally posted in NIST 2014.

Identify your mass spectra with the new NIST 14 Mass Spectral Library and Search Software.

NIST 14 - The successor to NIST 11 (2011) - Is a collection of:


  • Electron ionization (EI) mass spectra
  • Tandem MS/MS spectra (ion trap and collision cell)
  • GC method and retention data
  • Chemical structures and names
  • Software for searching and identifying your mass spectra
  • NIST 14 is integrated with most mass spectral data systems, including Agilent ChemStation/MassHunter, Thermo Xcalibur, and others. The NIST Library is known for its high quality, broad coverage, and accessibility. It is a product of a three decade, comprehensive evaluation and expansion of the world's most widely used and trusted mass spectral reference library compiled by a team of experienced mass spectrometrists in which each spectrum was examined for correctness.


Improvements from 2011 version:


  • Increased coverage in all libraries: 32,355 more EI spectra; 138,875 more MS/MS spectra; 37,706 more GC data sets
  • Retention index usable in spectral match scoring
  • Improved derivative naming, user library features, links to InChIKey, and other metadata.
  • Upgrade discount for any previous version
  • Lowest Agilent format price available

MS/MS and GC libraries may now be optionally purchased separately at very low cost
Learn what`s new http://www.sisweb.com/software/ms/nist.htm#whatsnew

 Pick related PDFs

Monday, 20 January 2014

Some of the most cited manuscripts in Proteomics and Computational Proteomics (2013)

Some of the most cited manuscripts in 2013 in the field of Proteomics and Computational Proteomics (no order):







     The PRoteomics IDEntifications (PRIDE, http://www.ebi.ac.uk/pride) database 
     at the European Bioinformatics Institute is one of the most prominent data 
     repositories of mass spectrometry (MS)-based proteomics data. Here, we 
     summarize recent developments in the PRIDE database and related tools. 
     First, we provide up-to-date statistics in data content, splitting the figures by 
     groups of organisms and species, including peptide and protein 
     identifications, and post-translational modifications. We then describe the 
     tools that are part of the PRIDE submission pipeline, especially the recently 
     developed PRIDE Converter 2 (new submission tool) and PRIDE Inspector 
     (visualization and analysis tool). We also give an update about the integration 
     of PRIDE with other MS proteomics resources in the context of the 
     ProteomeXchange consortium. Finally, we briefly review the quality control 
     efforts that are ongoing at present and outline our future plans.

Wednesday, 8 January 2014

News: The first version of PRIDE Inspector 2.0 is now available

PRIDE Inspector 2.0 is an integrated desktop application for MS Proteomics data analysis and visualization. 
- The current version support PRIDE XML, mzIdentML, as well as providing direct access to PRIDE public database. 
- The new version also support of Mass Spectra Formats such as mzxml, mgf, pkl, ms2, dta.
- Some of the new features are: Fragmentation Visualization.
- Protein and Peptide Group Visualization.
- Visualization of Peptide and Protein Properties (Scores, pI, etc)

- New Chart Options.
- Others ...



Links:

 

Thursday, 24 October 2013

Creating an Open Source Revolution in Computational Proteomics

First of all, I don’t want to discuss in this post about Open-Source, its strengths & strengths. This post is about the most useful Open-Source packages, frameworks or libraries in the field of computational proteomics (a short version of our manuscript “Open source libraries and frameworks for Mass Spectrometry based Proteomics: A developer’s perspective”). 


Schema of the possible computational processing steps of a proteomics data set.

In proteomics like other Omics, the bioinformatics efforts can be divided in three major fields: data processing, storage and visualization. From MS/MS preprocessing to post-processing of the identifications results, even though the objectives of these libraries and packages can vary significantly, they usually share a number of features. Common use cases include the handling of protein and peptide sequences, the parsing of results from various proteomics search engines output files, and the visualization of MS-related information (including mass spectra and chromatograms).

Tuesday, 1 October 2013

Celebrating Ten Years of Mann and Aebersold’s “Mass spectrometry-based proteomics” review.

In 2003 Mann & Aebersold reviewed on the pages of Nature the challenges and perspectives of the then-nascent field of MS-based proteomics. Mass spectrometry (MS) has since entrenched itself as the method of choice for analyzing complex protein samples, and MS-based proteomics has become an indispensable technology for interpreting genomic data and performing protein analyses (primary sequence, post-translational modifications (PTMs) or protein–protein interactions).

" The ability of mass spectrometry to identify and, increasingly, to precisely quantify thousands of proteins from complex samples can be expected to impact broadly on biology and medicine."
The manuscript by Mann & Aebersold is one of the most cited manuscripts in the field of MS proteomics, For this reason is one of the “core papers” in the field of proteomics and computational proteomics, outlining most of the basic concepts required to understand the fundamentals of this discipline.

Ten years after its publication the main workflow described in the manuscript do not change dramatically. In this period major advances are related to the development of the Thermo’s Orbitrap Mass Spectrometer (Velos, LTQ, Exactive, etc) and new fragmentations types (ETD, HCD). Separation techniques (electrophoretic and chromatographic) were explored extensively in these ten years. Aebersold pioneered in 2005 the use of OFFGEL electrophoresis and electrophoresis fragmentation at peptide level (Heller 2005) and Mann’s group developed the FASP method for sample preparation before protein digestion (Wiśniewski JR et al 2009),both of which have contributed significantly to the dramatic increase in the number of identified proteins characterizing today’s proteomic projects. Surprisingly, the development of electrophoretic methods in the last 3 years looks like a “passed-on topic”. In ten years we moved from identifying at most 500 species in complex samples to identifying 60% of the human proteome.