pmd: Paired Mass Distance Analysis for GC/LC-MS Based Non-Targeted Analysis and Reactomics Analysis

CRAN status Download counter Project Status: Active - The project has reached a stable, usable state and is being actively developed.

Paired mass distance (PMD) analysis proposed in Yu, Olkowicz and Pawliszyn (2018) and PMD based reactomics proposed in Yu and Petrick (2020) for gas/liquid chromatography–mass spectrometry (GC/LC-MS) based non-targeted analysis. PMD analysis including GlobalStd algorithm and structure/reaction directed analysis. GlobalStd algorithm could found independent peaks in m/z-retention time profiles based on retention time hierarchical cluster analysis and frequency analysis of paired mass distances within retention time groups. Structure directed analysis could be used to find potential relationship among those independent peaks in different retention time groups based on frequency of paired mass distances. Reactomics analysis could also be performed to build PMD network, assign sources and make biomarker reaction discovery. GUIs for PMD analysis is also included as ‘shiny’ applications.

Installation

You can install package from this GitHub repository:

{r} remotes::install_github("yufree/pmd")

Or find a stable version from CRAN:

{r} install.packages('pmd')

Usage

To perform GlobalStd algorithm, use the following code:

{r} library(pmd) data("spmeinvivo") pmd <- getpaired(spmeinvivo) std <- getstd(pmd) To perform structure/reaction directed analysis, use the following code:

{r} sda <- getsda(std)

To perform GlobalStd algorithem along with structure/reaction directed analysis, use the following code:

{r} result <- globalstd(spmeinvivo)

To use the shiny application within the package, use the following code:

{r} runPMD()

To find homologous series (e.g., repeated addition of CH2) or specific reaction sequences:

```{r} # Find homologous series of CH2 (14.0157) with at least 3 nodes homolog_series <- gethomolog(spmeinvivo, unit = 14.0157, min_len = 3)

Find specific sequences (e.g., glycosylation followed by dehydration)

seqs <- getchainseq(spmeinvivo, c(162.0528, -18.0106))


To screen the whole feature list against a curated database of known reaction chains (e.g. desaturation-elongation, hydroxylation-glucuronidation, sequential oxidation), use the built-in chain database:

```{r}
data("pmdchain")
r <- getchainseq(spmeinvivo, db = pmdchain)
# one row per database chain: which named chains are present in this sample,
# how many matched paths were found, and the chain's `specificity` (the number
# of KEGG reaction paths sharing the chain's PMD signature: lower = rarer =
# more specific, so a rare-chain match carries more weight than a common edit)
head(r$chainsearch)

A match here is a relational annotation (co-varying, chromatographically resolved features related by the chain’s mass differences), not compound identification. Both shiny applications (runPMD() and runPMDnet()) also include a one-click “Screen for known reaction chains” panel for this database screening.

To check the pmd reaction database:

{r} # all reaction data("omics") View(omics) # kegg reaction data("keggrall") View(keggrall) # literature reaction for mass spectrometry data("sda") View(sda) # curated known reaction-chain database for screening data("pmdchain") View(pmdchain)

To check the HMDB pmd database:

{r} data("hmdb") View(hmdb)

To cite related papers:

{r} citation('pmd')