Basics

The cmuts pipeline comprises five subcommands of the cmuts binary:

cmuts align
Purpose: Aligning raw sequencing data against the reference library
Requires: One or more FASTQ files, the FASTA library

cmuts hmm
Purpose: Computing reactivity rates from alignment files via the pair-HMM
Requires: One or more coordinate-sorted alignment (SAM, BAM, or CRAM) files, the FASTA library

cmuts sub
Purpose: Background subtraction of reactivity rates
Requires: Treated and untreated reactivity rates, in cmuts-compatible HDF5 files

cmuts div
Purpose: Normalization of reactivity rates against a denatured control
Requires: Reactivity rates and denatured control rates, in cmuts-compatible HDF5 files

cmuts norm
Purpose: Normalization of reactivity values across experiments
Requires: One or more sets of reactivity rates, in cmuts-compatible HDF5 files

This page goes over basic, end-to-end usage of these programs on standard data. For a full list of the arguments each command takes, please read their respective pages.

Standard Usage

The canonical use case of the cmuts pipeline is to generate reactivity profiles from a MaP-seq experiment, where the cDNA of treated and untreated RNA has been sequenced.

The first step is to align the reads against the reference library.

cmuts align -f references.fasta -x sr -o treated.bam treated.fastq.gz
cmuts align -f references.fasta -x sr -o untreated.bam untreated.fastq.gz

-x names the platform the reads were sequenced with, and is required. See the cmuts align page for the presets it accepts and the rest of its options.

Then, pass the alignments to the HMM in order to compute reactivity rates.

cmuts hmm -f references.fasta -o treated.h5 treated.bam
cmuts hmm -f references.fasta -o untreated.h5 untreated.bam

With this done, perform background subtraction.

cmuts sub -o reactivity.h5 treated.h5 untreated.h5

The final step is normalizing the reactivity.

cmuts norm -o normalized-reactivity.h5 reactivity.h5

All HDF5 files in cmuts have the same format, where n is the number of references and l the length of the longest of them.

Dataset

Shape

Type

Fill

coverage

(n, l)

float32

zero

reactivity

(n, l)

float32

NaN

error

(n, l)

float32

NaN

reads/lengths

(n, 2l)

uint64

zero

reads/counted

(n,)

uint64

zero

reads/rejected

(n,)

uint64

zero

reads/unmapped

()

uint64

zero

norm

()

float32

NaN

See the output page for more detail on what each dataset contains.

Pre-Aligned Data

Skip running cmuts align. Ensure your data is sorted, which can be done with samtools sort. Paired-end mates must have been merged before alignment; cmuts hmm refuses a paired read, since two mates would count their overlap twice.

Split Alignments

Multiple alignment files can be passed to the HMM, where they are treated as if they were one large alignment file.

cmuts hmm -f references.fasta -o counts.h5 lane1.bam lane2.bam lane3.bam

Denatured Control

To account for a denatured control, first use the HMM to get its mutation rates alongside the treated and untreated experiments,

cmuts hmm -f references.fasta -o denatured.h5 denatured.bam

and then divide the background-subtracted rates by it.

cmuts div -o normalized.h5 combined.h5 denatured.h5

No Control

With no control, skip background subtraction and use the output of cmuts hmm directly.

Correlation-Based Data

To analyze M2-seq, RING-MaP, or MOHCA-seq data, pass --pairwise to cmuts hmm with the statistics to compute, comma separated. --pairwise correlation stores the mutation correlation of every pair of positions, and --pairwise conditional stores the probability that one position was mutated in a read where the other was.