File Format

Every HDF5 file cmuts reads and writes follows a consistent format, containing or expecting a subset of the datasets below. See an individual program’s page for the datasets it reads or writes and how it produces them.

The size of each dataset depends on the number of references n and the length of the longest reference l. Ragged libraries contain left-aligned data, with the fill value where the reference did not have any bases.

The type indicated only applies when the corresponding dataset is written, as numeric types are automatically converted during loading.

Datasets

mismatches/rate

Shape (n, l) · Type float32 · Fill NaN

The rate of mismatches at each base, over the coverage.

mismatches/error

Shape (n, l) · Type float32 · Fill NaN

Binomial standard error of the mismatch rate.

insertions/rate

Shape (n, l) · Type float32 · Fill NaN

The rate of insertions opened after each base, over the reads pairing the base and continuing 5’.

insertions/error

Shape (n, l) · Type float32 · Fill NaN

Binomial standard error of the insertion rate.

deletions/rate

Shape (n, l) · Type float32 · Fill NaN

The rate of deletion runs ending at each base, over the reads pairing the base 3’ of it and continuing 5’.

deletions/error

Shape (n, l) · Type float32 · Fill NaN

Binomial standard error of the deletion rate.

terminations/rate

Shape (n, l) · Type float32 · Fill NaN

The rate of reads whose 5’-most paired base is each base, over the coverage.

terminations/error

Shape (n, l) · Type float32 · Fill NaN

Binomial standard error of the termination rate.

norm

Shape () · Type float32 · Fill NaN

The norm every rate in this file was divided by.

coverage

Shape (n, l) · Type float32 · Fill 0

The number of reads in which each base was present.

sequence

Shape (n, l) · Type int8 · Fill -1

The reference sequence: 0 for A, 1 for C, 2 for G, 3 for T, 4 for any other base, and -1 for every column past the reference’s end.

reads/primary/counted

Shape (n,) · Type uint64 · Fill 0

The number of reads contributing to the rates.

reads/primary/rejected

Shape (n,) · Type uint64 · Fill 0

The number of reads aligned to the reference but not contributing to the rates.

reads/supplementary/counted

Shape (n,) · Type uint64 · Fill 0

The number of further pieces of split reads contributing to the rates.

reads/supplementary/rejected

Shape (n,) · Type uint64 · Fill 0

The number of further pieces of split reads aligned to the reference but not contributing to the rates.

reads/length-histogram

Shape (n, 2l) · Type uint64 · Fill 0

The number of reads contributing to the rates, binned by length.

reads/quality-histogram

Shape (n, 255) · Type uint64 · Fill 0

The number of reads contributing to the rates, binned by mapping quality.

reads/unmapped

Shape () · Type uint64 · Fill 0

The number of reads not aligned to any reference.

pairwise/correlation

Shape (n, l, l) · Type float32 · Fill NaN

The Pearson correlation of mutations between this pair of bases.

pairwise/conditional

Shape (n, l, l) · Type float32 · Fill NaN

The probability that the base on the first axis was mutated in a read, given that the base on the second axis was.

pairwise/coverage

Shape (n, l, l) · Type float32 · Fill 0

The number of reads in which each pair of bases was present.

File Attributes

Every file records metadata on the root group. These are attributes and not datasets, so h5py reads them from f.attrs.

Attribute

Description

program

The name of the program that produced this file.

version

The version of cmuts that produced this file.

Reading With Python

import h5py

# Reading attributes

with h5py.File("output.h5") as f:
    print(f.attrs["program"], f.attrs["version"])

# Reading datasets

with h5py.File("output.h5") as f:
    mismatches = f["mismatches/rate"][:]
    reads = f["reads/primary/counted"][:]