File Format¶
Every HDF5 file cmuts reads and writes follows a consistent format, containing or expecting a subset of the datasets below. See an individual program’s page for the datasets it reads or writes and how it produces them.
The size of each dataset depends on the number of references n and the length of the longest reference l. Ragged libraries contain left-aligned data, with the fill value where the reference did not have any bases.
The type indicated only applies when the corresponding dataset is written, as numeric types are automatically converted during loading.
Datasets¶
mismatches/rate¶
Shape (n, l) · Type float32 · Fill NaN
The rate of mismatches at each base, over the coverage.
mismatches/error¶
Shape (n, l) · Type float32 · Fill NaN
Binomial standard error of the mismatch rate.
insertions/rate¶
Shape (n, l) · Type float32 · Fill NaN
The rate of insertions opened after each base, over the reads pairing the base and continuing 5’.
insertions/error¶
Shape (n, l) · Type float32 · Fill NaN
Binomial standard error of the insertion rate.
deletions/rate¶
Shape (n, l) · Type float32 · Fill NaN
The rate of deletion runs ending at each base, over the reads pairing the base 3’ of it and continuing 5’.
deletions/error¶
Shape (n, l) · Type float32 · Fill NaN
Binomial standard error of the deletion rate.
terminations/rate¶
Shape (n, l) · Type float32 · Fill NaN
The rate of reads whose 5’-most paired base is each base, over the coverage.
terminations/error¶
Shape (n, l) · Type float32 · Fill NaN
Binomial standard error of the termination rate.
norm¶
Shape () · Type float32 · Fill NaN
The norm every rate in this file was divided by.
coverage¶
Shape (n, l) · Type float32 · Fill 0
The number of reads in which each base was present.
sequence¶
Shape (n, l) · Type int8 · Fill -1
The reference sequence: 0 for A, 1 for C, 2 for G, 3 for T, 4 for any other base, and -1 for every column past the reference’s end.
reads/primary/counted¶
Shape (n,) · Type uint64 · Fill 0
The number of reads contributing to the rates.
reads/primary/rejected¶
Shape (n,) · Type uint64 · Fill 0
The number of reads aligned to the reference but not contributing to the rates.
reads/supplementary/counted¶
Shape (n,) · Type uint64 · Fill 0
The number of further pieces of split reads contributing to the rates.
reads/supplementary/rejected¶
Shape (n,) · Type uint64 · Fill 0
The number of further pieces of split reads aligned to the reference but not contributing to the rates.
reads/length-histogram¶
Shape (n, 2l) · Type uint64 · Fill 0
The number of reads contributing to the rates, binned by length.
reads/quality-histogram¶
Shape (n, 255) · Type uint64 · Fill 0
The number of reads contributing to the rates, binned by mapping quality.
reads/unmapped¶
Shape () · Type uint64 · Fill 0
The number of reads not aligned to any reference.
pairwise/correlation¶
Shape (n, l, l) · Type float32 · Fill NaN
The Pearson correlation of mutations between this pair of bases.
pairwise/conditional¶
Shape (n, l, l) · Type float32 · Fill NaN
The probability that the base on the first axis was mutated in a read, given that the base on the second axis was.
pairwise/coverage¶
Shape (n, l, l) · Type float32 · Fill 0
The number of reads in which each pair of bases was present.
File Attributes¶
Every file records metadata on the root group. These are attributes and not datasets, so h5py reads them from f.attrs.
Attribute |
Description |
|---|---|
|
The name of the program that produced this file. |
|
The version of |
Reading With Python¶
import h5py
# Reading attributes
with h5py.File("output.h5") as f:
print(f.attrs["program"], f.attrs["version"])
# Reading datasets
with h5py.File("output.h5") as f:
mismatches = f["mismatches/rate"][:]
reads = f["reads/primary/counted"][:]