Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Phenotype preprocessing

Prepare molecular-phenotype matrices for xQTL analysis by imputing missing values when needed, adding genomic annotations, and formatting data for chromosome-, region-, TAD-, or sample-specific analyses.

Miniprotocol Timing

This is the total duration for the standard toy-data route; module-specific timings appear on their respective pages.

Timing: <12 min (on the toy dataset)

Overview

This mini-protocol walks through how molecular-phenotype matrices are prepared for downstream xQTL analysis. Each step calls a workflow from phenotype_imputation.ipynb, gene_annotation.ipynb, or phenotype_formatting.ipynb.

The commands are selectable routes rather than one mandatory chain. Impute only matrices containing missing values. Choose the annotation workflow matching gene, protein, or LeafCutter identifiers, then choose the formatting workflow required by the downstream analysis unit. GCT sample extraction and BAM subsetting are independent utilities.

Steps

Choose a route before running commands; the 11 commands are not one mandatory chain.

Analysis goalCommands to run, in orderInputs
Gene-expression or protein phenotype by chromosome2 → 6tests/fixtures/gene_annotation/protocol_example.rnaseq.bed.gz; input/reference_data/Homo_sapiens.GRCh38.103.chr.reformatted.collapse_only.gene.ERCC.gtf
Phenotype with missing values by chromosome1 → 2 → 6tests/fixtures/phenotype_imputation/protocol_example.protein.missing.bed.gz; gene-coordinate GTF above
Retrieve gene coordinates from Ensembl BioMart3tests/fixtures/gene_annotation/protocol_example.rnaseq.gene_ID.tsv; internet access
LeafCutter clusters mapped to genes4tests/fixtures/gene_annotation/protocol_example.leafcutter.phenotype.bed.gz; tests/fixtures/gene_annotation/protocol_example.leafcutter.intron_count.tsv; input/reference_data/Homo_sapiens.GRCh38.103.chr.gtf
LeafCutter isoforms annotated for QTL analysis5tests/fixtures/gene_annotation/protocol_example.leafcutter.phenotype.bed.gz; tests/fixtures/gene_annotation/protocol_example.leafcutter.intron_count.tsv; input/reference_data/Homo_sapiens.GRCh38.103.chr.gtf
Partition a GCT matrix by chromosome7input/rnaseq/protocol_example.rnaseq.gene_tpm.gct.gz
Partition a BED phenotype by predefined regions2 → 8tests/fixtures/gene_annotation/protocol_example.rnaseq.bed.gz; gene-coordinate GTF above; input/reference_data/TAD/protocol_example_protein.enhanced_cis_chr22.bed
Define TAD-based phenotype regions2 → 9tests/fixtures/gene_annotation/protocol_example.rnaseq.bed.gz; gene-coordinate GTF above; tests/fixtures/generalized_TADB/expected/TADB_enhanced_cis.bed
Extract selected samples from a GCT matrix10input/rnaseq/protocol_example.rnaseq.gene_tpm.gct.gz; tests/fixtures/phenotype_formatting/keep_samples.txt
Subset BAM files to selected genomic regions11input/rnaseq/bam_file_list.txt; referenced BAM files; no example BAM is currently bundled

Run only the row matching the intended analysis goal, following its commands in numerical order. Other imputation methods are available through phenotype_imputation.ipynb and its Command Interface.

1. Impute missing phenotype values

What it does: Use generalized empirical Bayes matrix factorization to complete the example protein matrix.

2. Add genomic coordinates to gene or protein phenotypes

What it does: Join phenotype identifiers to a supplied coordinate annotation and write a coordinate-aware BED matrix.

3. Retrieve gene coordinates from Ensembl BioMart

What it does: Query the selected Ensembl release when a local coordinate annotation is unavailable.

4. Map LeafCutter clusters to genes

What it does: Assign LeafCutter clusters to genes using splice-site overlap with the gene annotation.

5. Annotate LeafCutter isoforms

What it does: Convert LeafCutter intron-cluster phenotypes into annotated isoform features for QTL analysis.

6. Partition a BED phenotype by chromosome

What it does: Split a coordinate-annotated BED phenotype into chromosome-specific files.

7. Partition a GCT phenotype by chromosome

What it does: Split a coordinate-aware GCT matrix into chromosome-specific GCT files.

8. Partition a phenotype by predefined regions

What it does: Extract phenotype features falling within each region in a supplied region list.

9. Define TAD-based phenotype regions

What it does: Assign phenotype features to TAD windows and generate a region list for downstream analysis.

10. Extract selected samples from a GCT matrix

What it does: Retain only samples listed in a supplied keep file.

11. Subset BAM files by genomic region

What it does: Extract selected chromosomes or regions from every BAM listed in the input manifest.

Output

Route or stepProducts and relative paths
Step 1output/phenotype_imputation_uf/protocol_example.protein.missing.bed.imputed.bed.gz
Step 2tests/fixtures/phenotype_formatting/protocol_example.rnaseq.bed.bed.gz; tests/fixtures/gene_annotation/expected/protocol_example.rnaseq.bed.region_list.txt
Step 3output/gene_annotation/protocol_example.rnaseq.gene_ID.bed.gz
Step 4output/gene_annotation/protocol_example.leafcutter.phenotype.bed.gz.exon_list; output/gene_annotation/protocol_example.leafcutter.phenotype.bed.gz.leafcutter.clusters_to_genes.txt
Step 5tests/fixtures/gene_annotation/expected/protocol_example.leafcutter.phenotype.bed.formated.bed.gz; output/gene_annotation/protocol_example.leafcutter.phenotype.phenotype_group.txt
Step 6output/phenotype_uf/protocol_example.genotype.chr22.bed.gz; tests/fixtures/phenotype_formatting/expected/protocol_example.phenotype_by_chrom_files.txt; its region list
Step 7output/phenotype_gct/protocol_example.rnaseq.gene_tpm.chr21.gct; output/phenotype_gct/protocol_example.rnaseq.gene_tpm.chr22.gct
Step 8output/phenotype_by_region/protocol_example_protein.enhanced_cis_chr22_phenotype_by_region/*.bed.gz; output/phenotype_by_region/*.phenotype_by_region_files.txt
Step 9output/phenotype_by_region/*_pheno_per_region.region_list; TAD-annotated region list under output/phenotype_by_region/
Step 10output/phenotype_gct/protocol_example.rnaseq.gene_tpm.sample_matched.gct.gz
Step 11output/bam_subset/*.subsetted.bam

Anticipated Results

The selected route produces a molecular-phenotype matrix with the genomic annotation and layout required by its downstream xQTL analysis. Standard gene-expression routes typically end with chromosome-specific BED files; splicing routes end with gene- or isoform-annotated phenotypes; region and TAD routes end with per-region files and their manifest.

Continue with covariate preprocessing and then the appropriate association-testing workflow.

Command Interface

List the workflows and parameters available in each module used by this mini-protocol.