Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Alternative polyadenylation

This mini-protocol converts aligned RNA-seq reads into an analysis-ready alternative-polyadenylation (APA) phenotype matrix for apaQTL analysis.

Miniprotocol Timing

Timing: TBD

Overview

APA calling derives 3′UTR regions, converts transcriptome BAM files to per-base coverage, and runs DaPars2 to estimate percentage of distal polyadenylation site usage (PDUI). APA imputation and QC imputes missing PDUI values, quantile-normalizes the matrix, and can replace internal sample identifiers with analysis identifiers.

The full calling route requires transcriptome BAM files, which are not included with the example data. A precomputed chromosome-22 DaPars result is provided so that the imputation and renaming routes can be run independently.

Steps

Choose the route that matches the available starting data; the six commands are not always one mandatory chain.

Analysis goalCommands to run, in orderInputs
Call and process APA from transcriptome alignments1 → 2 → 3 → 4 → 5input/reference_data/protocol_example.apa.chr22.gtf; transcriptome BAM files under output/rnaseq/bam/ (not bundled)
Resume from per-sample coverage files3 → 4 → 5output/apa/wig/*.wig; matching depth files; output/apa/protocol_example.apa.chr22_3UTR.bed
Impute and normalize an existing chromosome-22 DaPars result5tests/fixtures/apa_impute/apa_chr22/Dapars_result_result_temp.chr22.txt
Impute, normalize and rename samples5 → 6DaPars result above; tests/fixtures/apa_impute/protocol_example.apa_matchtable.txt

Within the selected route, run the listed commands in numerical order.

1. Build the 3′UTR reference

What it does: UTR_reference extracts transcript-level 3′UTR intervals from the reference GTF for DaPars2.

2. Generate coverage and read-depth files

What it does: bam2tools converts each transcriptome BAM into a per-base WIG file and a matching read-depth summary.

3. Build the DaPars2 configuration

What it does: APAconfig creates the sample mapping and DaPars2 configuration files from the coverage directory and 3′UTR annotation.

4. Estimate chromosome-level PDUI

What it does: APAmain runs DaPars2 for chromosome 22 and writes the unprocessed PDUI result used by the QC module.

5. Impute and normalize PDUI

What it does: APAimpute imputes missing chromosome-level PDUI values and applies quantile normalization.

6. Rename samples and assemble the final matrix

What it does: APArename maps sample identifiers, combines requested chromosomes, and bgzip-indexes the final APA phenotype matrix.

Output

StepOutput filename and relative pathDescription
1output/apa/protocol_example.apa.chr22.bed; output/apa/protocol_example.apa.chr22.transcript_to_geneName.txt; output/apa/protocol_example.apa.chr22_3UTR.bedBED12 annotation, transcript-to-gene map and DaPars2 3′UTR intervals
2output/apa/*.wig; output/apa/*.depthPer-sample coverage and read-depth files
3output/apa/sample_mapping_files.txt; output/apa/sample_configuration_file.txtDaPars2 sample mapping and configuration
4tests/fixtures/apa_impute/apa_chr22/Dapars_result_result_temp.chr22.txtRaw chromosome-22 DaPars2 PDUI result
5output/apa/apa_chr22/Dapars_result_impute_chr22.bedImputed and normalized chromosome-22 PDUI matrix
6output/apa/apa_chr22/Dapars_result_impute_renamed_chr22.bed; output/apa/apa_chr22/Dapars_result_impute_renamed_chr22.bed.gz; output/apa/apa_chr22/Dapars_result_impute_renamed_chr22.bed.gz.tbi; output/apa/Dapars_allchrom_renamed.bed; output/apa/Dapars_allchrom_renamed.bed.gz; output/apa/Dapars_allchrom_renamed.bed.gz.tbiRenamed chromosome-level and combined analysis-ready APA phenotype matrices

Anticipated Results

The complete calling route produces 3′UTR intervals, per-sample coverage summaries, a DaPars2 configuration and one raw PDUI result per requested chromosome. Step 5 produces an imputed chromosome-level phenotype matrix. Step 6 is optional but recommended when the BAM identifiers differ from the identifiers used by covariates and genotypes; its combined bgzip-compressed BED file is the input for downstream apaQTL association testing.

The bundled example starts at step 5 because transcriptome BAM files are not included.

Command Interface