Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Reference data preparation

Prepare the genome, annotation, alignment indices and optional genomic reference products used by downstream xQTL workflows.

Miniprotocol Timing

This is the total duration for all phases; module-specific timings appear on their respective pages.

Timing: TBD

Overview

This mini-protocol walks through how the reference data used throughout the xQTL Protocol are downloaded, formatted, and indexed. It is an introductory, top-to-bottom guide: each step calls a single workflow from reference_data_preparation.ipynb, with additional workflows from VCF_QC.ipynb, generalized_TADB.ipynb, and rss_ld_sketch.ipynb.

The commands are organized as selectable routes rather than one mandatory 13-step chain. Steps 1–10 prepare the core genome and the RNA-seq, RSEM, or Picard resources needed by the selected route. Steps 11–13 are independent optional workflows for dbSNP annotation, TAD-based association windows, and RSS LD sketches.

Steps

Choose a route before running commands; the 13 commands are not one mandatory chain.

Analysis goalCommands to run, in orderInputs
Basic genome and gene annotation1 → 2 → 5 → 6-
Bulk RNA-seq alignment with STAR1 → 2 → 3 → 5 → 6 → 7 → 8-
Transcript quantification with RSEM1 → 2 → 3 → 5 → 6 → 7 → 8 → 9-
Picard RNA-seq quality control1 → 2 → 3 → 5 → 6 → 7 → 8 → 10-
Add dbSNP rsIDs to a study VCF4 → 11Internet access and a study VCF
TAD-based association windows12Tissue-specific TAD calls and gene coordinates
RSS fine-mapping or TWAS13LD-block definitions and genotype VCFs

Run only the row matching the intended analysis goal, following its commands in numerical order.

1. Download the human genome

What it does: Download the GRCh38 reference FASTA used by sequence-aware tools.

2. Download the gene annotation

What it does: Download the Ensembl gene annotation used to define genes and transcripts.

3. Download the ERCC reference

What it does: Download ERCC spike-in sequences and annotations for RNA-seq reference construction.

4. Download dbSNP

What it does: Download the dbSNP variant resource used when rsID annotation is required.

5. Format and index the genome

What it does: Remove unsupported alternate sequences, append ERCC sequences and create FASTA indices.

6. Format the gene annotation

What it does: Add chromosome prefixes and construct the protocol-compatible gene models.

7. Combine gene and ERCC annotations

What it does: Combine the processed human and ERCC annotations into the GTF used downstream.

8. Build the STAR index

What it does: Build the STAR genome index used for RNA-seq alignment.

9. Build the RSEM index

What it does: Build the RSEM reference used for transcript-level expression quantification.

10. Generate RefFlat annotation

What it does: Convert the processed GTF into the RefFlat annotation used by Picard RNA-seq QC.

11. Annotate a study VCF with dbSNP identifiers (optional)

What it does: Add rsIDs to a study VCF when downstream tools require named variants.

12. Construct generalized TAD boundaries (optional)

What it does: Combine tissue-specific TAD calls into customized cis-association windows.

13. Construct an LD sketch for RSS analysis (optional)

What it does: Generate, process and merge LD-sketch products for summary-statistics regression.

Output

RouteOutput products and relative paths
Downloadsoutput/reference_data/GRCh38_full_analysis_set_plus_decoy_hla.fa; output/reference_data/Homo_sapiens.GRCh38.103.chr.gtf; output/reference_data/ERCC92.fa; tests/fixtures/reference_data_preparation/ERCC92.gtf; output/reference_data/00-All.vcf.gz; output/reference_data/00-All.vcf.gz.tbi
Basic genomeoutput/reference_data/GRCh38_full_analysis_set_plus_decoy_hla.noALT_noHLA_noDecoy.fasta; output/reference_data/GRCh38_full_analysis_set_plus_decoy_hla.noALT_noHLA_noDecoy_ERCC.fasta; output/reference_data/GRCh38_full_analysis_set_plus_decoy_hla.noALT_noHLA_noDecoy_ERCC.dict and FASTA index
Basic annotationoutput/reference_data/Homo_sapiens.GRCh38.103.chr.reformatted.gtf; output/reference_data/Homo_sapiens.GRCh38.103.chr.reformatted.collapse_only.gene.gtf; output/reference_data/Homo_sapiens.GRCh38.103.chr.reformatted.ERCC.gtf
STARoutput/reference_data/STAR_Index/
RSEMoutput/reference_data/RSEM_Index/
Picard QCoutput/reference_data/Homo_sapiens.GRCh38.103.chr.reformatted.ERCC.ref.flat
dbSNP study-VCF annotationoutput/reference_data/00-All.variants.gz lookup plus the annotated VCF under output/vcf_qc/
TAD windowstests/fixtures/generalized_TADB/expected/generalized_TAD.tsv; tests/fixtures/generalized_TADB/expected/generalized_TADB.tsv; tests/fixtures/generalized_TADB/expected/TADB_enhanced_cis.bed; tests/fixtures/generalized_TADB/expected/extended_TADB.bed
RSS LD sketchoutput/rss_ld_sketch/W_B50.rds; block-level output/rss_ld_sketch/protocol_example.<block>.dosage.gz; chromosome-level output/rss_ld_sketch/protocol_example.22.pgen products

Anticipated Results

After the selected route among steps 1–10, the reference directory contains the formatted genome and the annotations required by that analysis. The basic route produces the processed genome and gene annotation; the RNA-seq routes additionally produce the combined ERCC annotation and, as selected, the STAR index, RSEM index, or RefFlat annotation.

Optional step 11 produces a study VCF annotated with dbSNP rsIDs. Optional step 12 produces generalized TAD and TAD-boundary files such as generalized_TAD.tsv, generalized_TADB.tsv, TADB_enhanced_cis.bed, and extended_TADB.bed. Optional step 13 produces the random-projection matrix and the block- and chromosome-level LD-sketch products consumed by RSS workflows.

Proceed to the molecular-phenotype or genotype-preprocessing mini-protocol required by the selected analysis route.

Command Interface

List the workflows and parameters available in the reference data preparation module.