Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Genotype preprocessing

Convert raw genotype VCFs into quality-controlled, sample-matched PLINK datasets and genetic principal components for downstream xQTL analysis.

Miniprotocol Timing

This is the total duration for all phases; module-specific timings appear on their respective pages.

Timing: ~3-5 min (on the toy dataset)

Overview

This mini-protocol walks through how raw genotype VCFs are normalized, converted to PLINK, quality-controlled, matched to molecular-phenotype samples, and prepared for population-structure adjustment. Each step calls a workflow from VCF_QC.ipynb, genotype_formatting.ipynb, GWAS_QC.ipynb, or PCA.ipynb.

The commands are organized as selectable routes rather than one mandatory 12-step chain. The standard route estimates principal components in unrelated individuals. Use the conditional extension only when related individuals will remain in the analysis and must be projected into the same PCA space before the datasets are recombined.

Steps

Choose a route before running commands; the 12 commands are not one mandatory chain.

Analysis goalCommands to run, in orderInputs
Produce a QC-passed, chromosome-partitioned PLINK dataset1 → 2 → 3 → 4tests/fixtures/vcf_qc/protocol_example.genotype.chr22.vcf.gz
input/reference_data/00-All.add_chr.variants.gz
input/reference_data/GRCh38_full_analysis_set_plus_decoy_hla.noALT_noHLA_noDecoy_ERCC.fasta
Match samples and estimate PCA in unrelated individuals1 → 2 → 3 → 4 → 5 → 6 → 7 → 8Basic-route inputs plus tests/fixtures/gene_annotation/protocol_example.rnaseq.bed.gz
Retain related individuals by PCA projection1 → 2 → 3 → 4 → 5 → 6 → 7 → 8 → 9 → 10 → 11 → 12Standard-PCA inputs plus tests/fixtures/pca/protocol_example.pca_pheno.txt

Run only the row matching the intended analysis goal, following its commands in numerical order.

1. Quality-control the input VCF

What it does: Normalize and filter the toy VCF against dbSNP and GRCh38.

What it does: Convert the output of step 1 to PLINK. The merge command generalizes to multiple chromosome-level files.

What it does: Apply genotype-, sample- and Hardy-Weinberg-equilibrium filters to the merged PLINK dataset.

4. Partition the QC-passed genotype data by chromosome

What it does: Create chromosome-specific PLINK files required by chromosome-oriented downstream workflows.

5. Match genotype and molecular-phenotype samples

What it does: Retain the sample intersection between the QC-passed genotype data and molecular phenotype.

What it does: Use KING to identify related pairs and produce related and unrelated subsets.

7. Prepare the unrelated, LD-pruned PCA subset

What it does: Apply the minor-allele-count filter and LD pruning used to estimate ancestry axes.

8. Estimate principal components in unrelated individuals

What it does: Estimate the PCA model and scores in unrelated individuals.

What it does: Restrict the related subset to the variants used by the unrelated-sample PCA model.

What it does: Project related individuals into the PCA space and identify ancestry-space outliers.

11. Remove projected PCA outliers

What it does: Remove projected outliers before recombining samples.

What it does: Merge the unrelated PCA subset with the retained projected related samples for downstream analysis.

Output

Route or stepProducts
Step 1: VCF quality controloutput/vcf_qc/*.leftnorm.vcf.gz and its index
Step 2: PLINK conversion and mergeoutput/genotype_formatting/plink/protocol_example.genotype.merged.{bed,bim,fam}
Step 3: PLINK quality controloutput/gwas_qc/plink/*.plink_qc.{bed,bim,fam}
Step 4: Chromosome partitioningoutput/genotype_by_chrom/*.{bed,bim,fam}
Step 5: Sample matchingoutput/gwas_qc/genotype/*.sample_genotypes.txt
Step 6: KinshipKING results and related/unrelated PLINK subsets under output/gwas_qc/kinship/
Steps 7–8: Unrelated-sample PCALD-pruned PLINK data under output/gwas_qc/genotype/, plus the PCA model, scores and plots under output/pca_uf/
Steps 9–11: Related-sample projectionVariant-matched related PLINK data under output/pca_related/, projected PCA scores and the outlier list under output/pca_uf/, and an outlier-filtered related subset
Step 12: Recombined analysis datasetoutput/genotype_final/protocol_example.qced.{bed,bim,fam}

Anticipated Results

The QC route produces normalized genotype data in PLINK format, including chromosome-partitioned files for downstream xQTL workflows. The standard PCA route additionally produces sample-matched related and unrelated subsets, an LD-pruned unrelated dataset, and principal-component scores estimated without close relatives.

When related individuals are retained, steps 9–12 project them using the unrelated-sample PCA loadings, remove designated projection outliers, and recombine the retained samples into the final analysis dataset.

Proceed to phenotype and covariate preprocessing before association testing.

Command Interface

List the workflows and parameters available in each module used by this mini-protocol.