Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

6 Commits
 
 
 
 
 
 
 
 

Repository files navigation

Population Genetics

Handling genetic database information related to forensic DNA profiles. Each report builds up towards the culminating task: to identify the source of contamination in theoretical DTC Genetic testing company lab.

Jupyter notebooks used for visualization, core coding done in Python & R.

Case Background

A DTC genetic testing company launched has experienced a recent increase in contamination of the DNA sent in for processing from their buyers. They also have recently trained 25 new recruits who are fresh out of their undergraduate forensic science degree programs. The quality assurance/quality control (QA/QC) team has a hunch that contamination might be coming from a new hire but cannot be sure. The QA/QC team has contacted you, an expert in source attribution from biological material, to help identify the source of contamination.


Report 1 - Databases

  • Number of alleles per locus
  • Allele frequencies for 27 STR loci
  • Observed and expected heterozygosities for 27 STR loci
  • Hardy-Weinberg Equilibrium for 27 STR loci

Report 2 - PCA Data Visualization

  • Packages: tidyverse, ggplot2, factoextra, caret, randomForest
  • Sample size for each reference population
  • Mean, standard deviation, standard error, and 95% confidence interval of principal components 1-5 for each population
  • Proportion of variance explained by PCs 1-5 and the cumulative variance explained by the first 5 PCs. Please present a table and a scree plot
  • Scatter plots of PC1 vs. PC2 and PC2 vs. PC3 using the reference data only
  • Scatter plots of PC1 vs. PC2 and PC2 vs. PC3 with the unknown data projected onto the reference data
  • Confusion matrix to describe clustering of unknowns into genetically predicted ancestry groups

Report 3 - Match Statistics

  • Sample sizes for each reference population
  • Locus description for each population, number of alleles per locus, allele frequencies, observed and expected heterozygosities, Hardy-Weinberg Equilibrium
  • Correlation between observed heterozygosity and number of alleles per locus
  • Identities of the contaminators and random match statistics for each of their profiles

About

Python and R to handle genetic database information related to forensic DNA profiles.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages