Handling genetic database information related to forensic DNA profiles. Each report builds up towards the culminating task: to identify the source of contamination in theoretical DTC Genetic testing company lab.
Jupyter notebooks used for visualization, core coding done in Python & R.
A DTC genetic testing company launched has experienced a recent increase in contamination of the DNA sent in for processing from their buyers. They also have recently trained 25 new recruits who are fresh out of their undergraduate forensic science degree programs. The quality assurance/quality control (QA/QC) team has a hunch that contamination might be coming from a new hire but cannot be sure. The QA/QC team has contacted you, an expert in source attribution from biological material, to help identify the source of contamination.
- Number of alleles per locus
- Allele frequencies for 27 STR loci
- Observed and expected heterozygosities for 27 STR loci
- Hardy-Weinberg Equilibrium for 27 STR loci
- Packages: tidyverse, ggplot2, factoextra, caret, randomForest
- Sample size for each reference population
- Mean, standard deviation, standard error, and 95% confidence interval of principal components 1-5 for each population
- Proportion of variance explained by PCs 1-5 and the cumulative variance explained by the first 5 PCs. Please present a table and a scree plot
- Scatter plots of PC1 vs. PC2 and PC2 vs. PC3 using the reference data only
- Scatter plots of PC1 vs. PC2 and PC2 vs. PC3 with the unknown data projected onto the reference data
- Confusion matrix to describe clustering of unknowns into genetically predicted ancestry groups
- Sample sizes for each reference population
- Locus description for each population, number of alleles per locus, allele frequencies, observed and expected heterozygosities, Hardy-Weinberg Equilibrium
- Correlation between observed heterozygosity and number of alleles per locus
- Identities of the contaminators and random match statistics for each of their profiles