Mendelian violations in the CEU and YRI Pilot 2 Trios

Size: px
Start display at page:

Download "Mendelian violations in the CEU and YRI Pilot 2 Trios"

Transcription

1 Mendelian violations in the CEU and YRI Pilot 2 Trios Mark DePristo and Mark Daly Manager, Medical and Population Genetics Analysis Medical and Population Genetics Program Broad Institute of Harvard and MIT March 8, 2010

2 Summary Multi-sample (but trio-unaware) calls made with GATK Unified Genotyper for each Pilot 2 trio Part of the official release intersected with UMich calls This presentation focuses on Mendelian violations in general Next presentation will analyze highly-validated classic de novo mutations (novel het where parents are hom-ref) Violations fall into three distinct classes described in more detail here: True de novo mutations Cryptic copy number variation Many likely due to cell line degeneration but many segregating too FP due to sequencing artifacts

3 The Broad Unified Genotyper SNP caller was applied to all three wings of the project Sample-associated reads Individual 1 Genotype likelihoods Allele frequency Individual 2 Joint estimate across samples SNPs Individual N Genotype frequencies This approach allows us to combine weak single sample calls to discover variation among samples with high confidence See h%p:// for more informa3on

4 Pilot 2 yielded ~2.7B genotyped sites, discovering ~4-5M variants in each trio CEU trio YRI trio All variants SNPs 3.6M dbsnp % 89% Ti/Tv 2.07 All variants SNPs 4.5M dbsnp % 77% Ti/Tv 2.09 Novel variants 409K SNPs Ti/Tv = 2.04 Novel variants 1.05M SNPs Ti/Tv = 2.08

5 Mendelian violations identify germ-line and somatic de novo mutations Algorithm to identify putative Mendelian violations Autosomal Mendel violations in the CEU and CEU trios 1 De novo SNP or 2 somatic parent CNV Genotyping error or somatic parent CNV Genotype (A is ref, B is alt, * is either) CEU Trio YRI Trio A/A A/A B/B B/B Child Parents A/B A/B 1 2 A/B A/A and A/A 1, A/B B/B and B/B 1, Segregating CNV* A/A */* B/B */* 3 3 A/A B/B and */* B/B A/A and */* Total 3,769 2,174 B/B *Equivalent violations for mom A/A Initial validation suggests high validation rate for the de novo SNPs, but 80+% are non-inherited somatic events

6 Mendelian violations by known/novel Mendelian by genotype (A is ref, B is alt, * is either) CEU YRI Known Novel Known Novel Child genotype Parent genotypes A/B A/A and A/A A/B B/B and B/B A/A B/B and */* B/B A/A and */* Total Consistent with classic de novo SNP Autosomal Mendelian viola3ons where GQ >Q30 for all three individuals

7 Mom Dad No evidence in parents 454 Child SLX Consistent in all three technologies SOLid Work of Andrew Kernytsky Validated as a true de novo mutation

8 Many of the lost homozygous variant alleles in the parents are clustered For example, the 10 sites in 8Kb at 1: Consistent with somatic deletions in the parent (in this case Mom) Overlaps CNV discovered in pilot 1

9 Log10 of inter- SNP distance Clustering of violations on YRI chr1 Colored clustered have distance < 10Kb Cluster of violations marks 10 Mb deletion of chromosome arm Rare, randomly occurring de novo SNPs Inter-SNP distance for regular variation Distribu3on from all SNPs Histogram of log10 of inter-snp distances

10 Genome-wide perspective for YRI Top level junctions are among chromosomes CNVs are common and clear Signature of CNVs

11 Opposite homozygous SNPs are cluster on the chromosome whereas hets are not Opposite homozygote violations are the most clustered Consistent with cryptic copy number variation Few clusters of violations where the child is heterozygous Almost no SNP clusters Kid is heterozygous, parents are homozygous reference Mostly clustered No SNP clusters Kid and one parent are opposite homozygotes Kid is heterozygous, parents are homozygous variant

12 CEU cell line are apparently degrading with large somatic deletions throughout the genome Opposite homozygote Mendelian violations YRI trio CEU trio YRI samples prepared more recently, fewer cell line passages CEU samples prepared less recently, more cell line passages Older, more passaged CEU samples exhibit more and larger apparently somatic cell line artifacts

13 Many Mendelian violations overlap CNVs Overlaps Pilot1-only CNV Overlaps Pilot CNV Overlaps Pilot2-only CNV Not in CNV Overlaps Pilot1-only CNV Overlaps Pilot CNV Overlaps Pilot2-only CNV Not in CNV OppositeHoms CNV false negatives? Chromosome KidHet_ParentsHomRef 0.6 Freq Freq 0.8 Real de novo SNPs, mostly somatic False negative in pilot 1 calls? Likely all real CNVs Somatic CNVs in parents Violation index on chromosome KidHet_ParentsHomRef OppositeHoms CNVs from validated list of deletions in mastervalidationtables_pilot1_pilot2_

14 Violations where child is A/B and parents are B/B are homopolymer-induced genotyping errors 1200 in CEU and 400 in YRI of which ~95% are in dbsnp, suggesting they are real variant sites Consistent with somatic deletions, but do not cluster Visual inspection shows that these are true homozygous non-reference sites All neighboring long homopolymer runs Genotyping failure due to misaligned 454 and miscalled SOLiD reads SOLiD recalibration partially fixes reference bias in this area Difficult-to-fix failure mode for 454; even fixable?

15 Example

16 How reliable are NA12878 chr1 calls? Comparison to Complete Genomics sequencing for NA12878 CEU Chromosome 1 violations only Mendelian viola@ons by genotype (A is ref, B is alt, * is either) CEU % in CG Known Novel Known Novel Child genotype Parent genotypes A/B A/A and A/A A/B B/B and B/B Child and parent are opposite homozygotes NA NA Total CG: 60x genome-wide Complete Genomics sequence, called with CG analysis software

17 Conclusions Violations fall into three distinct classes described in more detail here: True de novo mutations which look excellent, also observed in CG data Cryptic copy number variation Many opposite homozygotes overlap with CNVs Class of FP sequencing artifacts Mendelian violations in trios May provide false negatives for SV calls evaluation Some CNV calls may be be somatic not segregating One value of trio sequencing is to differentiate between segregating and somatic variation

18 Appendix

19 Example 1

20 Mendelian violations consistent with cryptic copy number variation and rare de novo mutations Most clustered violations are from opposite homozygotes

Text file One header line meta information lines One line : variant/position

Text file One header line meta information lines One line : variant/position Software Calling: GATK SAMTOOLS mpileup Varscan SOAP VCF format Text file One header line meta information lines One line : variant/position ##fileformat=vcfv4.1! ##filedate=20090805! ##source=myimputationprogramv3.1!

More information

SeattleSNPs Interactive Tutorial: Web Tools for Site Selection, Linkage Disequilibrium and Haplotype Analysis

SeattleSNPs Interactive Tutorial: Web Tools for Site Selection, Linkage Disequilibrium and Haplotype Analysis SeattleSNPs Interactive Tutorial: Web Tools for Site Selection, Linkage Disequilibrium and Haplotype Analysis Goal: This tutorial introduces several websites and tools useful for determining linkage disequilibrium

More information

Paternity Testing. Chapter 23

Paternity Testing. Chapter 23 Paternity Testing Chapter 23 Kinship and Paternity DNA analysis can also be used for: Kinship testing determining whether individuals are related Paternity testing determining the father of a child Missing

More information

Disease gene identification with exome sequencing

Disease gene identification with exome sequencing Disease gene identification with exome sequencing Christian Gilissen Dept. of Human Genetics Radboud University Nijmegen Medical Centre c.gilissen@antrg.umcn.nl Contents Infrastructure Exome sequencing

More information

t e c h n i c a l r e p o r t s

t e c h n i c a l r e p o r t s t e c h n i c a l r e p o r t s A framework for variation discovery and genotyping using next-generation DNA sequencing data 2 Nature America, Inc. All rights reserved. Mark A DePristo, Eric Banks, Ryan

More information

A map of human genome variation from population-scale sequencing

A map of human genome variation from population-scale sequencing doi:1.138/nature9534 A map of human genome variation from population-scale sequencing The 1 Genomes Project Consortium* The 1 Genomes Project aims to provide a deep characterization of human genome sequence

More information

Globally, about 9.7% of cancers in men are prostate cancers, and the risk of developing the

Globally, about 9.7% of cancers in men are prostate cancers, and the risk of developing the Chapter 5 Analysis of Prostate Cancer Association Study Data 5.1 Risk factors for Prostate Cancer Globally, about 9.7% of cancers in men are prostate cancers, and the risk of developing the disease has

More information

Mendelian inheritance and the

Mendelian inheritance and the Mendelian inheritance and the most common genetic diseases Cornelia Schubert, MD, University of Goettingen, Dept. Human Genetics EUPRIM-Net course Genetics, Immunology and Breeding Mangement German Primate

More information

DNA-Analytik III. Genetische Variabilität

DNA-Analytik III. Genetische Variabilität DNA-Analytik III Genetische Variabilität Genetische Variabilität Lexikon Scherer et al. Nat Genet Suppl 39:s7 (2007) Genetische Variabilität Sequenzvariation Mutationen (Mikro~) Basensubstitution Insertion

More information

Heredity. Sarah crosses a homozygous white flower and a homozygous purple flower. The cross results in all purple flowers.

Heredity. Sarah crosses a homozygous white flower and a homozygous purple flower. The cross results in all purple flowers. Heredity 1. Sarah is doing an experiment on pea plants. She is studying the color of the pea plants. Sarah has noticed that many pea plants have purple flowers and many have white flowers. Sarah crosses

More information

Agilent CytoGenomics Software A Complete Solution for Cytogenetic Research Data Analysis

Agilent CytoGenomics Software A Complete Solution for Cytogenetic Research Data Analysis Agilent CytoGenomics Software A Complete Solution for Cytogenetic Research Data Analysis Technical Overview Streamlines the cytogenetic research workflow for finding CNCs, LOH, and UPD Enables manual sample

More information

(1-p) 2. p(1-p) From the table, frequency of DpyUnc = ¼ (p^2) = #DpyUnc = p^2 = 0.0004 ¼(1-p)^2 + ½(1-p)p + ¼(p^2) #Dpy + #DpyUnc

(1-p) 2. p(1-p) From the table, frequency of DpyUnc = ¼ (p^2) = #DpyUnc = p^2 = 0.0004 ¼(1-p)^2 + ½(1-p)p + ¼(p^2) #Dpy + #DpyUnc Advanced genetics Kornfeld problem set_key 1A (5 points) Brenner employed 2-factor and 3-factor crosses with the mutants isolated from his screen, and visually assayed for recombination events between

More information

ASSIsT: An Automatic SNP ScorIng Tool for in and out-breeding species Reference Manual

ASSIsT: An Automatic SNP ScorIng Tool for in and out-breeding species Reference Manual ASSIsT: An Automatic SNP ScorIng Tool for in and out-breeding species Reference Manual Di Guardo M, Micheletti D, Bianco L, Koehorst-van Putten HJJ, Longhi S, Costa F, Aranzana MJ, Velasco R, Arús P, Troggio

More information

Name: Class: Date: ID: A

Name: Class: Date: ID: A Name: Class: _ Date: _ Meiosis Quiz 1. (1 point) A kidney cell is an example of which type of cell? a. sex cell b. germ cell c. somatic cell d. haploid cell 2. (1 point) How many chromosomes are in a human

More information

Genotyping and quality control of UK Biobank, a large- scale, extensively phenotyped prospective resource

Genotyping and quality control of UK Biobank, a large- scale, extensively phenotyped prospective resource Genotyping and quality control of UK Biobank, a large- scale, extensively phenotyped prospective resource Information for researchers Interim Data Release, 2015 1 Introduction... 3 1.1 UK Biobank... 3

More information

Accelerating variant calling

Accelerating variant calling Accelerating variant calling Mauricio Carneiro GSA Broad Institute Intel Genomic Sequencing Pipeline Workshop Mount Sinai 12/10/2013 This is the work of many Genome sequencing and analysis team Mark DePristo

More information

Answer Key Problem Set 5

Answer Key Problem Set 5 7.03 Fall 2003 1 of 6 1. a) Genetic properties of gln2- and gln 3-: Answer Key Problem Set 5 Both are uninducible, as they give decreased glutamine synthetase (GS) activity. Both are recessive, as mating

More information

Chapter 8: Recombinant DNA 2002 by W. H. Freeman and Company Chapter 8: Recombinant DNA 2002 by W. H. Freeman and Company

Chapter 8: Recombinant DNA 2002 by W. H. Freeman and Company Chapter 8: Recombinant DNA 2002 by W. H. Freeman and Company Genetic engineering: humans Gene replacement therapy or gene therapy Many technical and ethical issues implications for gene pool for germ-line gene therapy what traits constitute disease rather than just

More information

Leading Genomics. Diagnostic. Discove. Collab. harma. Shanghai Cambridge, MA Reykjavik

Leading Genomics. Diagnostic. Discove. Collab. harma. Shanghai Cambridge, MA Reykjavik Leading Genomics Diagnostic harma Discove Collab Shanghai Cambridge, MA Reykjavik Global leadership for using the genome to create better medicine WuXi NextCODE provides a uniquely proven and integrated

More information

Delivering the power of the world s most successful genomics platform

Delivering the power of the world s most successful genomics platform Delivering the power of the world s most successful genomics platform NextCODE Health is bringing the full power of the world s largest and most successful genomics platform to everyday clinical care NextCODE

More information

DNA Copy Number and Loss of Heterozygosity Analysis Algorithms

DNA Copy Number and Loss of Heterozygosity Analysis Algorithms DNA Copy Number and Loss of Heterozygosity Analysis Algorithms Detection of copy-number variants and chromosomal aberrations in GenomeStudio software. Introduction Illumina has developed several algorithms

More information

Data Analysis for Ion Torrent Sequencing

Data Analysis for Ion Torrent Sequencing IFU022 v140202 Research Use Only Instructions For Use Part III Data Analysis for Ion Torrent Sequencing MANUFACTURER: Multiplicom N.V. Galileilaan 18 2845 Niel Belgium Revision date: August 21, 2014 Page

More information

Focusing on results not data comprehensive data analysis for targeted next generation sequencing

Focusing on results not data comprehensive data analysis for targeted next generation sequencing Focusing on results not data comprehensive data analysis for targeted next generation sequencing Daniel Swan, Jolyon Holdstock, Angela Matchan, Richard Stark, John Shovelton, Duarte Mohla and Simon Hughes

More information

Simplifying Data Interpretation with Nexus Copy Number

Simplifying Data Interpretation with Nexus Copy Number Simplifying Data Interpretation with Nexus Copy Number A WHITE PAPER FROM BIODISCOVERY, INC. Rapid technological advancements, such as high-density acgh and SNP arrays as well as next-generation sequencing

More information

CHROMOSOMES AND INHERITANCE

CHROMOSOMES AND INHERITANCE SECTION 12-1 REVIEW CHROMOSOMES AND INHERITANCE VOCABULARY REVIEW Distinguish between the terms in each of the following pairs of terms. 1. sex chromosome, autosome 2. germ-cell mutation, somatic-cell

More information

PRACTICE PROBLEMS - PEDIGREES AND PROBABILITIES

PRACTICE PROBLEMS - PEDIGREES AND PROBABILITIES PRACTICE PROBLEMS - PEDIGREES AND PROBABILITIES 1. Margaret has just learned that she has adult polycystic kidney disease. Her mother also has the disease, as did her maternal grandfather and his younger

More information

Step-by-Step Guide to Bi-Parental Linkage Mapping WHITE PAPER

Step-by-Step Guide to Bi-Parental Linkage Mapping WHITE PAPER Step-by-Step Guide to Bi-Parental Linkage Mapping WHITE PAPER JMP Genomics Step-by-Step Guide to Bi-Parental Linkage Mapping Introduction JMP Genomics offers several tools for the creation of linkage maps

More information

IGV Hands-on Exercise: UI basics and data integration

IGV Hands-on Exercise: UI basics and data integration IGV Hands-on Exercise: UI basics and data integration Verhaak, R.G. et al. Integrated Genomic Analysis Identifies Clinically Relevant Subtypes of Glioblastoma Characterized by Abnormalities in PDGFRA,

More information

SNPbrowser Software v3.5

SNPbrowser Software v3.5 Product Bulletin SNP Genotyping SNPbrowser Software v3.5 A Free Software Tool for the Knowledge-Driven Selection of SNP Genotyping Assays Easily visualize SNPs integrated with a physical map, linkage disequilibrium

More information

Single-Cell Whole Genome Sequencing on the C1 System: a Performance Evaluation

Single-Cell Whole Genome Sequencing on the C1 System: a Performance Evaluation PN 100-9879 A1 TECHNICAL NOTE Single-Cell Whole Genome Sequencing on the C1 System: a Performance Evaluation Introduction Cancer is a dynamic evolutionary process of which intratumor genetic and phenotypic

More information

Towards Integrating the Detection of Genetic Variants into an In-Memory Database

Towards Integrating the Detection of Genetic Variants into an In-Memory Database Towards Integrating the Detection of Genetic Variants into an 2nd International Workshop on Big Data in Bioinformatics and Healthcare Oct 27, 2014 Motivation Genome Data Analysis Process DNA Sample Base

More information

Breast cancer and the role of low penetrance alleles: a focus on ATM gene

Breast cancer and the role of low penetrance alleles: a focus on ATM gene Modena 18-19 novembre 2010 Breast cancer and the role of low penetrance alleles: a focus on ATM gene Dr. Laura La Paglia Breast Cancer genetic Other BC susceptibility genes TP53 PTEN STK11 CHEK2 BRCA1

More information

SeqScape Software Version 2.5 Comprehensive Analysis Solution for Resequencing Applications

SeqScape Software Version 2.5 Comprehensive Analysis Solution for Resequencing Applications Product Bulletin Sequencing Software SeqScape Software Version 2.5 Comprehensive Analysis Solution for Resequencing Applications Comprehensive reference sequence handling Helps interpret the role of each

More information

GWAS Data Cleaning. GENEVA Coordinating Center Department of Biostatistics University of Washington. January 13, 2016.

GWAS Data Cleaning. GENEVA Coordinating Center Department of Biostatistics University of Washington. January 13, 2016. GWAS Data Cleaning GENEVA Coordinating Center Department of Biostatistics University of Washington January 13, 2016 Contents 1 Overview 2 2 Preparing Data 3 2.1 Data formats used in GWASTools............................

More information

Evolution (18%) 11 Items Sample Test Prep Questions

Evolution (18%) 11 Items Sample Test Prep Questions Evolution (18%) 11 Items Sample Test Prep Questions Grade 7 (Evolution) 3.a Students know both genetic variation and environmental factors are causes of evolution and diversity of organisms. (pg. 109 Science

More information

Roberto Ciccone, Orsetta Zuffardi Università di Pavia

Roberto Ciccone, Orsetta Zuffardi Università di Pavia Roberto Ciccone, Orsetta Zuffardi Università di Pavia XIII Corso di Formazione Malformazioni Congenite dalla Diagnosi Prenatale alla Terapia Postnatale unipv.eu Carrara, 24 ottobre 2014 Legend:Bluebars

More information

Wissenschaftliche Highlights der GSF 2007

Wissenschaftliche Highlights der GSF 2007 H Forschungszentrum für Umwelt und Gesundheit GmbH in der Helmholtzgemeinschaft Wissenschaftlich-Technische Abteilung Wissenschaftliche Highlights der GSF 2007 Abfrage Oktober 2007 Institut / Selbst. Abteilung

More information

Basics of Marker Assisted Selection

Basics of Marker Assisted Selection asics of Marker ssisted Selection Chapter 15 asics of Marker ssisted Selection Julius van der Werf, Department of nimal Science rian Kinghorn, Twynam Chair of nimal reeding Technologies University of New

More information

5 GENETIC LINKAGE AND MAPPING

5 GENETIC LINKAGE AND MAPPING 5 GENETIC LINKAGE AND MAPPING 5.1 Genetic Linkage So far, we have considered traits that are affected by one or two genes, and if there are two genes, we have assumed that they assort independently. However,

More information

Overview One of the promises of studies of human genetic variation is to learn about human history and also to learn about natural selection.

Overview One of the promises of studies of human genetic variation is to learn about human history and also to learn about natural selection. Technical design document for a SNP array that is optimized for population genetics Yontao Lu, Nick Patterson, Yiping Zhan, Swapan Mallick and David Reich Overview One of the promises of studies of human

More information

Genomes and SNPs in Malaria and Sickle Cell Anemia

Genomes and SNPs in Malaria and Sickle Cell Anemia Genomes and SNPs in Malaria and Sickle Cell Anemia Introduction to Genome Browsing with Ensembl Ensembl The vast amount of information in biological databases today demands a way of organising and accessing

More information

Mendelian and Non-Mendelian Heredity Grade Ten

Mendelian and Non-Mendelian Heredity Grade Ten Ohio Standards Connection: Life Sciences Benchmark C Explain the genetic mechanisms and molecular basis of inheritance. Indicator 6 Explain that a unit of hereditary information is called a gene, and genes

More information

Chapter 9 Patterns of Inheritance

Chapter 9 Patterns of Inheritance Bio 100 Patterns of Inheritance 1 Chapter 9 Patterns of Inheritance Modern genetics began with Gregor Mendel s quantitative experiments with pea plants History of Heredity Blending theory of heredity -

More information

An example of bioinformatics application on plant breeding projects in Rijk Zwaan

An example of bioinformatics application on plant breeding projects in Rijk Zwaan An example of bioinformatics application on plant breeding projects in Rijk Zwaan Xiangyu Rao 17-08-2012 Introduction of RZ Rijk Zwaan is active worldwide as a vegetable breeding company that focuses on

More information

DRAGON GENETICS LAB -- Principles of Mendelian Genetics

DRAGON GENETICS LAB -- Principles of Mendelian Genetics DragonGeneticsProtocol Mendelian Genetics lab Student.doc DRAGON GENETICS LAB -- Principles of Mendelian Genetics Dr. Pamela Esprivalo Harrell, University of North Texas, developed an earlier version of

More information

Commonly Used STR Markers

Commonly Used STR Markers Commonly Used STR Markers Repeats Satellites 100 to 1000 bases repeated Minisatellites VNTR variable number tandem repeat 10 to 100 bases repeated Microsatellites STR short tandem repeat 2 to 6 bases repeated

More information

Report. A Note on Exact Tests of Hardy-Weinberg Equilibrium. Janis E. Wigginton, 1 David J. Cutler, 2 and Gonçalo R. Abecasis 1

Report. A Note on Exact Tests of Hardy-Weinberg Equilibrium. Janis E. Wigginton, 1 David J. Cutler, 2 and Gonçalo R. Abecasis 1 Am. J. Hum. Genet. 76:887 883, 2005 Report A Note on Exact Tests of Hardy-Weinberg Equilibrium Janis E. Wigginton, 1 David J. Cutler, 2 and Gonçalo R. Abecasis 1 1 Center for Statistical Genetics, Department

More information

Chromosomes, Mapping, and the Meiosis Inheritance Connection

Chromosomes, Mapping, and the Meiosis Inheritance Connection Chromosomes, Mapping, and the Meiosis Inheritance Connection Carl Correns 1900 Chapter 13 First suggests central role for chromosomes Rediscovery of Mendel s work Walter Sutton 1902 Chromosomal theory

More information

7A The Origin of Modern Genetics

7A The Origin of Modern Genetics Life Science Chapter 7 Genetics of Organisms 7A The Origin of Modern Genetics Genetics the study of inheritance (the study of how traits are inherited through the interactions of alleles) Heredity: the

More information

Lecture 6: Single nucleotide polymorphisms (SNPs) and Restriction Fragment Length Polymorphisms (RFLPs)

Lecture 6: Single nucleotide polymorphisms (SNPs) and Restriction Fragment Length Polymorphisms (RFLPs) Lecture 6: Single nucleotide polymorphisms (SNPs) and Restriction Fragment Length Polymorphisms (RFLPs) Single nucleotide polymorphisms or SNPs (pronounced "snips") are DNA sequence variations that occur

More information

Genetic Mutations Cause Many Birth Defects:

Genetic Mutations Cause Many Birth Defects: Genetic Mutations Cause Many Birth Defects: What We Learned from the FORGE Canada Project Jan M. Friedman, MD, PhD University it of British Columbia Vancouver, Canada I have no conflicts of interest related

More information

Forensic Statistics. From the ground up. 15 th International Symposium on Human Identification

Forensic Statistics. From the ground up. 15 th International Symposium on Human Identification Forensic Statistics 15 th International Symposium on Human Identification From the ground up UNTHSC John V. Planz, Ph.D. UNT Health Science Center at Fort Worth Why so much attention to statistics? Exclusions

More information

SeqArray: an R/Bioconductor Package for Big Data Management of Genome-Wide Sequencing Variants

SeqArray: an R/Bioconductor Package for Big Data Management of Genome-Wide Sequencing Variants SeqArray: an R/Bioconductor Package for Big Data Management of Genome-Wide Sequencing Variants Xiuwen Zheng Department of Biostatistics University of Washington Seattle Introduction Thousands of gigabyte

More information

Genetics Module B, Anchor 3

Genetics Module B, Anchor 3 Genetics Module B, Anchor 3 Key Concepts: - An individual s characteristics are determines by factors that are passed from one parental generation to the next. - During gamete formation, the alleles for

More information

How-To: SNP and INDEL detection

How-To: SNP and INDEL detection How-To: SNP and INDEL detection April 23, 2014 Lumenogix NGS SNP and INDEL detection Mutation Analysis Identifying known, and discovering novel genomic mutations, has been one of the most popular applications

More information

Computational Requirements

Computational Requirements Workshop on Establishing a Central Resource of Data from Genome Sequencing Projects Computational Requirements Steve Sherry, Lisa Brooks, Paul Flicek, Anton Nekrutenko, Kenna Shaw, Heidi Sofia High-density

More information

2. True or False? The sequence of nucleotides in the human genome is 90.9% identical from one person to the next. False (it s 99.

2. True or False? The sequence of nucleotides in the human genome is 90.9% identical from one person to the next. False (it s 99. 1. True or False? A typical chromosome can contain several hundred to several thousand genes, arranged in linear order along the DNA molecule present in the chromosome. True 2. True or False? The sequence

More information

Biology Behind the Crime Scene Week 4: Lab #4 Genetics Exercise (Meiosis) and RFLP Analysis of DNA

Biology Behind the Crime Scene Week 4: Lab #4 Genetics Exercise (Meiosis) and RFLP Analysis of DNA Page 1 of 5 Biology Behind the Crime Scene Week 4: Lab #4 Genetics Exercise (Meiosis) and RFLP Analysis of DNA Genetics Exercise: Understanding how meiosis affects genetic inheritance and DNA patterns

More information

Combining Data from Different Genotyping Platforms. Gonçalo Abecasis Center for Statistical Genetics University of Michigan

Combining Data from Different Genotyping Platforms. Gonçalo Abecasis Center for Statistical Genetics University of Michigan Combining Data from Different Genotyping Platforms Gonçalo Abecasis Center for Statistical Genetics University of Michigan The Challenge Detecting small effects requires very large sample sizes Combined

More information

How To Find Rare Variants In The Human Genome

How To Find Rare Variants In The Human Genome UNIVERSITÀ DEGLI STUDI DI SASSARI Scuola di Dottorato in Scienze Biomediche XXV CICLO DOTTORATO DI RICERCA IN SCIENZE BIOMEDICHE INDIRIZZO DI GENETICA MEDICA, MALATTIE METABOLICHE E NUTRIGENOMICA Direttore:

More information

Integration of genomic data into electronic health records

Integration of genomic data into electronic health records Integration of genomic data into electronic health records Daniel Masys, MD Affiliate Professor Biomedical & Health Informatics University of Washington, Seattle Major portion of today s lecture is based

More information

Drosophila Genetics by Michael Socolich May, 2003

Drosophila Genetics by Michael Socolich May, 2003 Drosophila Genetics by Michael Socolich May, 2003 I. General Information and Fly Husbandry II. Nomenclature III. Genetic Tools Available to the Fly Geneticists IV. Example Crosses V. P-element Transformation

More information

Consistent Assay Performance Across Universal Arrays and Scanners

Consistent Assay Performance Across Universal Arrays and Scanners Technical Note: Illumina Systems and Software Consistent Assay Performance Across Universal Arrays and Scanners There are multiple Universal Array and scanner options for running Illumina DASL and GoldenGate

More information

GENETIC CROSSES. Monohybrid Crosses

GENETIC CROSSES. Monohybrid Crosses GENETIC CROSSES Monohybrid Crosses Objectives Explain the difference between genotype and phenotype Explain the difference between homozygous and heterozygous Explain how probability is used to predict

More information

The correct answer is c A. Answer a is incorrect. The white-eye gene must be recessive since heterozygous females have red eyes.

The correct answer is c A. Answer a is incorrect. The white-eye gene must be recessive since heterozygous females have red eyes. 1. Why is the white-eye phenotype always observed in males carrying the white-eye allele? a. Because the trait is dominant b. Because the trait is recessive c. Because the allele is located on the X chromosome

More information

Overview of Genetic Testing and Screening

Overview of Genetic Testing and Screening Integrating Genetics into Your Practice Webinar Series Overview of Genetic Testing and Screening Genetic testing is an important tool in the screening and diagnosis of many conditions. New technology is

More information

Hardy-Weinberg Equilibrium Problems

Hardy-Weinberg Equilibrium Problems Hardy-Weinberg Equilibrium Problems 1. The frequency of two alleles in a gene pool is 0.19 (A) and 0.81(a). Assume that the population is in Hardy-Weinberg equilibrium. (a) Calculate the percentage of

More information

Milk protein genetic variation in Butana cattle

Milk protein genetic variation in Butana cattle Milk protein genetic variation in Butana cattle Ammar Said Ahmed Züchtungsbiologie und molekulare Genetik, Humboldt Universität zu Berlin, Invalidenstraβe 42, 10115 Berlin, Deutschland 1 Outline Background

More information

SeqArray: an R/Bioconductor Package for Big Data Management of Genome-Wide Sequence Variants

SeqArray: an R/Bioconductor Package for Big Data Management of Genome-Wide Sequence Variants SeqArray: an R/Bioconductor Package for Big Data Management of Genome-Wide Sequence Variants 1 Dr. Xiuwen Zheng Department of Biostatistics University of Washington Seattle Introduction Thousands of gigabyte

More information

Lecture 3: Mutations

Lecture 3: Mutations Lecture 3: Mutations Recall that the flow of information within a cell involves the transcription of DNA to mrna and the translation of mrna to protein. Recall also, that the flow of information between

More information

A trait is a variation of a particular character (e.g. color, height). Traits are passed from parents to offspring through genes.

A trait is a variation of a particular character (e.g. color, height). Traits are passed from parents to offspring through genes. 1 Biology Chapter 10 Study Guide Trait A trait is a variation of a particular character (e.g. color, height). Traits are passed from parents to offspring through genes. Genes Genes are located on chromosomes

More information

Single Nucleotide Polymorphisms (SNPs)

Single Nucleotide Polymorphisms (SNPs) Single Nucleotide Polymorphisms (SNPs) Additional Markers 13 core STR loci Obtain further information from additional markers: Y STRs Separating male samples Mitochondrial DNA Working with extremely degraded

More information

A guide to the analysis of KASP genotyping data using cluster plots

A guide to the analysis of KASP genotyping data using cluster plots extraction sequencing genotyping extraction sequencing genotyping extraction sequencing genotyping extraction sequencing A guide to the analysis of KASP genotyping data using cluster plots Contents of

More information

CCR Biology - Chapter 7 Practice Test - Summer 2012

CCR Biology - Chapter 7 Practice Test - Summer 2012 Name: Class: Date: CCR Biology - Chapter 7 Practice Test - Summer 2012 Multiple Choice Identify the choice that best completes the statement or answers the question. 1. A person who has a disorder caused

More information

Comment on Widespread RNA and DNA sequence differences in the human transcriptome

Comment on Widespread RNA and DNA sequence differences in the human transcriptome omment on Widespread RN and DN sequence differences in the human transcriptome Joseph K. Pickrell 1, Yoav ilad 1, Jonathan K. Pritchard 1,2 1 Department of Human enetics and 2 Howard Hughes Medical Institute

More information

somatic cell egg genotype gamete polar body phenotype homologous chromosome trait dominant autosome genetics recessive

somatic cell egg genotype gamete polar body phenotype homologous chromosome trait dominant autosome genetics recessive CHAPTER 6 MEIOSIS AND MENDEL Vocabulary Practice somatic cell egg genotype gamete polar body phenotype homologous chromosome trait dominant autosome genetics recessive CHAPTER 6 Meiosis and Mendel sex

More information

Chapter 3. Chapter Outline. Chapter Outline 9/11/10. Heredity and Evolu4on

Chapter 3. Chapter Outline. Chapter Outline 9/11/10. Heredity and Evolu4on Chapter 3 Heredity and Evolu4on Chapter Outline The Cell DNA Structure and Function Cell Division: Mitosis and Meiosis The Genetic Principles Discovered by Mendel Mendelian Inheritance in Humans Misconceptions

More information

Got Lactase? The Co-evolution of Genes and Culture

Got Lactase? The Co-evolution of Genes and Culture The Making of the Fittest: Natural The Making Selection of the and Fittest: Adaptation Natural Selection and Adaptation OVERVIEW PEDIGREES AND THE INHERITANCE OF LACTOSE INTOLERANCE This activity serves

More information

School of Nursing. Presented by Yvette Conley, PhD

School of Nursing. Presented by Yvette Conley, PhD Presented by Yvette Conley, PhD What we will cover during this webcast: Briefly discuss the approaches introduced in the paper: Genome Sequencing Genome Wide Association Studies Epigenomics Gene Expression

More information

CHAPTER 10 BLOOD GROUPS: ABO AND Rh

CHAPTER 10 BLOOD GROUPS: ABO AND Rh CHAPTER 10 BLOOD GROUPS: ABO AND Rh The success of human blood transfusions requires compatibility for the two major blood group antigen systems, namely ABO and Rh. The ABO system is defined by two red

More information

17. A testcross A.is used to determine if an organism that is displaying a recessive trait is heterozygous or homozygous for that trait. B.

17. A testcross A.is used to determine if an organism that is displaying a recessive trait is heterozygous or homozygous for that trait. B. ch04 Student: 1. Which of the following does not inactivate an X chromosome? A. Mammals B. Drosophila C. C. elegans D. Humans 2. Who originally identified a highly condensed structure in the interphase

More information

SAP HANA Enabling Genome Analysis

SAP HANA Enabling Genome Analysis SAP HANA Enabling Genome Analysis Joanna L. Kelley, PhD Postdoctoral Scholar, Stanford University Enakshi Singh, MSc HANA Product Management, SAP Labs LLC Outline Use cases Genomics review Challenges in

More information

Using Illumina BaseSpace Apps to Analyze RNA Sequencing Data

Using Illumina BaseSpace Apps to Analyze RNA Sequencing Data Using Illumina BaseSpace Apps to Analyze RNA Sequencing Data The Illumina TopHat Alignment and Cufflinks Assembly and Differential Expression apps make RNA data analysis accessible to any user, regardless

More information

NGS and complex genetics

NGS and complex genetics NGS and complex genetics Robert Kraaij Genetic Laboratory Department of Internal Medicine r.kraaij@erasmusmc.nl Gene Hunting Rotterdam Study and GWAS Next Generation Sequencing Gene Hunting Mendelian gene

More information

Cystic Fibrosis Webquest Sarah Follenweider, The English High School 2009 Summer Research Internship Program

Cystic Fibrosis Webquest Sarah Follenweider, The English High School 2009 Summer Research Internship Program Cystic Fibrosis Webquest Sarah Follenweider, The English High School 2009 Summer Research Internship Program Introduction: Cystic fibrosis (CF) is an inherited chronic disease that affects the lungs and

More information

The Variant Call Format (VCF) Version 4.2 Specification

The Variant Call Format (VCF) Version 4.2 Specification The Variant Call Format (VCF) Version 4.2 Specification 26 Jan 2015 The master version of this document can be found at https://github.com/samtools/hts-specs. This printing is version cbd60fe from that

More information

Molecular typing of VTEC: from PFGE to NGS-based phylogeny

Molecular typing of VTEC: from PFGE to NGS-based phylogeny Molecular typing of VTEC: from PFGE to NGS-based phylogeny Valeria Michelacci 10th Annual Workshop of the National Reference Laboratories for E. coli in the EU Rome, November 5 th 2015 Molecular typing

More information

Overview of Next Generation Sequencing platform technologies

Overview of Next Generation Sequencing platform technologies Overview of Next Generation Sequencing platform technologies Dr. Bernd Timmermann Next Generation Sequencing Core Facility Max Planck Institute for Molecular Genetics Berlin, Germany Outline 1. Technologies

More information

Single-Cell DNA Sequencing with the C 1. Single-Cell Auto Prep System. Reveal hidden populations and genetic diversity within complex samples

Single-Cell DNA Sequencing with the C 1. Single-Cell Auto Prep System. Reveal hidden populations and genetic diversity within complex samples DATA Sheet Single-Cell DNA Sequencing with the C 1 Single-Cell Auto Prep System Reveal hidden populations and genetic diversity within complex samples Single-cell sensitivity Discover and detect SNPs,

More information

Heredity - Patterns of Inheritance

Heredity - Patterns of Inheritance Heredity - Patterns of Inheritance Genes and Alleles A. Genes 1. A sequence of nucleotides that codes for a special functional product a. Transfer RNA b. Enzyme c. Structural protein d. Pigments 2. Genes

More information

I. Genes found on the same chromosome = linked genes

I. Genes found on the same chromosome = linked genes Genetic recombination in Eukaryotes: crossing over, part 1 I. Genes found on the same chromosome = linked genes II. III. Linkage and crossing over Crossing over & chromosome mapping I. Genes found on the

More information

Marker-Assisted Backcrossing. Marker-Assisted Selection. 1. Select donor alleles at markers flanking target gene. Losing the target allele

Marker-Assisted Backcrossing. Marker-Assisted Selection. 1. Select donor alleles at markers flanking target gene. Losing the target allele Marker-Assisted Backcrossing Marker-Assisted Selection CS74 009 Jim Holland Target gene = Recurrent parent allele = Donor parent allele. Select donor allele at markers linked to target gene.. Select recurrent

More information

CNV Univariate Analysis Tutorial

CNV Univariate Analysis Tutorial CNV Univariate Analysis Tutorial Release 8.1 Golden Helix, Inc. March 18, 2014 Contents 1. Overview 2 2. CNAM Optimal Segmenting 4 A. Performing CNAM Optimal Segmenting..................................

More information

Lesson Plan: GENOTYPE AND PHENOTYPE

Lesson Plan: GENOTYPE AND PHENOTYPE Lesson Plan: GENOTYPE AND PHENOTYPE Pacing Two 45- minute class periods RATIONALE: According to the National Science Education Standards, (NSES, pg. 155-156), In the middle-school years, students should

More information

Tutorial for Windows and Macintosh. Preparing Your Data for NGS Alignment

Tutorial for Windows and Macintosh. Preparing Your Data for NGS Alignment Tutorial for Windows and Macintosh Preparing Your Data for NGS Alignment 2015 Gene Codes Corporation Gene Codes Corporation 775 Technology Drive, Ann Arbor, MI 48108 USA 1.800.497.4939 (USA) 1.734.769.7249

More information

Simulation Model of Mating Behavior in Flies

Simulation Model of Mating Behavior in Flies Simulation Model of Mating Behavior in Flies MEHMET KAYIM & AYKUT Ecological and Evolutionary Genetics Lab. Department of Biology, Middle East Technical University International Workshop on Hybrid Systems

More information

Tutorial on gplink. http://pngu.mgh.harvard.edu/~purcell/plink/gplink.shtml. PLINK tutorial, December 2006; Shaun Purcell, shaun@pngu.mgh.harvard.

Tutorial on gplink. http://pngu.mgh.harvard.edu/~purcell/plink/gplink.shtml. PLINK tutorial, December 2006; Shaun Purcell, shaun@pngu.mgh.harvard. Tutorial on gplink http://pngu.mgh.harvard.edu/~purcell/plink/gplink.shtml Basic gplink analyses Data management Summary statistics Association analysis Population stratification IBD-based analysis gplink

More information

Mendelian Genetics in Drosophila

Mendelian Genetics in Drosophila Mendelian Genetics in Drosophila Lab objectives: 1) To familiarize you with an important research model organism,! Drosophila melanogaster. 2) Introduce you to normal "wild type" and various mutant phenotypes.

More information

Next Generation Sequencing: Technology, Mapping, and Analysis

Next Generation Sequencing: Technology, Mapping, and Analysis Next Generation Sequencing: Technology, Mapping, and Analysis Gary Benson Computer Science, Biology, Bioinformatics Boston University gbenson@bu.edu http://tandem.bu.edu/ The Human Genome Project took

More information

Forensic DNA Testing Terminology

Forensic DNA Testing Terminology Forensic DNA Testing Terminology ABI 310 Genetic Analyzer a capillary electrophoresis instrument used by forensic DNA laboratories to separate short tandem repeat (STR) loci on the basis of their size.

More information