Introduction to the BioConductor framework Algorithmic Analysis of Flow Cytometry Data (Part 1)

Size: px
Start display at page:

Download "Introduction to the BioConductor framework Algorithmic Analysis of Flow Cytometry Data (Part 1)"

Transcription

1 Introduction to the BioConductor framework Algorithmic Analysis of Flow Cytometry Data (Part 1) Ryan Brinkman Senior Scientist, BC Cancer Agency Associate Professor, Department of Medical Genetics, UBC Vancouver, British Columbia, Canada May 17, 2014

2 Outline BioConductor for Automated Flow Cytometry Analysis Mini-tutorial on BioConductor won t teach you how to fish It will teach you that there is such a thing as fishing And that fish are tasty Even when eaten raw and wriggling

3 What Automated Analysis Needs to Deal With Large number of dimensions, events, samples Mutifactorial formats Need quick, robust processing Need to maintain data & metadata relationships No commercially available software solves these issues* Bashashti et al. Adv Bioinformatics (2009); PMC *Robinson et al. Expert Opinion Drug Discovery (2012); PMID Le Meur Curr Opin Biotechnol (2013); PMID O Neil et al. PLoS Computational Biology (2014); PMC

4 Solution: Free, Open Source Statistical Programming R is a free/libre open source, robust statistical programming environment for Windows, Mac & Linux that offers a wide range of statistical and visualization methods BioConductor provides R software modules for biological and clinical data analysis A scripted approach to high throughput data analysis Non-interactive, self-documented, reproducible Breaks problem into smaller pieces (packages) Modules can plug-in & swap-out Integrates with other software tools via open data standards Collaborative development

5 38 R packages for Flow Analysis Data processing & visualization (19/38) flowcore* platecore* flowutils* flowq* flowstats* ncdfflow QUALIFIER flowviz flowplots* flowworkspace* flowtrans* OpenCyto flowbeads flowcl flowcybar flowfit flowmap flowmatch* FCSClean Read/write & process flow data Analyze multiwell plates Import gates, transformation and compensation Quality control of ungated data Advanced statistical methods and functions Advanced methods for large dataset processing Quality control and assessment of gated data Visualization (e.g., histograms, dot plots, density plots) Graphical displays with statistical tests Importing FlowJo workspaces Estimates parameters for data transformation Simplifies data processing Analysis of flow bead data Semantic labelling of cell populations Visualize correlations between cell number& abiotic parameters Estimate proliferation of a cell population in cell-tracking dye studies Match and compare multiple flow cytometry samples Matching and meta-clustering Fluorescence vs. time gating for stream issues *Peer-reviewed manuscript available

6 15 R packages for Automated Gating flowclust* Clustering using t-mixture model with Box-Cox transformation flowmerge* flowclust + entropy-based merging flowmeans* k-means clustering and merging using the Mahalanobis distance SamSpectral* Efficient spectral clustering using density-based down-sampling flowqb Q&B analysis flowbin Combining multitube data by binning flowpeaks* Unsupervised clustering using k-means & mixture model flowfp* Fingerprint generation flowphyto* Analysis of marine biology data FLAME* Multivariate finite mixtures of skew & tailed distributions flowkoh Self-organizing maps NMF-curvHDR* Density-based clustering and non-negative matrix factorization flowcore/stats* Sequential gating and normalization w/ Beta-Binomial model PRAMS* 2D Clustering and logistic regression SPADE* Density-based sampling, k-means clustering & minimum spanning trees *Peer-reviewed manuscript available

7 4 Packages for Post-Gating Significance Assessment flowtype* Automated phenotyping using 1D gates extrapolated to multiple dimensions RchyOptimyx* Cellular hierarchies correlated with outcome of interest COMPASS Unbiased analysis of antigen-specic T-cell subsets MIMOSA* Mixture Models to model count data *Peer-reviewed manuscript available

8 BioConductor s Open, Extensible Infrastructure Packages are Interoperable & Interchangeable & Separable

9 QA: One of These Samples is Not Like the Others Le Meur et al., Cytometry A, 2007 Hahne et al., BMC Bioinformatics,

10 flowq: Summary web page

11 QA with QUALIFIER: : Flourescence Stability Finak et al., BMC Bioinformatics

12 FCSClean Remove outlier events due to stream issues Kipper Fletez-Brant & Pratip

13 Normalization Examples: Laser Change or Multiple Centres Normalization will also remove difference due to biology Movement of a subset of populations -> gating and cell populations matching problem

14 Data Normalization raw data CD3 gaussnorm CD3 fdanorm CD raw data gaussnorm fdanorm Hahne et al., Cytometry A,

15 Automated Cell Population Identification 15 different R-based algorithms for gating Many different approaches available to the problem No best solution; more is better Can be as accurate (more?!) than human gating Better chance of finding interesting populations in high-d data Allow scientists to do valuable science See Part 2 of Tutorial 5 for details of some methods also FlowCAP Workshop & Aghaeepour et al. FlowCAP Nature Methods

16 Getting Started: r-project.org

17 Getting Started: bioconductor.org

18 bioconductor.org/install OSX, Windows, Linux

19 bioconductor.org.org/help

20 bioconductor.org/help/workflows/high-throughput-assays/

21 BioConductor Vignettes Each Bioconductor package contains at least one vignette Vignettes provide a task-oriented description of functionality Vignettes contain interactive, executable examples You can access the PDF version of a vignette from R: browsevignettes(package = flowmeans ) Opens browser with links to the vignette PDF & plain-text R file containing the code used in the vignette.

22 Example Package Page

23 Vignettes: Peer-reviewed Executable Documentation

24 Documentation Peer-reviewed by Scientists

25 Working With R For Real: Getting Started RStudio

26 FCM data: Reading FCS files in R

27 Plotting FCM Data in R This is simplest example, it can mimic commercial software plots

28 Manipulating Multiple FCS Files fsapply

29 What Next? Bioinformatics.ca: Free 2-day in-depth walkthrough tutorial BioConductor.org: Mailing list of friendly people GenePattern.org: BioConductor packages in a webpage Collaborate

30 Acknowledgements Genentech Worldwide BCCA Funding R/BioConductor.org flow cytometry infrastructure Robert Gentleman All BioConductor contributors Bioinformatics.ca Tutorial Radina Droumeva $ NIH (NIBIB, NIAID), HIP-C, TFRI & TFF, CCS, MSFHR, WHCF, NSERC

THE BIOCONDUCTOR PACKAGE FLOWCORE, A SHARED DEVELOPMENT PLATFORM FOR FLOW CYTOMETRY DATA ANALYSIS IN R

THE BIOCONDUCTOR PACKAGE FLOWCORE, A SHARED DEVELOPMENT PLATFORM FOR FLOW CYTOMETRY DATA ANALYSIS IN R THE BIOCONDUCTOR PACKAGE FLOWCORE, A SHARED DEVELOPMENT PLATFORM FOR FLOW CYTOMETRY DATA ANALYSIS IN R N. Le Meur 1,2, F. Hahne 1, R. Brinkman 3, B. Ellis 5, P. Haaland 4, D. Sarkar 1, J. Spidlen 3, E.

More information

Deep profiling of multitube flow cytometry data Supplemental information

Deep profiling of multitube flow cytometry data Supplemental information Deep profiling of multitube flow cytometry data Supplemental information Kieran O Neill et al December 19, 2014 1 Table S1: Markers in simulated multitube data. The data was split into three tubes, each

More information

flowtrans: A Package for Optimizing Data Transformations for Flow Cytometry

flowtrans: A Package for Optimizing Data Transformations for Flow Cytometry flowtrans: A Package for Optimizing Data Transformations for Flow Cytometry Greg Finak, Raphael Gottardo October 13, 2014 [email protected], [email protected] Contents 1 Licensing 2 2 Overview

More information

FlowMergeCluster Documentation

FlowMergeCluster Documentation FlowMergeCluster Documentation Description: Author: Clustering of flow cytometry data using the FlowMerge algorithm. Josef Spidlen, [email protected] Please see the gp-flowcyt-help Google Group (https://groups.google.com/a/broadinstitute.org/forum/#!forum/gpflowcyt-help)

More information

Analyzing Flow Cytometry Data with Bioconductor

Analyzing Flow Cytometry Data with Bioconductor Introduction Data Analysis Analyzing Flow Cytometry Data with Bioconductor Nolwenn Le Meur, Deepayan Sarkar, Errol Strain, Byron Ellis, Perry Haaland, Florian Hahne Fred Hutchinson Cancer Research Center

More information

Exploratory Data Analysis with MATLAB

Exploratory Data Analysis with MATLAB Computer Science and Data Analysis Series Exploratory Data Analysis with MATLAB Second Edition Wendy L Martinez Angel R. Martinez Jeffrey L. Solka ( r ec) CRC Press VV J Taylor & Francis Group Boca Raton

More information

Using CyTOF Data with FlowJo Version 10.0.7. Revised 2/3/14

Using CyTOF Data with FlowJo Version 10.0.7. Revised 2/3/14 Using CyTOF Data with FlowJo Version 10.0.7 Revised 2/3/14 Table of Contents 1. Background 2. Scaling and Display Preferences 2.1 Cytometer Based Preferences 2.2 Useful Display Preferences 3. Scale and

More information

WebFOCUS RStat. RStat. Predict the Future and Make Effective Decisions Today. WebFOCUS RStat

WebFOCUS RStat. RStat. Predict the Future and Make Effective Decisions Today. WebFOCUS RStat Information Builders enables agile information solutions with business intelligence (BI) and integration technologies. WebFOCUS the most widely utilized business intelligence platform connects to any enterprise

More information

Why Taking This Course? Course Introduction, Descriptive Statistics and Data Visualization. Learning Goals. GENOME 560, Spring 2012

Why Taking This Course? Course Introduction, Descriptive Statistics and Data Visualization. Learning Goals. GENOME 560, Spring 2012 Why Taking This Course? Course Introduction, Descriptive Statistics and Data Visualization GENOME 560, Spring 2012 Data are interesting because they help us understand the world Genomics: Massive Amounts

More information

#jenkinsconf. Jenkins as a Scientific Data and Image Processing Platform. Jenkins User Conference Boston #jenkinsconf

#jenkinsconf. Jenkins as a Scientific Data and Image Processing Platform. Jenkins User Conference Boston #jenkinsconf Jenkins as a Scientific Data and Image Processing Platform Ioannis K. Moutsatsos, Ph.D., M.SE. Novartis Institutes for Biomedical Research www.novartis.com June 18, 2014 #jenkinsconf Life Sciences are

More information

Automated Quadratic Characterization of Flow Cytometer Instrument Sensitivity (flowqb Package: Introductory Processing Using Data NIH))

Automated Quadratic Characterization of Flow Cytometer Instrument Sensitivity (flowqb Package: Introductory Processing Using Data NIH)) Automated Quadratic Characterization of Flow Cytometer Instrument Sensitivity (flowqb Package: Introductory Processing Using Data NIH)) October 14, 2013 1 Licensing Under the Artistic License, you are

More information

Immunophenotyping peripheral blood cells

Immunophenotyping peripheral blood cells IMMUNOPHENOTYPING Attune Accoustic Focusing Cytometer Immunophenotyping peripheral blood cells A no-lyse, no-wash, no cell loss method for immunophenotyping nucleated peripheral blood cells using the Attune

More information

Data Mining. SPSS Clementine 12.0. 1. Clementine Overview. Spring 2010 Instructor: Dr. Masoud Yaghini. Clementine

Data Mining. SPSS Clementine 12.0. 1. Clementine Overview. Spring 2010 Instructor: Dr. Masoud Yaghini. Clementine Data Mining SPSS 12.0 1. Overview Spring 2010 Instructor: Dr. Masoud Yaghini Introduction Types of Models Interface Projects References Outline Introduction Introduction Three of the common data mining

More information

Introduction to Data Mining and Machine Learning Techniques. Iza Moise, Evangelos Pournaras, Dirk Helbing

Introduction to Data Mining and Machine Learning Techniques. Iza Moise, Evangelos Pournaras, Dirk Helbing Introduction to Data Mining and Machine Learning Techniques Iza Moise, Evangelos Pournaras, Dirk Helbing Iza Moise, Evangelos Pournaras, Dirk Helbing 1 Overview Main principles of data mining Definition

More information

Grid Density Clustering Algorithm

Grid Density Clustering Algorithm Grid Density Clustering Algorithm Amandeep Kaur Mann 1, Navneet Kaur 2, Scholar, M.Tech (CSE), RIMT, Mandi Gobindgarh, Punjab, India 1 Assistant Professor (CSE), RIMT, Mandi Gobindgarh, Punjab, India 2

More information

Instructions for Use. CyAn ADP. High-speed Analyzer. Summit 4.3. 0000050G June 2008. Beckman Coulter, Inc. 4300 N. Harbor Blvd. Fullerton, CA 92835

Instructions for Use. CyAn ADP. High-speed Analyzer. Summit 4.3. 0000050G June 2008. Beckman Coulter, Inc. 4300 N. Harbor Blvd. Fullerton, CA 92835 Instructions for Use CyAn ADP High-speed Analyzer Summit 4.3 0000050G June 2008 Beckman Coulter, Inc. 4300 N. Harbor Blvd. Fullerton, CA 92835 Overview Summit software is a Windows based application that

More information

Computational Statistics: A Crash Course using R for Biologists (and Their Friends)

Computational Statistics: A Crash Course using R for Biologists (and Their Friends) Computational Statistics: A Crash Course using R for Biologists (and Their Friends) Randall Pruim Calvin College Michigan NExT 2011 Computational Statistics Using R Why Use R? Statistics + Computation

More information

not possible or was possible at a high cost for collecting the data.

not possible or was possible at a high cost for collecting the data. Data Mining and Knowledge Discovery Generating knowledge from data Knowledge Discovery Data Mining White Paper Organizations collect a vast amount of data in the process of carrying out their day-to-day

More information

Statistics for BIG data

Statistics for BIG data Statistics for BIG data Statistics for Big Data: Are Statisticians Ready? Dennis Lin Department of Statistics The Pennsylvania State University John Jordan and Dennis K.J. Lin (ICSA-Bulletine 2014) Before

More information

Introduction to Pattern Recognition

Introduction to Pattern Recognition Introduction to Pattern Recognition Selim Aksoy Department of Computer Engineering Bilkent University [email protected] CS 551, Spring 2009 CS 551, Spring 2009 c 2009, Selim Aksoy (Bilkent University)

More information

DHL Data Mining Project. Customer Segmentation with Clustering

DHL Data Mining Project. Customer Segmentation with Clustering DHL Data Mining Project Customer Segmentation with Clustering Timothy TAN Chee Yong Aditya Hridaya MISRA Jeffery JI Jun Yao 3/30/2010 DHL Data Mining Project Table of Contents Introduction to DHL and the

More information

Today's Topics. COMP 388/441: Human-Computer Interaction. simple 2D plotting. 1D techniques. Ancient plotting techniques. Data Visualization:

Today's Topics. COMP 388/441: Human-Computer Interaction. simple 2D plotting. 1D techniques. Ancient plotting techniques. Data Visualization: COMP 388/441: Human-Computer Interaction Today's Topics Overview of visualization techniques 1D charts, 2D plots, 3D+ techniques, maps A few guidelines for scientific visualization methods, guidelines,

More information

Using self-organizing maps for visualization and interpretation of cytometry data

Using self-organizing maps for visualization and interpretation of cytometry data 1 Using self-organizing maps for visualization and interpretation of cytometry data Sofie Van Gassen, Britt Callebaut and Yvan Saeys Ghent University September, 2014 Abstract The FlowSOM package provides

More information

A MULTIVARIATE OUTLIER DETECTION METHOD

A MULTIVARIATE OUTLIER DETECTION METHOD A MULTIVARIATE OUTLIER DETECTION METHOD P. Filzmoser Department of Statistics and Probability Theory Vienna, AUSTRIA e-mail: [email protected] Abstract A method for the detection of multivariate

More information

Clustering & Visualization

Clustering & Visualization Chapter 5 Clustering & Visualization Clustering in high-dimensional databases is an important problem and there are a number of different clustering paradigms which are applicable to high-dimensional data.

More information

DELPHI 27 V 2016 CYTOMETRY STRATEGIES IN THE DIAGNOSIS OF HEMATOLOGICAL DISEASES

DELPHI 27 V 2016 CYTOMETRY STRATEGIES IN THE DIAGNOSIS OF HEMATOLOGICAL DISEASES DELPHI 27 V 2016 CYTOMETRY STRATEGIES IN THE DIAGNOSIS OF HEMATOLOGICAL DISEASES CLAUDIO ORTOLANI UNIVERSITY OF URBINO - ITALY SUN TZU (544 b.c. 496 b.c) SUN TZU (544 b.c. 496 b.c.) THE ART OF CYTOMETRY

More information

The Scientific Data Mining Process

The Scientific Data Mining Process Chapter 4 The Scientific Data Mining Process When I use a word, Humpty Dumpty said, in rather a scornful tone, it means just what I choose it to mean neither more nor less. Lewis Carroll [87, p. 214] In

More information

Using multiple models: Bagging, Boosting, Ensembles, Forests

Using multiple models: Bagging, Boosting, Ensembles, Forests Using multiple models: Bagging, Boosting, Ensembles, Forests Bagging Combining predictions from multiple models Different models obtained from bootstrap samples of training data Average predictions or

More information

BIG DATA What it is and how to use?

BIG DATA What it is and how to use? BIG DATA What it is and how to use? Lauri Ilison, PhD Data Scientist 21.11.2014 Big Data definition? There is no clear definition for BIG DATA BIG DATA is more of a concept than precise term 1 21.11.14

More information

Unsupervised Data Mining (Clustering)

Unsupervised Data Mining (Clustering) Unsupervised Data Mining (Clustering) Javier Béjar KEMLG December 01 Javier Béjar (KEMLG) Unsupervised Data Mining (Clustering) December 01 1 / 51 Introduction Clustering in KDD One of the main tasks in

More information

Introduction to Data Mining

Introduction to Data Mining Introduction to Data Mining 1 Why Data Mining? Explosive Growth of Data Data collection and data availability Automated data collection tools, Internet, smartphones, Major sources of abundant data Business:

More information

CyTOF2. Mass cytometry system. Unveil new cell types and function with high-parameter protein detection

CyTOF2. Mass cytometry system. Unveil new cell types and function with high-parameter protein detection CyTOF2 Mass cytometry system Unveil new cell types and function with high-parameter protein detection DISCOVER MORE. IMAGINE MORE. MASS CYTOMETRY. THE FUTURE OF CYTOMETRY TODAY. Mass cytometry resolves

More information

flowcore: data structures package for flow cytometry data

flowcore: data structures package for flow cytometry data flowcore: data structures package for flow cytometry data N. Le Meur F. Hahne B. Ellis P. Haaland May 4, 2010 Abstract Background The recent application of modern automation technologies to staining and

More information

Azure Machine Learning, SQL Data Mining and R

Azure Machine Learning, SQL Data Mining and R Azure Machine Learning, SQL Data Mining and R Day-by-day Agenda Prerequisites No formal prerequisites. Basic knowledge of SQL Server Data Tools, Excel and any analytical experience helps. Best of all:

More information

Structural Health Monitoring Tools (SHMTools)

Structural Health Monitoring Tools (SHMTools) Structural Health Monitoring Tools (SHMTools) Getting Started LANL/UCSD Engineering Institute LA-CC-14-046 c Copyright 2014, Los Alamos National Security, LLC All rights reserved. May 30, 2014 Contents

More information

Potency Assays for an Autologous Active Immunotherapy (Sipuleucel-T) Pocheng Liu, Ph.D. Senior Scientist of Product Development Dendreon Corporation

Potency Assays for an Autologous Active Immunotherapy (Sipuleucel-T) Pocheng Liu, Ph.D. Senior Scientist of Product Development Dendreon Corporation Potency Assays for an Autologous Active Immunotherapy (Sipuleucel-T) Pocheng Liu, Ph.D. Senior Scientist of Product Development Dendreon Corporation Sipuleucel-T Manufacturing Process Day 1 Leukapheresis

More information

Flow Data Analysis. Qianjun Zhang Application Scientist, Tree Star Inc. Oregon, USA FLOWJO CYTOMETRY DATA ANALYSIS SOFTWARE

Flow Data Analysis. Qianjun Zhang Application Scientist, Tree Star Inc. Oregon, USA FLOWJO CYTOMETRY DATA ANALYSIS SOFTWARE Flow Data Analysis Qianjun Zhang Application Scientist, Tree Star Inc. Oregon, USA Flow Data Analysis From Cells to Data - data measurement and standard Data display and gating - Data Display options -

More information

Statistical Analysis. NBAF-B Metabolomics Masterclass. Mark Viant

Statistical Analysis. NBAF-B Metabolomics Masterclass. Mark Viant Statistical Analysis NBAF-B Metabolomics Masterclass Mark Viant 1. Introduction 2. Univariate analysis Overview of lecture 3. Unsupervised multivariate analysis Principal components analysis (PCA) Interpreting

More information

Final Project Report

Final Project Report CPSC545 by Introduction to Data Mining Prof. Martin Schultz & Prof. Mark Gerstein Student Name: Yu Kor Hugo Lam Student ID : 904907866 Due Date : May 7, 2007 Introduction Final Project Report Pseudogenes

More information

An Overview of Knowledge Discovery Database and Data mining Techniques

An Overview of Knowledge Discovery Database and Data mining Techniques An Overview of Knowledge Discovery Database and Data mining Techniques Priyadharsini.C 1, Dr. Antony Selvadoss Thanamani 2 M.Phil, Department of Computer Science, NGM College, Pollachi, Coimbatore, Tamilnadu,

More information

Course Syllabus. Purposes of Course:

Course Syllabus. Purposes of Course: Course Syllabus Eco 5385.701 Predictive Analytics for Economists Summer 2014 TTh 6:00 8:50 pm and Sat. 12:00 2:50 pm First Day of Class: Tuesday, June 3 Last Day of Class: Tuesday, July 1 251 Maguire Building

More information

KNIME TUTORIAL. Anna Monreale KDD-Lab, University of Pisa Email: [email protected]

KNIME TUTORIAL. Anna Monreale KDD-Lab, University of Pisa Email: annam@di.unipi.it KNIME TUTORIAL Anna Monreale KDD-Lab, University of Pisa Email: [email protected] Outline Introduction on KNIME KNIME components Exercise: Market Basket Analysis Exercise: Customer Segmentation Exercise:

More information

BIOS 6660: Analysis of Biomedical Big Data Using R and Bioconductor, Fall 2015 Computer Lab: Education 2 North Room 2201DE (TTh 10:30 to 11:50 am)

BIOS 6660: Analysis of Biomedical Big Data Using R and Bioconductor, Fall 2015 Computer Lab: Education 2 North Room 2201DE (TTh 10:30 to 11:50 am) BIOS 6660: Analysis of Biomedical Big Data Using R and Bioconductor, Fall 2015 Computer Lab: Education 2 North Room 2201DE (TTh 10:30 to 11:50 am) Course Instructor: Dr. Tzu L. Phang, Assistant Professor

More information

Data Mining and Visualization

Data Mining and Visualization Data Mining and Visualization Jeremy Walton NAG Ltd, Oxford Overview Data mining components Functionality Example application Quality control Visualization Use of 3D Example application Market research

More information

FACSCanto RUO Special Order QUICK REFERENCE GUIDE

FACSCanto RUO Special Order QUICK REFERENCE GUIDE FACSCanto RUO Special Order QUICK REFERENCE GUIDE INSTRUMENT: 1. The computer is left on at all times. Note: If not Username: Administrator Password: BDIS 2. Unlock the screen with your PPMS account (UTSW

More information

Practical Data Science with Azure Machine Learning, SQL Data Mining, and R

Practical Data Science with Azure Machine Learning, SQL Data Mining, and R Practical Data Science with Azure Machine Learning, SQL Data Mining, and R Overview This 4-day class is the first of the two data science courses taught by Rafal Lukawiecki. Some of the topics will be

More information

Selected Topics in Electrical Engineering: Flow Cytometry Data Analysis

Selected Topics in Electrical Engineering: Flow Cytometry Data Analysis Selected Topics in Electrical Engineering: Flow Cytometry Data Analysis Bilge Karaçalı, PhD Department of Electrical and Electronics Engineering Izmir Institute of Technology Outline Compensation and gating

More information

Introduction to Flow Cytometry

Introduction to Flow Cytometry Introduction to Flow Cytometry presented by: Flow Cytometry y Core Facility Biomedical Instrumentation Center Uniformed Services University Topics Covered in this Lecture What is flow cytometry? Flow cytometer

More information

Service courses for graduate students in degree programs other than the MS or PhD programs in Biostatistics.

Service courses for graduate students in degree programs other than the MS or PhD programs in Biostatistics. Course Catalog In order to be assured that all prerequisites are met, students must acquire a permission number from the education coordinator prior to enrolling in any Biostatistics course. Courses are

More information

What is Data mining?

What is Data mining? STAT : DATA MIIG Javier Cabrera Fall Business Question Answer Business Question What is Data mining? Find Data Data Processing Extract Information Data Analysis Internal Databases Data Warehouses Internet

More information

Data Mining: Overview. What is Data Mining?

Data Mining: Overview. What is Data Mining? Data Mining: Overview What is Data Mining? Recently * coined term for confluence of ideas from statistics and computer science (machine learning and database methods) applied to large databases in science,

More information

Immunophenotyping Flow Cytometry Tutorial. Contents. Experimental Requirements...1. Data Storage...2. Voltage Adjustments...3. Compensation...

Immunophenotyping Flow Cytometry Tutorial. Contents. Experimental Requirements...1. Data Storage...2. Voltage Adjustments...3. Compensation... Immunophenotyping Flow Cytometry Tutorial Contents Experimental Requirements...1 Data Storage...2 Voltage Adjustments...3 Compensation...5 Experimental Requirements For immunophenotyping with FITC and

More information

COPYRIGHTED MATERIAL. Contents. List of Figures. Acknowledgments

COPYRIGHTED MATERIAL. Contents. List of Figures. Acknowledgments Contents List of Figures Foreword Preface xxv xxiii xv Acknowledgments xxix Chapter 1 Fraud: Detection, Prevention, and Analytics! 1 Introduction 2 Fraud! 2 Fraud Detection and Prevention 10 Big Data for

More information

123count ebeads Catalog Number: 01-1234 Also known as: Absolute cell count beads GPR: General Purpose Reagents. For Laboratory Use.

123count ebeads Catalog Number: 01-1234 Also known as: Absolute cell count beads GPR: General Purpose Reagents. For Laboratory Use. Page 1 of 1 Catalog Number: 01-1234 Also known as: Absolute cell count beads GPR: General Purpose Reagents. For Laboratory Use. Normal human peripheral blood was stained with Anti- Human CD45 PE (cat.

More information

OplAnalyzer: A Toolbox for MALDI-TOF Mass Spectrometry Data Analysis

OplAnalyzer: A Toolbox for MALDI-TOF Mass Spectrometry Data Analysis OplAnalyzer: A Toolbox for MALDI-TOF Mass Spectrometry Data Analysis Thang V. Pham and Connie R. Jimenez OncoProteomics Laboratory, Cancer Center Amsterdam, VU University Medical Center De Boelelaan 1117,

More information

HT2015: SC4 Statistical Data Mining and Machine Learning

HT2015: SC4 Statistical Data Mining and Machine Learning HT2015: SC4 Statistical Data Mining and Machine Learning Dino Sejdinovic Department of Statistics Oxford http://www.stats.ox.ac.uk/~sejdinov/sdmml.html Bayesian Nonparametrics Parametric vs Nonparametric

More information

Gates/filters in Flow Cytometry Data Visualization

Gates/filters in Flow Cytometry Data Visualization Gates/filters in Flow Cytometry Data Visualization October 3, Abstract The flowviz package provides tools for visualization of flow cytometry data. This document describes the support for visualizing gates

More information

2015 Workshops for Professors

2015 Workshops for Professors SAS Education Grow with us Offered by the SAS Global Academic Program Supporting teaching, learning and research in higher education 2015 Workshops for Professors 1 Workshops for Professors As the market

More information

Using Data Mining for Mobile Communication Clustering and Characterization

Using Data Mining for Mobile Communication Clustering and Characterization Using Data Mining for Mobile Communication Clustering and Characterization A. Bascacov *, C. Cernazanu ** and M. Marcu ** * Lasting Software, Timisoara, Romania ** Politehnica University of Timisoara/Computer

More information

Spherotech, Inc. 27845 Irma Lee Circle, Unit 101, Lake Forest, Illinois 60045 1

Spherotech, Inc. 27845 Irma Lee Circle, Unit 101, Lake Forest, Illinois 60045 1 78 Irma Lee Circle, Unit 0, Lake Forest, Illinois 00 SPHERO TM Technical Note STN- Rev B 008 DETERMINING PMT LINEARITY IN FLOW CYTOMETERS USING THE SPHERO TM PMT QUALITY CONTROL EXCEL TEMPLATE Introduction

More information

How To Cluster

How To Cluster Data Clustering Dec 2nd, 2013 Kyrylo Bessonov Talk outline Introduction to clustering Types of clustering Supervised Unsupervised Similarity measures Main clustering algorithms k-means Hierarchical Main

More information

The Open2Dprot Proteomics Project for n-dimensional Protein Expression Data Analysis

The Open2Dprot Proteomics Project for n-dimensional Protein Expression Data Analysis The Open2Dprot Proteomics Project for n-dimensional Protein Expression Data Analysis http://open2dprot.sourceforge.net/ Revised 2-05-2006 * (cf. 2D-LC) Introduction There is a need for integrated proteomics

More information

Cytell Cell Imaging System

Cytell Cell Imaging System Cytell Cell Imaging System Simplify your cell analysis and start changing tomorrow today www.gelifesciences.com/cytell The Cytell Cell Imaging System is an intuitive cell analysis instrument that streamlines

More information

Machine Learning with MATLAB David Willingham Application Engineer

Machine Learning with MATLAB David Willingham Application Engineer Machine Learning with MATLAB David Willingham Application Engineer 2014 The MathWorks, Inc. 1 Goals Overview of machine learning Machine learning models & techniques available in MATLAB Streamlining the

More information

Russell K. Anderson. Visual Data Mining THE VISMINER APPROACH

Russell K. Anderson. Visual Data Mining THE VISMINER APPROACH Russell K. Anderson Visual Data Mining THE VISMINER APPROACH Visual Data Mining Visual Data Mining The VisMiner Approach RUSSELL K. ANDERSON VisTech, USA This edition first published 2013 # 2013 John

More information

How To Perform An Ensemble Analysis

How To Perform An Ensemble Analysis Charu C. Aggarwal IBM T J Watson Research Center Yorktown, NY 10598 Outlier Ensembles Keynote, Outlier Detection and Description Workshop, 2013 Based on the ACM SIGKDD Explorations Position Paper: Outlier

More information

Non-negative Matrix Factorization (NMF) in Semi-supervised Learning Reducing Dimension and Maintaining Meaning

Non-negative Matrix Factorization (NMF) in Semi-supervised Learning Reducing Dimension and Maintaining Meaning Non-negative Matrix Factorization (NMF) in Semi-supervised Learning Reducing Dimension and Maintaining Meaning SAMSI 10 May 2013 Outline Introduction to NMF Applications Motivations NMF as a middle step

More information

COC131 Data Mining - Clustering

COC131 Data Mining - Clustering COC131 Data Mining - Clustering Martin D. Sykora [email protected] Tutorial 05, Friday 20th March 2009 1. Fire up Weka (Waikako Environment for Knowledge Analysis) software, launch the explorer window

More information

Customer and Business Analytic

Customer and Business Analytic Customer and Business Analytic Applied Data Mining for Business Decision Making Using R Daniel S. Putler Robert E. Krider CRC Press Taylor &. Francis Group Boca Raton London New York CRC Press is an imprint

More information

Hierarchical Cluster Analysis Some Basics and Algorithms

Hierarchical Cluster Analysis Some Basics and Algorithms Hierarchical Cluster Analysis Some Basics and Algorithms Nethra Sambamoorthi CRMportals Inc., 11 Bartram Road, Englishtown, NJ 07726 (NOTE: Please use always the latest copy of the document. Click on this

More information

DATA MINING CLUSTER ANALYSIS: BASIC CONCEPTS

DATA MINING CLUSTER ANALYSIS: BASIC CONCEPTS DATA MINING CLUSTER ANALYSIS: BASIC CONCEPTS 1 AND ALGORITHMS Chiara Renso KDD-LAB ISTI- CNR, Pisa, Italy WHAT IS CLUSTER ANALYSIS? Finding groups of objects such that the objects in a group will be similar

More information

Primetime for KNIME:

Primetime for KNIME: Primetime for KNIME: Towards an Integrated Analysis and Visualization Environment for RNAi Screening Data F. Oliver Gathmann, Ph. D. Director IT, Cenix BioScience Presentation for: KNIME User Group Meeting

More information

Chapter ML:XI (continued)

Chapter ML:XI (continued) Chapter ML:XI (continued) XI. Cluster Analysis Data Mining Overview Cluster Analysis Basics Hierarchical Cluster Analysis Iterative Cluster Analysis Density-Based Cluster Analysis Cluster Evaluation Constrained

More information

Hierarchical Clustering Analysis

Hierarchical Clustering Analysis Hierarchical Clustering Analysis What is Hierarchical Clustering? Hierarchical clustering is used to group similar objects into clusters. In the beginning, each row and/or column is considered a cluster.

More information

Lavastorm Analytic Library Predictive and Statistical Analytics Node Pack FAQs

Lavastorm Analytic Library Predictive and Statistical Analytics Node Pack FAQs 1.1 Introduction Lavastorm Analytic Library Predictive and Statistical Analytics Node Pack FAQs For brevity, the Lavastorm Analytics Library (LAL) Predictive and Statistical Analytics Node Pack will be

More information

Data Mining Clustering (2) Sheets are based on the those provided by Tan, Steinbach, and Kumar. Introduction to Data Mining

Data Mining Clustering (2) Sheets are based on the those provided by Tan, Steinbach, and Kumar. Introduction to Data Mining Data Mining Clustering (2) Toon Calders Sheets are based on the those provided by Tan, Steinbach, and Kumar. Introduction to Data Mining Outline Partitional Clustering Distance-based K-means, K-medoids,

More information

What is Data Mining? MS4424 Data Mining & Modelling. MS4424 Data Mining & Modelling. MS4424 Data Mining & Modelling. MS4424 Data Mining & Modelling

What is Data Mining? MS4424 Data Mining & Modelling. MS4424 Data Mining & Modelling. MS4424 Data Mining & Modelling. MS4424 Data Mining & Modelling MS4424 Data Mining & Modelling MS4424 Data Mining & Modelling Lecturer : Dr Iris Yeung Room No : P7509 Tel No : 2788 8566 Email : [email protected] 1 Aims To introduce the basic concepts of data mining

More information

Using the Grid for the interactive workflow management in biomedicine. Andrea Schenone BIOLAB DIST University of Genova

Using the Grid for the interactive workflow management in biomedicine. Andrea Schenone BIOLAB DIST University of Genova Using the Grid for the interactive workflow management in biomedicine Andrea Schenone BIOLAB DIST University of Genova overview background requirements solution case study results background A multilevel

More information

MS1b Statistical Data Mining

MS1b Statistical Data Mining MS1b Statistical Data Mining Yee Whye Teh Department of Statistics Oxford http://www.stats.ox.ac.uk/~teh/datamining.html Outline Administrivia and Introduction Course Structure Syllabus Introduction to

More information

Identification of rheumatoid arthritis and osteoarthritis patients by transcriptome-based rule set generation

Identification of rheumatoid arthritis and osteoarthritis patients by transcriptome-based rule set generation Identification of rheumatoid arthritis and osterthritis patients by transcriptome-based rule set generation Bering Limited Report generated on September 19, 2014 Contents 1 Dataset summary 2 1.1 Project

More information

BD FACSComp Software Tutorial

BD FACSComp Software Tutorial BD FACSComp Software Tutorial This tutorial guides you through a BD FACSComp software lyse/no-wash assay setup run. If you are already familiar with previous versions of BD FACSComp software on Mac OS

More information

LabKey Server: An open source platform for scientific data integration, analysis, and collaboration

LabKey Server: An open source platform for scientific data integration, analysis, and collaboration LabKey Server: An open source platform for scientific data integration, analysis, and collaboration Mark Igra Partner, LabKey Software [email protected] Presentation Topics Why Scientific Data Integration

More information

Comparative Analysis of EM Clustering Algorithm and Density Based Clustering Algorithm Using WEKA tool.

Comparative Analysis of EM Clustering Algorithm and Density Based Clustering Algorithm Using WEKA tool. International Journal of Engineering Research and Development e-issn: 2278-067X, p-issn: 2278-800X, www.ijerd.com Volume 9, Issue 8 (January 2014), PP. 19-24 Comparative Analysis of EM Clustering Algorithm

More information

Thermo Scientific CellInsight CX5 HCS Platform. discover more uncover the possibilities

Thermo Scientific CellInsight CX5 HCS Platform. discover more uncover the possibilities Thermo Scientific CellInsight CX5 HCS Platform discover more uncover the possibilities high content screening Driving your Productivity with Innovated Technology What is High Content Screening (HCS)? High

More information

What s New in SPSS 16.0

What s New in SPSS 16.0 SPSS 16.0 New capabilities What s New in SPSS 16.0 SPSS Inc. continues its tradition of regularly enhancing this family of powerful but easy-to-use statistical software products with the release of SPSS

More information

Imaging and Bioinformatics Software

Imaging and Bioinformatics Software Imaging and Bioinformatics Software Software Overview 242 Gel Analysis Software 243 Ordering Information 246 Software Overview Software Overview See Also Imaging systems: pages 232 237. Bio-Plex Manager

More information

Data Mining Cluster Analysis: Basic Concepts and Algorithms. Lecture Notes for Chapter 8. Introduction to Data Mining

Data Mining Cluster Analysis: Basic Concepts and Algorithms. Lecture Notes for Chapter 8. Introduction to Data Mining Data Mining Cluster Analysis: Basic Concepts and Algorithms Lecture Notes for Chapter 8 Introduction to Data Mining by Tan, Steinbach, Kumar Tan,Steinbach, Kumar Introduction to Data Mining 4/8/2004 Hierarchical

More information

Compensation Basics - Bagwell. Compensation Basics. C. Bruce Bagwell MD, Ph.D. Verity Software House, Inc.

Compensation Basics - Bagwell. Compensation Basics. C. Bruce Bagwell MD, Ph.D. Verity Software House, Inc. Compensation Basics C. Bruce Bagwell MD, Ph.D. Verity Software House, Inc. 2003 1 Intrinsic or Autofluorescence p2 ac 1,2 c 1 ac 1,1 p1 In order to describe how the general process of signal cross-over

More information

Data Mining. Cluster Analysis: Advanced Concepts and Algorithms

Data Mining. Cluster Analysis: Advanced Concepts and Algorithms Data Mining Cluster Analysis: Advanced Concepts and Algorithms Tan,Steinbach, Kumar Introduction to Data Mining 4/18/2004 1 More Clustering Methods Prototype-based clustering Density-based clustering Graph-based

More information

AGILENT S BIOINFORMATICS ANALYSIS SOFTWARE

AGILENT S BIOINFORMATICS ANALYSIS SOFTWARE ACCELERATING PROGRESS IS IN OUR GENES AGILENT S BIOINFORMATICS ANALYSIS SOFTWARE GENESPRING GENE EXPRESSION (GX) MASS PROFILER PROFESSIONAL (MPP) PATHWAY ARCHITECT (PA) See Deeper. Reach Further. BIOINFORMATICS

More information

Clustering. Adrian Groza. Department of Computer Science Technical University of Cluj-Napoca

Clustering. Adrian Groza. Department of Computer Science Technical University of Cluj-Napoca Clustering Adrian Groza Department of Computer Science Technical University of Cluj-Napoca Outline 1 Cluster Analysis What is Datamining? Cluster Analysis 2 K-means 3 Hierarchical Clustering What is Datamining?

More information

Information Management course

Information Management course Università degli Studi di Milano Master Degree in Computer Science Information Management course Teacher: Alberto Ceselli Lecture 01 : 06/10/2015 Practical informations: Teacher: Alberto Ceselli ([email protected])

More information

Standardization, Calibration and Quality Control

Standardization, Calibration and Quality Control Standardization, Calibration and Quality Control Ian Storie Flow cytometry has become an essential tool in the research and clinical diagnostic laboratory. The range of available flow-based diagnostic

More information

Lecture 9: Introduction to Pattern Analysis

Lecture 9: Introduction to Pattern Analysis Lecture 9: Introduction to Pattern Analysis g Features, patterns and classifiers g Components of a PR system g An example g Probability definitions g Bayes Theorem g Gaussian densities Features, patterns

More information