Development of Big Data Infrastructure in NCHC: From Sensing to Understanding

Size: px
Start display at page:

Download "Development of Big Data Infrastructure in NCHC: From Sensing to Understanding"

Transcription

1 Development of Big Data Infrastructure in NCHC: From Sensing to Understanding Fang-Pang Lin National Center for High-Performance Computing National Applied Research Laboratories, Taiwan ASC 2015

2 BIG DATA ECOSYSTEM : FROM DATA TO DECISIONS Source: Craig Stires DATA CREATION DATA ACQUISITION INFO PROCESSING BUSINESS PROCESS PRODUCERS ARCHITECTS / ENGINEERS ANALYSTS / SCIENTISTS END USERS Machine and Sensors Shared Nothing Scale-out Storage + SSD Data Exploration Access-Anywhere Analytics Services ON-DEMAND Transaction and Usage Logs Geolocation Mobile Apps Data and Messaging VOLUME VELOCITY VARIETY MPP + In-Memory Compute Non-relational DWH Converged Infrastructure Hi-Speed / - Resiliency Networking Hadoop DEEP INSIGHTS REAL-TIME EVENTS Contextualized Data Modeling / Scenarios Forecasting Stream Processing PUSH Context-Aware Business Applications Location-Based Services Alert and Respond Workflow and Interaction Automation VALUE Relationships and Social Influence SYSTEMS INTEGRATION Cloud OBJECTIVES Event Management EMBEDDED DELIVERY MODELS Smart devices and systems 2 2 Copyright IDC (2013)

3 BIG DATA ECOSYSTEM : FROM DATA TO DECISIONS Source: Craig Stires DATA CREATION DATA ACQUISITION INFO PROCESSING BUSINESS PROCESS PRODUCERS ARCHITECTS / ENGINEERS ANALYSTS / SCIENTISTS END USERS User-gen Text 68% (Data collected) Machine and Sensors Transaction and Usage Logs Geolocation Transactional 67% (Data collected) Mobile Apps Data and Messaging Machine or device 41% (Data collected) Relationships and Social Influence VOLUME VELOCITY VARIETY SYSTEMS INTEGRATION Shared Nothing Scale-out Storage + SSD MPP + In-Memory Compute Non-relational DWH Cloud Managing >5TB data 36% Converged Infrastructure Deployed/ing Hadoop 21% Hi-Speed / - Resiliency Networking Keep/discard which data Hadoop 35% (#1 BD Challenge) DEEP INSIGHTS REAL-TIME EVENTS OBJECTIVES Deployed/ing Data Exploration Text Analytics 26% Contextualized Data Modeling / Scenarios Managing Data Quality (#1 IT Forecasting Challenge) Stream Processing Deployed/ing Event Processing 22% Event Management ON-DEMAND PUSH EMBEDDED DELIVERY MODELS Customer Engagement Access-Anywhere Analytics Services [microsegment, net promoter, personalization] (2013 BD Hotspots) Context-Aware Business Applications Geofencing [retail offers, asset Location-Based tracking] Services (GPS Innovations) Alert and Respond Smart Cities Workflow and Interaction Automation [power, water, traffic, safety] (M2M Innovations) Smart devices and systems VALUE 3 3 Copyright IDC (2013)

4 APeJ BDA Maturity by Country Market Size (USD) Source: Craig Stires Source: IDC APeJ Big Data Maturity Assessment and Benchmark 2013 (n=802) Source: IDC APeJ Big Data Market Analysis & Forecast , Oct 2013 IDC Visit us at IDC.com and follow us on 4

5 Trends among Big Data leaders Hyper-competition, regulation, analytics sophistication Source: Craig Stires HK - IaaS investments, Telco asset utilization, FSI risk SGP - FSI microsegment, Telco LBS & data monetization, ICA AUS - FSI personalization, Telco psychographic, Data privacy NZ - Emergency svc, SME SaaS, Fleet mgmt, Farm asset mgmt IDC Visit us at IDC.com and follow us on 5

6 Trends among Big Data midstream Mass urbanization, first-measured, neighborly pressure Source: Craig Stires KOR - M2M automation, Manu QC, Govt NLP audio analytics PRC - Web 2.0 [ABT] xacts, City CCTV, FSI microfinance TWN - Smart Tourism, Data Privacy, Manu process control IND - Citizen registration, banking the unbanked, Big Data insource IDC Visit us at IDC.com and follow us on 6

7 The Path from Infrastructure to Data Sensing for Understanding Sensing: (Networks change the game!) Evolve since 10 years ago: Ecogrid, SARS Grid, etc Institutional missions based on special vehicles: Satellites, Research Ships & Aircrafts, Met Stations etc. It is growing even larger and broader, e.g. IOT, social network. Understanding: Modeling from hypothesis to discovery Phenomena, Behavior etc

8 Ecogrid: Sense the nature in a new way (10 years on)

9 The Path from Infrastructure to Data Sensing for Understanding Sensing: Evolve since 10 years ago: Ecogrid, SARS Grid, etc Institutional missions based on special vehicles: Satellites, Research Ships & Aircrafts, Met Stations etc. It is growing even larger and broader, e.g. IOT, social network. Understanding: Modeling from hypothesis to discovery Phenomena, Behavior.. etc.

10 Fish4Knowledge human level query for Marine Biology U. of Edinburgh (UK) Video Data: 10 camera years = 10 cameraframes = 112 Tb massive data storage Descriptive data: 20 Tb Fish: 10^10 detections Summary data: 500 Gb Target: 1 sec query answering Centrum Wiskunde & Informatica (Netherlands) UCATANIA (Italy) NCHC NCHC: sustainable system for data acquisition, storage, and computing. U. of Catania in Italy: fish detection and tracking. U. of Edinburgh in UK: workflow, fish recognition, and fish behavior. CWI in Netherlands: user interfaces. 10

11 Global Lake Ecological Observational Network (GLEON) Lakebase: harvest quality data from internet and collect more than ~25,000 lakes across the world. Global compute service through CONDOR >10 major real time observational data from selected GLEON sites. Source: Pau Hanson 11

12 Infrastructure enables Fish4Knowledge Cloud of Resources Clients Storage platform Service Frontend Video Server Data Source Computing platform Storage Nodes Computing Nodes Marine Biologists Application Developers 2015/5/25 12

13 2015/5/25 Work On Linking Open Data (e.g. FAO) 13

14 Video Classification (~350K videos from 2009 to 2013) Algae: 9.2% Blurred: 33.5% Complex Scenes: 4.3% Encoding: 23.9% Highly Blurred: 13.9% Normal: 12.9% Unknown: 2.2% 14

15 The SEATANK Event & Other SEATANK event It happened again! Require scientific evidences for Chinese Taipei to win the lawsuit. (TORI, NSPO, NCHC) Freighter SEATANK stranded in Penhu water 2006 Cargo ship Tzini stranded in Yilan water and leaks >100 tons fuel Oil 15

16 The Path from Infrastructure to Data Sensing for Understanding Sensing: (Networks change the game!) Evolve since 10 years ago: Ecogrid, SARS Grid, etc Institutional missions based on special vehicles: Satellites, Research Ships & Aircrafts, Met Stations etc. It is growing even larger and broader, e.g. IOT, social network. Understanding: (Data change the game!) Modeling from hypothesis to discovery Phenomena, Behavior etc

17 CC image by Sharyn Morrow on Flickr Natural disaster Facilities infrastructure failure Storage failure Server hardware/software failure Application software failure External dependencies (e.g. PKI failure) Format obsolescence Legal encumbrance Human error Malicious attack by human or automated agents Loss of staffing competencies Loss of institutional commitment Loss of financial stability Changes in user expectations and requirements Source: Source: DataONE

18 Info. content Time for Publication Specific Details Data Management General Details Retirement or Career Change Accident Death Time (From Michener et al 1997)

19 The Treasure of NARLabs: Earth Science Observational Data 19

20 Big Data Infrastructure Challenge in ESOD Data Features: high complexity, large scale, frequency, real-time and stream Big Data Data Size: 155TB/yr. (Actual data size will 2~3x) Currently, 103 TB for Q1. Simulation data 200TB from TTFRI. Real time, High frequency Data: Process > 18,000 records/sec. Currently 6,000 records/sec. 20

21 ESOD Big Data Service ~ 1PB External data sources Data-intensive computing system (e.g. Hadoop) Parallel data server Parallel compute server Parallel query server Continuous query stream Query Result Continuous query results External client d 1 d 2 d 3 Source dataset Derived datasets Parallel file system (e.g., GFS, HDFS, GPFS) Virtually Centralized Resources Source: David O Hallaron Characteristics: Small queries and results Massive data and computation performed on server Examples: Search Photo scene completion Log processing Science analytics 21

22 ESOD Hardware Architecture Challenges of use of distributed resources Exmaples of Datasets 1 TTFRI NCREE NSPO 1 2 1G - 10G Fiber TORI 8G Optical Fiber WindRider: 40 G infinitband (Lustre FS) GPFS Preload/Stage ESOD Storage NCHC 10G Bridge Gateway GPFS Disk 918TB Tape 433TB 22

23 ESOD Data Archiving and System Infrastructure HAProxy https ftps NCHC firewall Failover between 3 sites. Automatic load balancing. Transparent file systems: direct read/write of files across different sites. Automatic storage backup. Load balance server Slave Load balance server Master High IOPS WindRider Formosa 2/3/5 Paralllel DBMS Scale out > 1 PB 2 X 10GE 2 X 10GE 2 X 10GE HSM General Parallel File System Tape library HS M Tape library General Parallel File System HSM General Parallel File System Tape library 23 23

24 Data discovery: Metadata Catalog Use relational database (mysql) with multi-dimensional schema design to speed up searching (migrating to NOSQL solutions now) System Metadata User name space - Address / / telephone number - Role (administrator, curator, user) - File name space Example: geonetwork - Creation date / size / location / checksum - Owner / access controls Storage resource name space - Capacity / quotas / Type (archive, disk, fast cache) Domain Metadata User-given metadata - Key-Value-Unit Triplets, Annotation - Relational / XML Metadata - Domain-specific Schema Adopt OGC standards 24

25 Smart query and answer Develop a set of control vocabulary based on int l standards, e.g. HDF, NetCDF, OGC etc. Derive RDF triple dataset from common query tasks of ESOD. Combination of visual data plus metadata to support a specific high-level information seeking user task. Design of an interface for graph & visual comparison search. Selected one specific task, that of comparing sets of objects, and designed a prototype interface on top of linked data sets used by the experts to support this task explicitly Develop methods for data provenance. 25

26 The Path from Infrastructure to Data Sensing for Understanding Sensing: (Infrastructure-Centric ) Evolve since 10 years ago: Ecogrid, SARS Grid, etc Institutional missions based on special vehicles: Satellites, Research Ships & Aircrafts, Met Stations etc. It is growing even larger and broader, e.g. IOT, social network. Understanding: (Data-Centric) Modeling from hypothesis to discovery Phenomena, Behavior etc in a sophisticated manner.

27 Government Big Data Government Clouds & Open data Big Data 10 government clouds: Health, Food, e-invoice, Transportation, Environment, Finance, Education, Culture, Disaster Prevention, Geospatial Information, Agriculture and e-government, focusing on societal impact via Research & innovation. 3 kinds of data: Infrastructural Data, Public data and Personal data, open but required protection. Open, but Protected Licensing models: Open Government License (OGL), Non-Commercial Government License, Charged License. Access Platform Facility and Systems to enable Data store, Management and Analytics. Access Control with the licensing models and use scenarios.

28 Government Big Data NCHC/NARLabs provides Big Data Platform, Bridging Excellent Research Societal Impact Culture: Animation Finance: e-invoice Health: Electronic Medical Record Network: Openflow Food Security: Food Traceability Transportation: etag system

Development of Earth Science Observational Data Infrastructure of Taiwan. Fang-Pang Lin National Center for High-Performance Computing, Taiwan

Development of Earth Science Observational Data Infrastructure of Taiwan. Fang-Pang Lin National Center for High-Performance Computing, Taiwan Development of Earth Science Observational Data Infrastructure of Taiwan Fang-Pang Lin National Center for High-Performance Computing, Taiwan GLIF 13, Singapore, 4 Oct 2013 The Path from Infrastructure

More information

End to End Solution to Accelerate Data Warehouse Optimization. Franco Flore Alliance Sales Director - APJ

End to End Solution to Accelerate Data Warehouse Optimization. Franco Flore Alliance Sales Director - APJ End to End Solution to Accelerate Data Warehouse Optimization Franco Flore Alliance Sales Director - APJ Big Data Is Driving Key Business Initiatives Increase profitability, innovation, customer satisfaction,

More information

The Enterprise Data Hub and The Modern Information Architecture

The Enterprise Data Hub and The Modern Information Architecture The Enterprise Data Hub and The Modern Information Architecture Dr. Amr Awadallah CTO & Co-Founder, Cloudera Twitter: @awadallah 1 2013 Cloudera, Inc. All rights reserved. Cloudera Overview The Leader

More information

Optimized for the Industrial Internet: GE s Industrial Data Lake Platform

Optimized for the Industrial Internet: GE s Industrial Data Lake Platform Optimized for the Industrial Internet: GE s Industrial Lake Platform Agenda The Opportunity The Solution The Challenges The Results Solutions for Industrial Internet, deep domain expertise 2 GESoftware.com

More information

Big Data Are You Ready? Jorge Plascencia Solution Architect Manager

Big Data Are You Ready? Jorge Plascencia Solution Architect Manager Big Data Are You Ready? Jorge Plascencia Solution Architect Manager Big Data: The Datafication Of Everything Thoughts Devices Processes Thoughts Things Processes Run the Business Organize data to do something

More information

EO Data by using SAP HANA Spatial Hinnerk Gildhoff, Head of HANA Spatial, SAP Satellite Masters Conference 21 th October 2015 Public

EO Data by using SAP HANA Spatial Hinnerk Gildhoff, Head of HANA Spatial, SAP Satellite Masters Conference 21 th October 2015 Public Leveraging Geospatial Technologies EO Data by using SAP HANA Spatial Hinnerk Gildhoff, Head of HANA Spatial, SAP Satellite Masters Conference 21 th October 2015 Public Disclaimer This presentation outlines

More information

Building a Datacenter Infrastructure to Support Your Big Data Plans

Building a Datacenter Infrastructure to Support Your Big Data Plans WHITE PAPER Building a Datacenter Infrastructure to Support Your Big Data Plans Sponsored by: Cisco in collaboration with Intel Richard L. Villars January 2014 Dan Vesset IDC OPINION Companies rely on

More information

Big Data, Cloud Computing, Spatial Databases Steven Hagan Vice President Server Technologies

Big Data, Cloud Computing, Spatial Databases Steven Hagan Vice President Server Technologies Big Data, Cloud Computing, Spatial Databases Steven Hagan Vice President Server Technologies Big Data: Global Digital Data Growth Growing leaps and bounds by 40+% Year over Year! 2009 =.8 Zetabytes =.08

More information

Big Data on AWS. Services Overview. Bernie Nallamotu Principle Solutions Architect

Big Data on AWS. Services Overview. Bernie Nallamotu Principle Solutions Architect on AWS Services Overview Bernie Nallamotu Principle Solutions Architect \ So what is it? When your data sets become so large that you have to start innovating around how to collect, store, organize, analyze

More information

Big Data and Analytics: Challenges and Opportunities

Big Data and Analytics: Challenges and Opportunities Big Data and Analytics: Challenges and Opportunities Dr. Amin Beheshti Lecturer and Senior Research Associate University of New South Wales, Australia (Service Oriented Computing Group, CSE) Talk: Sharif

More information

Apache Hadoop in the Enterprise. Dr. Amr Awadallah, CTO/Founder @awadallah, aaa@cloudera.com

Apache Hadoop in the Enterprise. Dr. Amr Awadallah, CTO/Founder @awadallah, aaa@cloudera.com Apache Hadoop in the Enterprise Dr. Amr Awadallah, CTO/Founder @awadallah, aaa@cloudera.com Cloudera The Leader in Big Data Management Powered by Apache Hadoop The Leading Open Source Distribution of Apache

More information

Sustainable Development with Geospatial Information Leveraging the Data and Technology Revolution

Sustainable Development with Geospatial Information Leveraging the Data and Technology Revolution Sustainable Development with Geospatial Information Leveraging the Data and Technology Revolution Steven Hagan, Vice President, Server Technologies 1 Copyright 2011, Oracle and/or its affiliates. All rights

More information

Optimized for the Industrial Internet: GE s Industrial Data Lake Platform

Optimized for the Industrial Internet: GE s Industrial Data Lake Platform Optimized for the Industrial Internet: GE s Industrial Lake Platform Agenda Opportunity Solution Challenges Result GE Lake 2 GESoftware.com @GESoftware #IndustrialInternet Big opportunities with Industrial

More information

SCALABLE FILE SHARING AND DATA MANAGEMENT FOR INTERNET OF THINGS

SCALABLE FILE SHARING AND DATA MANAGEMENT FOR INTERNET OF THINGS Sean Lee Solution Architect, SDI, IBM Systems SCALABLE FILE SHARING AND DATA MANAGEMENT FOR INTERNET OF THINGS Agenda Converging Technology Forces New Generation Applications Data Management Challenges

More information

Enabling Manufacturing Transformation in a Connected World. John Shewchuk Technical Fellow DX

Enabling Manufacturing Transformation in a Connected World. John Shewchuk Technical Fellow DX Enabling Manufacturing Transformation in a Connected World John Shewchuk Technical Fellow DX Internet of Things What is the Internet of Things? The network of physical objects that contain embedded technology

More information

Big Data Analytics Platform @ Nokia

Big Data Analytics Platform @ Nokia Big Data Analytics Platform @ Nokia 1 Selecting the Right Tool for the Right Workload Yekesa Kosuru Nokia Location & Commerce Strata + Hadoop World NY - Oct 25, 2012 Agenda Big Data Analytics Platform

More information

The 4 Pillars of Technosoft s Big Data Practice

The 4 Pillars of Technosoft s Big Data Practice beyond possible Big Use End-user applications Big Analytics Visualisation tools Big Analytical tools Big management systems The 4 Pillars of Technosoft s Big Practice Overview Businesses have long managed

More information

COMP9321 Web Application Engineering

COMP9321 Web Application Engineering COMP9321 Web Application Engineering Semester 2, 2015 Dr. Amin Beheshti Service Oriented Computing Group, CSE, UNSW Australia Week 11 (Part II) http://webapps.cse.unsw.edu.au/webcms2/course/index.php?cid=2411

More information

A Next-Generation Analytics Ecosystem for Big Data. Colin White, BI Research September 2012 Sponsored by ParAccel

A Next-Generation Analytics Ecosystem for Big Data. Colin White, BI Research September 2012 Sponsored by ParAccel A Next-Generation Analytics Ecosystem for Big Data Colin White, BI Research September 2012 Sponsored by ParAccel BIG DATA IS BIG NEWS The value of big data lies in the business analytics that can be generated

More information

White Paper. How Streaming Data Analytics Enables Real-Time Decisions

White Paper. How Streaming Data Analytics Enables Real-Time Decisions White Paper How Streaming Data Analytics Enables Real-Time Decisions Contents Introduction... 1 What Is Streaming Analytics?... 1 How Does SAS Event Stream Processing Work?... 2 Overview...2 Event Stream

More information

BIG Big Data Public Private Forum

BIG Big Data Public Private Forum DATA STORAGE Martin Strohbach, AGT International (R&D) THE DATA VALUE CHAIN Value Chain Data Acquisition Data Analysis Data Curation Data Storage Data Usage Structured data Unstructured data Event processing

More information

CAPITALIZE ON BIG DATA

CAPITALIZE ON BIG DATA CAPITALIZE ON BIG DATA SARA GARDNER, SENIOR DIRECTOR OF SOFTWARE PRODUCT MARKETING 1 Hitachi Data Systems Corporation 2013. All Rights Reserved. WEBTECH EDUCATIONAL SERIES CAPITALIZE ON BIG DATA We are

More information

Case Studies and Needs: Disaster Mitigation

Case Studies and Needs: Disaster Mitigation Taiwan 2012 Bridging Big Data Infrastructure Workshop Case Studies and Needs: Disaster Mitigation Whey-Fone Tsai National Center for High-performance Computing National Applied Research Laboratories Hsinchu,

More information

REAL-TIME OPERATIONAL INTELLIGENCE. Competitive advantage from unstructured, high-velocity log and machine Big Data

REAL-TIME OPERATIONAL INTELLIGENCE. Competitive advantage from unstructured, high-velocity log and machine Big Data REAL-TIME OPERATIONAL INTELLIGENCE Competitive advantage from unstructured, high-velocity log and machine Big Data 2 SQLstream: Our s-streaming products unlock the value of high-velocity unstructured log

More information

Data Refinery with Big Data Aspects

Data Refinery with Big Data Aspects International Journal of Information and Computation Technology. ISSN 0974-2239 Volume 3, Number 7 (2013), pp. 655-662 International Research Publications House http://www. irphouse.com /ijict.htm Data

More information

Big Data Analytics in Space Exploration and Entrepreneurship

Big Data Analytics in Space Exploration and Entrepreneurship Space Society of Silicon Valley Big Data Analytics in Space Exploration and Entrepreneurship Tiffani Crawford, PhD January 14, 2015 Big Data Analytics Data Characteristics Large quantities of many data

More information

Big Data, Physics, and the Industrial Internet! How Modeling & Analytics are Making the World Work Better."

Big Data, Physics, and the Industrial Internet! How Modeling & Analytics are Making the World Work Better. Big Data, Physics, and the Industrial Internet! How Modeling & Analytics are Making the World Work Better." Matt Denesuk! Chief Data Science Officer! GE Software! October 2014! Imagination at work. Contact:

More information

BIG DATA-AS-A-SERVICE

BIG DATA-AS-A-SERVICE White Paper BIG DATA-AS-A-SERVICE What Big Data is about What service providers can do with Big Data What EMC can do to help EMC Solutions Group Abstract This white paper looks at what service providers

More information

An Integrated Analytics & Big Data Infrastructure September 21, 2012 Robert Stackowiak, Vice President Data Systems Architecture Oracle Enterprise

An Integrated Analytics & Big Data Infrastructure September 21, 2012 Robert Stackowiak, Vice President Data Systems Architecture Oracle Enterprise An Integrated Analytics & Big Data Infrastructure September 21, 2012 Robert Stackowiak, Vice President Data Systems Architecture Oracle Enterprise Solutions Group The following is intended to outline our

More information

Architecting for the Internet of Things & Big Data

Architecting for the Internet of Things & Big Data Architecting for the Internet of Things & Big Data Robert Stackowiak, Oracle North America, VP Information Architecture & Big Data September 29, 2014 Safe Harbor Statement The following is intended to

More information

Hortonworks & SAS. Analytics everywhere. Page 1. Hortonworks Inc. 2011 2014. All Rights Reserved

Hortonworks & SAS. Analytics everywhere. Page 1. Hortonworks Inc. 2011 2014. All Rights Reserved Hortonworks & SAS Analytics everywhere. Page 1 A change in focus. A shift in Advertising From mass branding A shift in Financial Services From Educated Investing A shift in Healthcare From mass treatment

More information

Industrial Internet @GE. Dr. Stefan Bungart

Industrial Internet @GE. Dr. Stefan Bungart Industrial Internet @GE Dr. Stefan Bungart The vision is clear The real opportunity for change surpassing the magnitude of the consumer Internet is the Industrial Internet, an open, global network that

More information

News and trends in Data Warehouse Automation, Big Data and BI. Johan Hendrickx & Dirk Vermeiren

News and trends in Data Warehouse Automation, Big Data and BI. Johan Hendrickx & Dirk Vermeiren News and trends in Data Warehouse Automation, Big Data and BI Johan Hendrickx & Dirk Vermeiren Extreme Agility from Source to Analysis DWH Appliances & DWH Automation Typical Architecture 3 What Business

More information

Oracle Big Data Strategy Simplified Infrastrcuture

Oracle Big Data Strategy Simplified Infrastrcuture Big Data Oracle Big Data Strategy Simplified Infrastrcuture Selim Burduroğlu Global Innovation Evangelist & Architect Education & Research Industry Business Unit Oracle Confidential Internal/Restricted/Highly

More information

HP Vertica OnDemand. Vertica OnDemand. Enterprise-class Big Data analytics in the cloud. Enterprise-class Big Data analytics for any size organization

HP Vertica OnDemand. Vertica OnDemand. Enterprise-class Big Data analytics in the cloud. Enterprise-class Big Data analytics for any size organization Data sheet HP Vertica OnDemand Enterprise-class Big Data analytics in the cloud Enterprise-class Big Data analytics for any size organization Vertica OnDemand Organizations today are experiencing a greater

More information

Datenverwaltung im Wandel - Building an Enterprise Data Hub with

Datenverwaltung im Wandel - Building an Enterprise Data Hub with Datenverwaltung im Wandel - Building an Enterprise Data Hub with Cloudera Bernard Doering Regional Director, Central EMEA, Cloudera Cloudera Your Hadoop Experts Founded 2008, by former employees of Employees

More information

Microsoft Big Data Solutions. Anar Taghiyev P-TSP E-mail: b-anarta@microsoft.com;

Microsoft Big Data Solutions. Anar Taghiyev P-TSP E-mail: b-anarta@microsoft.com; Microsoft Big Data Solutions Anar Taghiyev P-TSP E-mail: b-anarta@microsoft.com; Why/What is Big Data and Why Microsoft? Options of storage and big data processing in Microsoft Azure. Real Impact of Big

More information

Big Data Executive Survey

Big Data Executive Survey Big Data Executive Full Questionnaire Big Date Executive Full Questionnaire Appendix B Questionnaire Welcome The survey has been designed to provide a benchmark for enterprises seeking to understand the

More information

APPROACHABLE ANALYTICS MAKING SENSE OF DATA

APPROACHABLE ANALYTICS MAKING SENSE OF DATA APPROACHABLE ANALYTICS MAKING SENSE OF DATA AGENDA SAS DELIVERS PROVEN SOLUTIONS THAT DRIVE INNOVATION AND IMPROVE PERFORMANCE. About SAS SAS Business Analytics Framework Approachable Analytics SAS for

More information

Big Data Challenges and Success Factors. Deloitte Analytics Your data, inside out

Big Data Challenges and Success Factors. Deloitte Analytics Your data, inside out Big Data Challenges and Success Factors Deloitte Analytics Your data, inside out Big Data refers to the set of problems and subsequent technologies developed to solve them that are hard or expensive to

More information

A New Era Of Analytic

A New Era Of Analytic Penang egovernment Seminar 2014 A New Era Of Analytic Megat Anuar Idris Head, Project Delivery, Business Analytics & Big Data Agenda Overview of Big Data Case Studies on Big Data Big Data Technology Readiness

More information

BIG DATA + ANALYTICS

BIG DATA + ANALYTICS An IDC InfoBrief for SAP and Intel + USING BIG DATA + ANALYTICS TO DRIVE BUSINESS TRANSFORMATION 1 In this Study Industry IDC recently conducted a survey sponsored by SAP and Intel to discover how organizations

More information

Data-intensive HPC: opportunities and challenges. Patrick Valduriez

Data-intensive HPC: opportunities and challenges. Patrick Valduriez Data-intensive HPC: opportunities and challenges Patrick Valduriez Big Data Landscape Multi-$billion market! Big data = Hadoop = MapReduce? No one-size-fits-all solution: SQL, NoSQL, MapReduce, No standard,

More information

BIG DATA What it is and how to use?

BIG DATA What it is and how to use? BIG DATA What it is and how to use? Lauri Ilison, PhD Data Scientist 21.11.2014 Big Data definition? There is no clear definition for BIG DATA BIG DATA is more of a concept than precise term 1 21.11.14

More information

Why big data? Lessons from a Decade+ Experiment in Big Data

Why big data? Lessons from a Decade+ Experiment in Big Data Why big data? Lessons from a Decade+ Experiment in Big Data David Belanger PhD Senior Research Fellow Stevens Institute of Technology dbelange@stevens.edu 1 What Does Big Look Like? 7 Image Source Page:

More information

Wrangler: A New Generation of Data-intensive Supercomputing. Christopher Jordan, Siva Kulasekaran, Niall Gaffney

Wrangler: A New Generation of Data-intensive Supercomputing. Christopher Jordan, Siva Kulasekaran, Niall Gaffney Wrangler: A New Generation of Data-intensive Supercomputing Christopher Jordan, Siva Kulasekaran, Niall Gaffney Project Partners Academic partners: TACC Primary system design, deployment, and operations

More information

Turn your information into a competitive advantage

Turn your information into a competitive advantage INDLÆG 03 Data Driven Business Value Turn your information into a competitive advantage Jonas Linders 04.10.2015 (dato) CGI Group Inc. 2015 Jonas Linders Education Role Industries M.Sc Informatics Experience

More information

Oracle Big Data SQL Technical Update

Oracle Big Data SQL Technical Update Oracle Big Data SQL Technical Update Jean-Pierre Dijcks Oracle Redwood City, CA, USA Keywords: Big Data, Hadoop, NoSQL Databases, Relational Databases, SQL, Security, Performance Introduction This technical

More information

What do Big Data & HAVEn mean? Robert Lejnert HP Autonomy

What do Big Data & HAVEn mean? Robert Lejnert HP Autonomy What do Big Data & HAVEn mean? Robert Lejnert HP Autonomy Much higher Volumes. Processed with more Velocity. With much more Variety. Is Big Data so big? Big Data Smart Data Project HAVEn: Adaptive Intelligence

More information

5 Keys to Unlocking the Big Data Analytics Puzzle. Anurag Tandon Director, Product Marketing March 26, 2014

5 Keys to Unlocking the Big Data Analytics Puzzle. Anurag Tandon Director, Product Marketing March 26, 2014 5 Keys to Unlocking the Big Data Analytics Puzzle Anurag Tandon Director, Product Marketing March 26, 2014 1 A Little About Us A global footprint. A proven innovator. A leader in enterprise analytics for

More information

Enterprise Application Enablement for the Internet of Things

Enterprise Application Enablement for the Internet of Things Enterprise Application Enablement for the Internet of Things Prof. Dr. Uwe Kubach VP Internet of Things Platform, P&I Technology, SAP SE Public Internet of Things (IoT) Trends 12 50 bn 40 50 % Devices

More information

The Future of Data Management

The Future of Data Management The Future of Data Management with Hadoop and the Enterprise Data Hub Amr Awadallah (@awadallah) Cofounder and CTO Cloudera Snapshot Founded 2008, by former employees of Employees Today ~ 800 World Class

More information

Aligning Your Strategic Initiatives with a Realistic Big Data Analytics Roadmap

Aligning Your Strategic Initiatives with a Realistic Big Data Analytics Roadmap Aligning Your Strategic Initiatives with a Realistic Big Data Analytics Roadmap 3 key strategic advantages, and a realistic roadmap for what you really need, and when 2012, Cognizant Topics to be discussed

More information

Oracle Database - Engineered for Innovation. Sedat Zencirci Teknoloji Satış Danışmanlığı Direktörü Türkiye ve Orta Asya

Oracle Database - Engineered for Innovation. Sedat Zencirci Teknoloji Satış Danışmanlığı Direktörü Türkiye ve Orta Asya Oracle Database - Engineered for Innovation Sedat Zencirci Teknoloji Satış Danışmanlığı Direktörü Türkiye ve Orta Asya Oracle Database 11g Release 2 Shipping since September 2009 11.2.0.3 Patch Set now

More information

Big Data Analytics. Prof. Dr. Lars Schmidt-Thieme

Big Data Analytics. Prof. Dr. Lars Schmidt-Thieme Big Data Analytics Prof. Dr. Lars Schmidt-Thieme Information Systems and Machine Learning Lab (ISMLL) Institute of Computer Science University of Hildesheim, Germany 33. Sitzung des Arbeitskreises Informationstechnologie,

More information

Is a Data Scientist the New Quant? Stuart Kozola MathWorks

Is a Data Scientist the New Quant? Stuart Kozola MathWorks Is a Data Scientist the New Quant? Stuart Kozola MathWorks 2015 The MathWorks, Inc. 1 Facts or information used usually to calculate, analyze, or plan something Information that is produced or stored by

More information

HDP Hadoop From concept to deployment.

HDP Hadoop From concept to deployment. HDP Hadoop From concept to deployment. Ankur Gupta Senior Solutions Engineer Rackspace: Page 41 27 th Jan 2015 Where are you in your Hadoop Journey? A. Researching our options B. Currently evaluating some

More information

Luncheon Webinar Series May 13, 2013

Luncheon Webinar Series May 13, 2013 Luncheon Webinar Series May 13, 2013 InfoSphere DataStage is Big Data Integration Sponsored By: Presented by : Tony Curcio, InfoSphere Product Management 0 InfoSphere DataStage is Big Data Integration

More information

The IoT Inc Business Meetup Silicon Valley

The IoT Inc Business Meetup Silicon Valley The IoT Inc Business Meetup Silicon Valley Meeting 6 February 2015 Bruce Sinclair (Organizer): bruce@iot-inc.com Target of Meetup For business people selling products and services into IoT but of course

More information

Big Data and Cloud Computing for GHRSST

Big Data and Cloud Computing for GHRSST Big Data and Cloud Computing for GHRSST Jean-Francois Piollé (jfpiolle@ifremer.fr) Frédéric Paul, Olivier Archer CERSAT / Institut Français de Recherche pour l Exploitation de la Mer Facing data deluge

More information

BIG DATA & ANALYTICS. Transforming the business and driving revenue through big data and analytics

BIG DATA & ANALYTICS. Transforming the business and driving revenue through big data and analytics BIG DATA & ANALYTICS Transforming the business and driving revenue through big data and analytics Collection, storage and extraction of business value from data generated from a variety of sources are

More information

Big Data Are You Ready? Thomas Kyte http://asktom.oracle.com

Big Data Are You Ready? Thomas Kyte http://asktom.oracle.com Big Data Are You Ready? Thomas Kyte http://asktom.oracle.com The following is intended to outline our general product direction. It is intended for information purposes only, and may not be incorporated

More information

ANALYTICS IN BIG DATA ERA

ANALYTICS IN BIG DATA ERA ANALYTICS IN BIG DATA ERA ANALYTICS TECHNOLOGY AND ARCHITECTURE TO MANAGE VELOCITY AND VARIETY, DISCOVER RELATIONSHIPS AND CLASSIFY HUGE AMOUNT OF DATA MAURIZIO SALUSTI SAS Copyr i g ht 2012, SAS Ins titut

More information

The Internet of Things

The Internet of Things The Internet of Things The Power of Actionable Insight An introduction to the Internet of Things Chris Vetor Business Unit Executive, WW Programs cvetor@us.ibm.com More and more of the world s activity

More information

The future of Big Data A United Hitachi View

The future of Big Data A United Hitachi View The future of Big Data A United Hitachi View Alex van Die Pre-Sales Consultant 1 Oktober 2014 1 Agenda Evolutie van Data en Analytics Internet of Things Hitachi Social Innovation Vision and Solutions 2

More information

Big Data Big Deal for Public Sector Organizations

Big Data Big Deal for Public Sector Organizations Big Data Big Deal for Public Sector Organizations Hoàng Xuân Hiếu Director, FAB & Government Business Indochina & Myanmar 1 Copyright 2013, Oracle and/or its affiliates. All rights reserved. The following

More information

Surfing the Data Tsunami: A New Paradigm for Big Data Processing and Analytics

Surfing the Data Tsunami: A New Paradigm for Big Data Processing and Analytics Surfing the Data Tsunami: A New Paradigm for Big Data Processing and Analytics Dr. Liangxiu Han Future Networks and Distributed Systems Group (FUNDS) School of Computing, Mathematics and Digital Technology,

More information

The distribution of marine OpenData via distributed data networks and Web APIs. The example of ERDDAP, the message broker and data mediator from NOAA

The distribution of marine OpenData via distributed data networks and Web APIs. The example of ERDDAP, the message broker and data mediator from NOAA The distribution of marine OpenData via distributed data networks and Web APIs. The example of ERDDAP, the message broker and data mediator from NOAA Dr. Conor Delaney 9 April 2014 GeoMaritime, London

More information

Smart Cities Solution Overview Innovation Center Network, Research & Innovation. SAP SE Reiner Bildmayer

Smart Cities Solution Overview Innovation Center Network, Research & Innovation. SAP SE Reiner Bildmayer Smart Cities Solution Overview Innovation Center Network, Research & Innovation SAP SE Reiner Bildmayer Why Cities need to be Run Better Challenges and Opportunities ~50% of the world s population currently

More information

HADOOP SOLUTION USING EMC ISILON AND CLOUDERA ENTERPRISE Efficient, Flexible In-Place Hadoop Analytics

HADOOP SOLUTION USING EMC ISILON AND CLOUDERA ENTERPRISE Efficient, Flexible In-Place Hadoop Analytics HADOOP SOLUTION USING EMC ISILON AND CLOUDERA ENTERPRISE Efficient, Flexible In-Place Hadoop Analytics ESSENTIALS EMC ISILON Use the industry's first and only scale-out NAS solution with native Hadoop

More information

Ganzheitliches Datenmanagement

Ganzheitliches Datenmanagement Ganzheitliches Datenmanagement für Hadoop Michael Kohs, Senior Sales Consultant @mikchaos The Problem with Big Data Projects in 2016 Relational, Mainframe Documents and Emails Data Modeler Data Scientist

More information

Unlocking the Intelligence in. Big Data. Ron Kasabian General Manager Big Data Solutions Intel Corporation

Unlocking the Intelligence in. Big Data. Ron Kasabian General Manager Big Data Solutions Intel Corporation Unlocking the Intelligence in Big Data Ron Kasabian General Manager Big Data Solutions Intel Corporation Volume & Type of Data What s Driving Big Data? 10X Data growth by 2016 90% unstructured 1 Lower

More information

Investor Presentation. Second Quarter 2015

Investor Presentation. Second Quarter 2015 Investor Presentation Second Quarter 2015 Note to Investors Certain non-gaap financial information regarding operating results may be discussed during this presentation. Reconciliations of the differences

More information

Surak Thammarak. Advisory Systems Engineer EMC. surak.thammarak@emc.com +668-1700-6333

Surak Thammarak. Advisory Systems Engineer EMC. surak.thammarak@emc.com +668-1700-6333 Surak Thammarak Advisory Systems Engineer EMC surak.thammarak@emc.com +668-1700-6333 1 2 ?? Today s Life?? - ส งคมก มหน า 3 4 Gartner and IDC Mobile Platform 5 The DIGITAL UNIVERSE of OPPORTUNITIES 6 Digital

More information

Are You Ready for Big Data?

Are You Ready for Big Data? Are You Ready for Big Data? Jim Gallo National Director, Business Analytics April 10, 2013 Agenda What is Big Data? How do you leverage Big Data in your company? How do you prepare for a Big Data initiative?

More information

WINDOWS AZURE DATA MANAGEMENT AND BUSINESS ANALYTICS

WINDOWS AZURE DATA MANAGEMENT AND BUSINESS ANALYTICS WINDOWS AZURE DATA MANAGEMENT AND BUSINESS ANALYTICS Managing and analyzing data in the cloud is just as important as it is anywhere else. To let you do this, Windows Azure provides a range of technologies

More information

2015 Analyst and Advisor Summit. Advanced Data Analytics Dr. Rod Fontecilla Vice President, Application Services, Chief Data Scientist

2015 Analyst and Advisor Summit. Advanced Data Analytics Dr. Rod Fontecilla Vice President, Application Services, Chief Data Scientist 2015 Analyst and Advisor Summit Advanced Data Analytics Dr. Rod Fontecilla Vice President, Application Services, Chief Data Scientist Agenda Key Facts Offerings and Capabilities Case Studies When to Engage

More information

INVENTING THE FUTURE HITACHI DATA SYSTEMS BIG DATA ROADMAP MICHAEL HAY

INVENTING THE FUTURE HITACHI DATA SYSTEMS BIG DATA ROADMAP MICHAEL HAY INVENTING THE FUTURE HITACHI DATA SYSTEMS BIG DATA ROADMAP MICHAEL HAY CTO AND VP, GLOBAL SOLUTIONS STRATEGY AND DEVELOPMENT CHIEF ENGINEER, INTEGRATED PLATFORM STRATEGY @ ITPD WEBTECH EDUCATIONAL SERIES

More information

Transforming the Telecoms Business using Big Data and Analytics

Transforming the Telecoms Business using Big Data and Analytics Transforming the Telecoms Business using Big Data and Analytics Event: ICT Forum for HR Professionals Venue: Meikles Hotel, Harare, Zimbabwe Date: 19 th 21 st August 2015 AFRALTI 1 Objectives Describe

More information

EOFS Workshop Paris Sept, 2011. Lustre at exascale. Eric Barton. CTO Whamcloud, Inc. eeb@whamcloud.com. 2011 Whamcloud, Inc.

EOFS Workshop Paris Sept, 2011. Lustre at exascale. Eric Barton. CTO Whamcloud, Inc. eeb@whamcloud.com. 2011 Whamcloud, Inc. EOFS Workshop Paris Sept, 2011 Lustre at exascale Eric Barton CTO Whamcloud, Inc. eeb@whamcloud.com Agenda Forces at work in exascale I/O Technology drivers I/O requirements Software engineering issues

More information

Big Data Keep it Simple

Big Data Keep it Simple RSM Leadership Summit Big Data Keep it Simple Rotterdam, October 3 rd 2014 Jens-Peter Seick, VP Head of Product Management and Development Fujitsu in Europe 0 Agenda Big Data phenomena Big Data technologies

More information

Big Data Storage Challenges for the Industrial Internet of Things

Big Data Storage Challenges for the Industrial Internet of Things Big Data Storage Challenges for the Industrial Internet of Things Shyam V Nath Diwakar Kasibhotla SDC September, 2014 Agenda Introduction to IoT and Industrial Internet Industrial & Sensor Data Big Data

More information

Johan Hallberg Research Manager / Industry Analyst IDC Nordic Services & Sourcing Digital Transformation Global CIO Agenda

Johan Hallberg Research Manager / Industry Analyst IDC Nordic Services & Sourcing Digital Transformation Global CIO Agenda IDC s Big Data Predictions 2015 Johan Hallberg Research Manager / Industry Analyst IDC Nordic Services & Sourcing Digital Transformation Global CIO Agenda Big Data Opportunity: The Need for Deep Personalization

More information

NIST Big Data PWG & RDA Big Data Infrastructure WG: Implementation Strategy: Best Practice Guideline for Big Data Application Development

NIST Big Data PWG & RDA Big Data Infrastructure WG: Implementation Strategy: Best Practice Guideline for Big Data Application Development NIST Big Data PWG & RDA Big Data Infrastructure WG: Implementation Strategy: Best Practice Guideline for Big Data Application Development Wo Chang wchang@nist.gov National Institute of Standards and Technology

More information

How to use Big Data in Industry 4.0 implementations. LAURI ILISON, PhD Head of Big Data and Machine Learning

How to use Big Data in Industry 4.0 implementations. LAURI ILISON, PhD Head of Big Data and Machine Learning How to use Big Data in Industry 4.0 implementations LAURI ILISON, PhD Head of Big Data and Machine Learning Big Data definition? Big Data is about structured vs unstructured data Big Data is about Volume

More information

The Challenge of Big Data Benchmarking Large-Scale Data Management Insights from Benchmark Research

The Challenge of Big Data Benchmarking Large-Scale Data Management Insights from Benchmark Research Benchmarking Large-Scale Data Management Insights from Presentation Confidentiality Statement The materials in this presentation are protected under the confidential agreement and/or are copyrighted materials

More information

Chapter 7. Using Hadoop Cluster and MapReduce

Chapter 7. Using Hadoop Cluster and MapReduce Chapter 7 Using Hadoop Cluster and MapReduce Modeling and Prototyping of RMS for QoS Oriented Grid Page 152 7. Using Hadoop Cluster and MapReduce for Big Data Problems The size of the databases used in

More information

Informix The Intelligent Database for IoT

Informix The Intelligent Database for IoT Informix The Intelligent Database for IoT Kiran Challapalli Informix Competitive Technology & Enablement challapalli@in.ibm.com +91-80431-91802 Agenda What is Internet of Things (IoT) Why it matters IoT

More information

Hybrid Cloud Architectures for Operational Performance Management

Hybrid Cloud Architectures for Operational Performance Management Hybrid Cloud Architectures for Operational Performance Management Delbert Murphy Solution Architect / Data Scientist Microsoft Corporation GPDIS_2014.ppt 1 Delbert Murphy and Microsoft s Data Insights

More information

BIG DATA AND THE ENTERPRISE DATA WAREHOUSE WORKSHOP

BIG DATA AND THE ENTERPRISE DATA WAREHOUSE WORKSHOP BIG DATA AND THE ENTERPRISE DATA WAREHOUSE WORKSHOP Business Analytics for All Amsterdam - 2015 Value of Big Data is Being Recognized Executives beginning to see the path from data insights to revenue

More information

Enabling the SmartGrid through Cloud Computing

Enabling the SmartGrid through Cloud Computing Enabling the SmartGrid through Cloud Computing April 2012 Creating Value, Delivering Results 2012 eglobaltech Incorporated. Tech, Inc. All rights reserved. 1 Overall Objective To deliver electricity from

More information

Bringing Big Data Modelling into the Hands of Domain Experts

Bringing Big Data Modelling into the Hands of Domain Experts Bringing Big Data Modelling into the Hands of Domain Experts David Willingham Senior Application Engineer MathWorks david.willingham@mathworks.com.au 2015 The MathWorks, Inc. 1 Data is the sword of the

More information

Big Data is Changing Business

Big Data is Changing Business Big Data is Changing Business Opportunities and Challenges The Regent Hotel, Beijing, China 20 th November, 2012 Presented by: Craig Stires, Research Director, IDC Asia/Pacific @Craig_IDC Copyright IDC.

More information

BIG DATA & SOCIAL INNOVATION KENNETH THOMAS, CLIENT MANAGER

BIG DATA & SOCIAL INNOVATION KENNETH THOMAS, CLIENT MANAGER BIG DATA & SOCIAL INNOVATION KENNETH THOMAS, CLIENT MANAGER 1 MAKING THE RIGHT DECISSION AT THE RIGHT PLACE AT THE RIGHT TIME 2 THE DATA MULTIPLIER EFFECT AT WORK BUSINESS DRIVEN HUMAN DRIVEN MACHINE DRIVEN

More information

Big Data Analytics. with EMC Greenplum and Hadoop. Big Data Analytics. Ofir Manor Pre Sales Technical Architect EMC Greenplum

Big Data Analytics. with EMC Greenplum and Hadoop. Big Data Analytics. Ofir Manor Pre Sales Technical Architect EMC Greenplum Big Data Analytics with EMC Greenplum and Hadoop Big Data Analytics with EMC Greenplum and Hadoop Ofir Manor Pre Sales Technical Architect EMC Greenplum 1 Big Data and the Data Warehouse Potential All

More information

VIEWPOINT. High Performance Analytics. Industry Context and Trends

VIEWPOINT. High Performance Analytics. Industry Context and Trends VIEWPOINT High Performance Analytics Industry Context and Trends In the digital age of social media and connected devices, enterprises have a plethora of data that they can mine, to discover hidden correlations

More information

Geospatial Technology Innovations and Convergence

Geospatial Technology Innovations and Convergence Geospatial Technology Innovations and Convergence Processing Big and Fast Data: Best with a Multi-Model Database Steven Hagan Vice President Oracle Database Server Technologies August, 2015 Data Volume

More information

Oracle s Big Data solutions. Roger Wullschleger.

Oracle s Big Data solutions. Roger Wullschleger. <Insert Picture Here> s Big Data solutions Roger Wullschleger DBTA Workshop on Big Data, Cloud Data Management and NoSQL 10. October 2012, Stade de Suisse, Berne 1 The following is intended to outline

More information

Big Data and Analytics: Getting Started with ArcGIS. Mike Park Erik Hoel

Big Data and Analytics: Getting Started with ArcGIS. Mike Park Erik Hoel Big Data and Analytics: Getting Started with ArcGIS Mike Park Erik Hoel Agenda Overview of big data Distributed computation User experience Data management Big data What is it? Big Data is a loosely defined

More information

Integrating Cloudera and SAP HANA

Integrating Cloudera and SAP HANA Integrating Cloudera and SAP HANA Version: 103 Table of Contents Introduction/Executive Summary 4 Overview of Cloudera Enterprise 4 Data Access 5 Apache Hive 5 Data Processing 5 Data Integration 5 Partner

More information