Aeromancer: A Workflow Manager for Large- Scale MapReduce-Based Scientific Workflows

Size: px

Start display at page:

Download "Aeromancer: A Workflow Manager for Large- Scale MapReduce-Based Scientific Workflows"

Roger McBride
6 years ago
Views:

1 Aeromancer: A Workflow Manager for Large- Scale MapReduce-Based Scientific Workflows Presented by Sarunya Pumma Supervisors: Dr. Wu-chun Feng, Dr. Mark Gardner, and Dr. Hao Wang synergy.cs.vt.edu

2 Outline Introduction & Motivation Cloud computing and its benefits SeqInCloud and scalability analysis Aeromancer Usage scenario Workflow and architecture Discussion Experiments & Initial Results

3 Introduction & Motivation

4 What is Cloud?

5 I would like 6 virtual machines for a month I would like 3 virtual machines for 2 months What is Cloud? user a user b vcpus vmemory vstorage vnetwork virtual resources : resource pool CPUs Memory Storage Network physical resources

6 How is Cloud beneficial for Us?

7 How is Cloud beneficial for Us? Next-generation sequencing (NGS) produces a huge amount of DNA sequence data Cost of sequencing a human-size genome decreases over time { 2k sequencers produce 15 PB } { 2011 $95M } genetic data/year 2012 $6.5k 2014 $1k

8 How is Cloud beneficial for Us? Big Data

9 How is Cloud beneficial for Us? Cloud can accelerate our ability to identify the cure for cancer

10 Big Data Challenge Provisioning compute and storage resources to analyze the exponentially increasing genomic data

11 SeqInCloud It is implemented based on GATK (Genome Analysis Toolkit) All six stages have to be executed in order Each stage is implemented as a small separated Hadoop program It can be broken down into small jobs and run on multiple computers/processors We expect a linear speed-up on every stage

12 SeqInCloud Genome variant analysis pipeline FASTQ file Align (BWA)... Align (BWA) Count Covariate... Count Covariate 3 Sort... Sort Merge Covariates 1 Index Local- Realignment & Sort Index Local- Realignment & Sort Table Recalibration Genotyper & Variant Filter Table Recalibration Genotyper & Variant Filter MarkDuplicate MarkDuplicate Combine Variants Merge BAM File 6 VCF file BAM file

13 SeqInCloud Scalability Analysis 30 1 node 2 nodes 4 nodes 8 nodes 16 nodes 32 nodes Speed up Stage 0 Stage 1 Stage 2 Stage 3 Stage 4 Stage 5 Stage 6

14 SeqInCloud Scalability Analysis To run SeqInCloud efficiently on Hadoop Number of compute nodes must be varied based on the scalability of each stage Challenges 1. How many nodes do we need for each stage? 2. How to rapidly adjust a number of nodes with negligible overhead? An automatic Hadoop workflow execution is needed!

15 Hadoop Workflow (SeqInCloud) Hadoop Workflow Manager Cloud

16 Aeromancer Cloud Hadoop Workflow (SeqInCloud)

17 Aeromancer Platform for executing the scientific Hadoop workflows Based on Hadoop MapReduce HDFS Build on top of CloudGene Workflow (DAG) Node/Stage Edge (represent dependency between two stages)

18 Aeromancer Usage Scenario 1 Add a new application to Aeromancer

19 Aeromancer Usage Scenario Application details Public cloud details Execution details - Execution command - Stage dependencies - Data dependencies YAML file

20 Aeromancer Usage Scenario 1 Add a new application to Aeromancer 2 Submit a new job

21 Aeromancer Usage Scenario User Portal

22 Aeromancer Usage Scenario 1 Add a new application to Aeromancer 2 Submit a new job 3 Wait for the execution to complete

23 Aeromancer Workflow & Architecture Job Queue Spawn Hadoop DAG Processor Step Queue Aeromancer Query Spawn Aeromancer Server Worker Thread Enqueue Sub- Worker Thread Templeton Hadoop Master Worker Hadoop Worker Hadoop Worker Cluster

24 Aeromancer Discussion Currently, Aeromancer worker nodes are virtual machines in the cloud Pro Con Number of worker nodes are flexibly adjustable upon demand The overhead of creating/deleting worker nodes is high Execution speed drops significantly Containerization may be able to improve Aeromancer s performance

25 Virtualization Containerization Reference:

26 Experiment Platforms Bare metal machines KVM Docker Benchmarks (HiBench) Sort WordCount TeraSort Performance Speed (execution time) HDFS throughput Nutch Indexing PageRank EnhancedDFSIO

27 Experiment: Hadoop on Bare Metal Environment DataNode NodeManager ApplicationHistoryServer SecondaryNameNode JobHistoryServer AmbariServer NameNode ResourceManager eth0 MiCloud3 (Slave Machine) DataNode NodeManager MiCloud2 (Master Machine) MiCloud5 (Slave Machine)

28 Experiment: Hadoop on Virtual Machine Environment br0 NodeManager DataNode AmbariServer ApplicationHistoryServer SecondaryNameNode JobHistoryServer br0 eth0 MiCloud3 (Slave VM) NameNode ResourceManager Micloud2 (Master VM) eth0: x/24 br0: x/24 br0 12 vcpus 16 GB-vMemory NodeManager DataNode MiCloud5 (Slave VM)

29 Experiment: Hadoop on Docker br0 NodeManager DataNode AmbariServer ApplicationHistoryServer SecondaryNameNode JobHistoryServer br0 eth0 MiCloud3 (Slave Container) NameNode ResourceManager Micloud2 (Master Container) eth0: x/24 br0: x/24 br0 12 vcpus 16 GB-vMemory NodeManager DataNode MiCloud5 (Slave Container)

30 Initial Results: Slow Down KVM Docker 100% 75% 95.4% 98.6% 92% 88% 97.4% 83% 94.8% 50% 62% 59.5% 56.8% 25% 0% 14% Sort Wordcount DFSIOE NutchIndexing TeraSort PageRank 9%

31 Future Work Automatic resources provisioning Determining the number of computing nodes needed for each stage Based on the scalability of each stage Based on user s constraints (budget/expected response time) Automatic task scheduling Determining the optimal execution location for each stage (either client or cloud)

32 Thank you 32 synergy.cs.vt.edu

Big Data for Engineers Spring Resource Management

Ghislain Fourny Big Data for Engineers Spring 2018 7. Resource Management artjazz / 123RF Stock Photo Data Technology Stack User interfaces Querying Data stores Indexing Processing Validation Data models