Chapter 8 Fault Tolerance

Similar documents
Chapter 8 Fault Tolerance

CprE Fault Tolerance. Dr. Yong Guan. Department of Electrical and Computer Engineering & Information Assurance Center Iowa State University

Fault Tolerance. Basic Concepts

CSE 5306 Distributed Systems. Fault Tolerance

CSE 5306 Distributed Systems

MYE017 Distributed Systems. Kostas Magoutis

MYE017 Distributed Systems. Kostas Magoutis

Fault Tolerance. Distributed Software Systems. Definitions

Fault Tolerance. Chapter 7

Fault Tolerance. Distributed Systems. September 2002

Fault Tolerance Part II. CS403/534 Distributed Systems Erkay Savas Sabanci University

Chapter 5: Distributed Systems: Fault Tolerance. Fall 2013 Jussi Kangasharju

Distributed Systems Fault Tolerance

Failure Tolerance. Distributed Systems Santa Clara University

Distributed Systems Principles and Paradigms. Chapter 08: Fault Tolerance

Today: Fault Tolerance. Failure Masking by Redundancy

Fault Tolerance. Distributed Systems IT332

Today: Fault Tolerance. Reliable One-One Communication

Fault Tolerance 1/64

Fault Tolerance. Fall 2008 Jussi Kangasharju

Failure Models. Fault Tolerance. Failure Masking by Redundancy. Agreement in Faulty Systems

Today: Fault Tolerance. Replica Management

Distributed Systems Principles and Paradigms

Distributed Systems Principles and Paradigms

Distributed Systems Principles and Paradigms

Dep. Systems Requirements

Today: Fault Tolerance. Fault Tolerance

Distributed Systems Reliable Group Communication

Fault Tolerance Part I. CS403/534 Distributed Systems Erkay Savas Sabanci University

Today: Fault Tolerance

Module 8 - Fault Tolerance

Module 8 Fault Tolerance CS655! 8-1!

Basic concepts in fault tolerance Masking failure by redundancy Process resilience Reliable communication. Distributed commit.

Distributed Systems. 09. State Machine Replication & Virtual Synchrony. Paul Krzyzanowski. Rutgers University. Fall Paul Krzyzanowski

Distributed Systems COMP 212. Lecture 19 Othon Michail

DISTRIBUTED SYSTEMS. Second Edition. Andrew S. Tanenbaum Maarten Van Steen. Vrije Universiteit Amsterdam, 7'he Netherlands PEARSON.

Distributed Systems (ICE 601) Fault Tolerance

Distributed Systems

Networked Systems and Services, Fall 2018 Chapter 4. Jussi Kangasharju Markku Kojo Lea Kutvonen

TWO-PHASE COMMIT ATTRIBUTION 5/11/2018. George Porter May 9 and 11, 2018

Distributed Systems. Fault Tolerance. Paul Krzyzanowski

PRIMARY-BACKUP REPLICATION

Distributed Systems. 19. Fault Tolerance Paul Krzyzanowski. Rutgers University. Fall 2013

CS 470 Spring Fault Tolerance. Mike Lam, Professor. Content taken from the following:

Recovering from a Crash. Three-Phase Commit

CA464 Distributed Programming

2017 Paul Krzyzanowski 1

Chapter 6 Synchronization

Fault Tolerance. o Basic Concepts o Process Resilience o Reliable Client-Server Communication o Reliable Group Communication. o Distributed Commit

Fault Tolerance Dealing with an imperfect world

Verteilte Systeme/Distributed Systems Ch. 5: Various distributed algorithms


Distributed Systems 11. Consensus. Paul Krzyzanowski

Viewstamped Replication to Practical Byzantine Fault Tolerance. Pradipta De

Today CSCI Recovery techniques. Recovery. Recovery CAP Theorem. Instructor: Abhishek Chandra

Practical Byzantine Fault Tolerance

Distributed Systems. Characteristics of Distributed Systems. Lecture Notes 1 Basic Concepts. Operating Systems. Anand Tripathi

Distributed Systems. Characteristics of Distributed Systems. Characteristics of Distributed Systems. Goals in Distributed System Designs

Process groups and message ordering

Distributed Deadlock

Distributed Systems Principles and Paradigms. Chapter 01: Introduction. Contents. Distributed System: Definition.

Consensus and related problems

Chapter 7 Consistency And Replication

Three Models. 1. Time Order 2. Distributed Algorithms 3. Nature of Distributed Systems1. DEPT. OF Comp Sc. and Engg., IIT Delhi

Distributed Systems Principles and Paradigms. Chapter 01: Introduction

Distributed Systems 24. Fault Tolerance

Middleware for Transparent TCP Connection Migration

Parallel and Distributed Systems. Programming Models. Why Parallel or Distributed Computing? What is a parallel computer?

Distributed Systems Principles and Paradigms

To do. Consensus and related problems. q Failure. q Raft

Rollback-Recovery p Σ Σ

Byzantine Fault Tolerance and Consensus. Adi Seredinschi Distributed Programming Laboratory

DISTRIBUTED COMPUTER SYSTEMS

Consensus in Distributed Systems. Jeff Chase Duke University

Distributed Algorithms Models

Distributed Systems 23. Fault Tolerance

MODELS OF DISTRIBUTED SYSTEMS

Network Protocols. Sarah Diesburg Operating Systems CS 3430

Basic vs. Reliable Multicast

CSE 486/586 Distributed Systems Reliable Multicast --- 1

Chapter 4 Communication

SYNCHRONIZATION. DISTRIBUTED SYSTEMS Principles and Paradigms. Second Edition. Chapter 6 ANDREW S. TANENBAUM MAARTEN VAN STEEN

Distributed systems. Lecture 6: distributed transactions, elections, consensus and replication. Malte Schwarzkopf

FAULT TOLERANCE. Fault Tolerant Systems. Faults Faults (cont d)

Beyond FLP. Acknowledgement for presentation material. Chapter 8: Distributed Systems Principles and Paradigms: Tanenbaum and Van Steen

Chapter 10 DISTRIBUTED OBJECT-BASED SYSTEMS

Chapter 11 DISTRIBUTED FILE SYSTEMS

Clock Synchronization. Synchronization. Clock Synchronization Algorithms. Physical Clock Synchronization. Tanenbaum Chapter 6 plus additional papers

Fault-Tolerant Computer Systems ECE 60872/CS Recovery

Exam 2 Review. October 29, Paul Krzyzanowski 1

CS 455: INTRODUCTION TO DISTRIBUTED SYSTEMS [DISTRIBUTED MUTUAL EXCLUSION] Frequently asked questions from the previous class survey

DISTRIBUTED SYSTEMS Principles and Paradigms Second Edition ANDREW S. TANENBAUM MAARTEN VAN STEEN. Chapter 7 Consistency And Replication

Introduction to Distributed Systems

CHAPTER 4: INTERPROCESS COMMUNICATION AND COORDINATION

CS455: Introduction to Distributed Systems [Spring 2018] Dept. Of Computer Science, Colorado State University

Reliable Distributed System Approaches

Distributed Algorithms Reliable Broadcast

Specifying and Proving Broadcast Properties with TLA

MODELS OF DISTRIBUTED SYSTEMS

Practical Byzantine Fault

Transcription:

DISTRIBUTED SYSTEMS Principles and Paradigms Second Edition ANDREW S. TANENBAUM MAARTEN VAN STEEN Chapter 8 Fault Tolerance

Fault Tolerance Basic Concepts Being fault tolerant is strongly related to what are called dependable systems Dependability implies the following: 1. Availability 2. Reliability 3. Safety 4. Maintainability

Failure Models Figure 8-1. Different types of failures.

Failure Masking by Redundancy Figure 8-2. Triple modular redundancy.

Flat Groups versus Hierarchical Groups Figure 8-3. (a) Communication in a flat group. (b) Communication in a simple hierarchical group.

Agreement in Faulty Systems (1) Possible cases: 1. Synchronous versus asynchronous systems. 2. Communication delay is bounded or not. 3. Message delivery is ordered or not. 4. Message transmission is done through unicasting or multicasting.

Agreement in Faulty Systems (2) Figure 8-4. Circumstances under which distributed agreement can be reached.

Agreement in Faulty Systems (3) Figure 8-5. The Byzantine agreement problem for three nonfaulty and one faulty process. (a) Each process sends their value to the others.

Agreement in Faulty Systems (4) Figure 8-5. The Byzantine agreement problem for three nonfaulty and one faulty process. (b) The vectors that each process assembles based on (a). (c) The vectors that each process receives in step 3.

Agreement in Faulty Systems (5) Figure 8-6. The same as Fig. 8-5, except now with two correct process and one faulty process.

RPC Semantics in the Presence of Failures Five different classes of failures that can occur in RPC systems: 1. The client is unable to locate the server. 2. The request message from the client to the server is lost. 3. The server crashes after receiving a request. 4. The reply message from the server to the client is lost. 5. The client crashes after sending a request.

Server Crashes (1) Figure 8-7. A server in client-server communication. (a) The normal case. (b) Crash after execution. (c) Crash before execution.

Server Crashes (2) Three events that can happen at the server: Send the completion message (M), Print the text (P), Crash (C).

Server Crashes (3) These events can occur in six different orderings: 1. M P C: A crash occurs after sending the completion message and printing the text. 2. M C ( P): A crash happens after sending the completion message, but before the text could be printed. 3. P M C: A crash occurs after sending the completion message and printing the text. 4. P C( M): The text printed, after which a crash occurs before the completion message could be sent. 5. C ( P M): A crash happens before the server could do anything. 6. C ( M P): A crash happens before the server could do anything.

Server Crashes (4) Figure 8-8. Different combinations of client and server strategies in the presence of server crashes.

Basic Reliable-Multicasting Schemes Figure 8-9. A simple solution to reliable multicasting when all receivers are known and are assumed not to fail. (a) Message transmission. (b) Reporting feedback.

Nonhierarchical Feedback Control Figure 8-10. Several receivers have scheduled a request for retransmission, but the first retransmission request leads to the suppression of others.

Hierarchical Feedback Control Figure 8-11. The essence of hierarchical reliable multicasting. Each local coordinator forwards the message to its children and later handles retransmission requests.

Virtual Synchrony (1) Figure 8-12. The logical organization of a distributed system to distinguish between message receipt and message delivery.

Virtual Synchrony (2) Figure 8-13. The principle of virtual synchronous multicast.

Message Ordering (1) Four different orderings are distinguished: Unordered multicasts FIFO-ordered multicasts Causally-ordered multicasts Totally-ordered multicasts

Message Ordering (2) Figure 8-14. Three communicating processes in the same group. The ordering of events per process is shown along the vertical axis.

Message Ordering (3) Figure 8-15. Four processes in the same group with two different senders, and a possible delivery order of messages under FIFO-ordered multicasting

Implementing Virtual Synchrony (1) Figure 8-16. Six different versions of virtually synchronous reliable multicasting.

Implementing Virtual Synchrony (2) Figure 8-17. (a) Process 4 notices that process 7 has crashed and sends a view change.

Implementing Virtual Synchrony (3) Figure 8-17. (b) Process 6 sends out all its unstable messages, followed by a flush message.

Implementing Virtual Synchrony (4) Figure 8-17. (c) Process 6 installs the new view when it has received a flush message from everyone else.

Two-Phase Commit (1) Figure 8-18. (a) The finite state machine for the coordinator in 2PC. (b) The finite state machine for a participant.

Two-Phase Commit (2) Figure 8-19. Actions taken by a participant P when residing in state READY and having contacted another participant Q.

Two-Phase Commit (3)... Figure 8-20. Outline of the steps taken by the coordinator in a two-phase commit protocol.

Two-Phase Commit (4)... Figure 8-20. Outline of the steps taken by the coordinator in a two-phase commit protocol.

Two-Phase Commit (5) Figure 8-21. (a) The steps taken by a participant process in 2PC.

Two-Phase Commit (7) Figure 8-21. (b) The steps for handling incoming decision requests..

Three-Phase Commit (1) The states of the coordinator and each participant satisfy the following two conditions: 1. There is no single state from which it is possible to make a transition directly to either a COMMIT or an ABORT state. 2. There is no state in which it is not possible to make a final decision, and from which a transition to a COMMIT state can be made.

Three-Phase Commit (2) Figure 8-22. (a) The finite state machine for the coordinator in 3PC. (b) The finite state machine for a participant.

Recovery Stable Storage Figure 8-23. (a) Stable storage. (b) Crash after drive 1 is updated. (c) Bad spot.

Checkpointing Figure 8-24. A recovery line.

Independent Checkpointing Figure 8-25. The domino effect.

Characterizing Message-Logging Schemes Figure 8-26. Incorrect replay of messages after recovery, leading to an orphan process.