# Fault tolerance

**URL:** <https://forum.yugabyte.com/t/fault-tolerance/784>\
**Category:** General\
**Created:** [August 3, 2020, 5:36pm UTC](https://forum.yugabyte.com/t/fault-tolerance/784 "2020-08-03T17:36:50Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ben\_Jiro](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.yugabyte.com/ben_jiro/32/242_2.png) [@Ben\_Jiro](https://forum.yugabyte.com/u/Ben_Jiro)\
**Post date:** [August 3, 2020, 5:36pm UTC](https://forum.yugabyte.com/t/fault-tolerance/784/1 "2020-08-03T17:36:50Z")

</div>

Is my logic correct in assuming ( with all nodes having plenty of free space ):

YugabyteDB with 3 nodes:

- Single node failure, the replication will distribute among the two remaining nodes. Read / Write remains available If another node dies, it enters read-only mode.
- Dual node ( both instantly unable ) failure, the remaining node goes into read-only mode.

YugabyteDB with 4 nodes:

- Single node failure, the replication will distribute among the three remaining nodes. If the replication has been successful, YugabyteDB enters a state of “3 nodes” ( See 3 node example how it will be handled upon future failures ).
- Dual node failure ( both failing withing 60 seconds )? Does YugabyteDB goes into read-only mode?

YugabyteDB with 5 nodes:

- Single node failure, the replication will distribute among the four remaining nodes. If the replication has been successful, YugabyteDB enters into 4 node behavior ( See 4 node example how it will be handled upon future failures ).
- Dual node failure ( both failing withing 60 seconds )? Replication will start between the 3 nodes. if successful, YugabyteDB enters into 3 node behavior ( See 3 node example how it will be handled upon future failures ). Able to handle another failure.
- Triple node failure ( all 3 failing withing 60 seconds )? Does YugabyteDB goes into read-only mode?

Is this understanding correct? If yes, please also also update the documentation because i see in a lot of DB’s like YugabyteDB, CRDB that talk about node failures, do not really mention “instantaneous” failures vs “slow” failure ( with a chance to rebuild and enter a different node setup ). And the effects on different type of node 3,4,5,6,7 with instant vs slow failures on the DB infrastructure are mostly glanced over.

---

<div class="post-metadata">

**Author:** ![dorian\_yugabyte](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.yugabyte.com/dorian_yugabyte/32/206_2.png) [@dorian\_yugabyte](https://forum.yugabyte.com/u/dorian_yugabyte)\
**Post date:** [August 3, 2020, 5:53pm UTC](https://forum.yugabyte.com/t/fault-tolerance/784/2 "2020-08-03T17:53:16Z")

</div>

Hi @Ben_Jiro

Welcome to YugabyteDB Forum!

Are we assuming replication factor = 3 in all those cases ?

Regards,  
Dorian  
Technical Support Engineer

---

<div class="post-metadata">

**Author:** ![Ben\_Jiro](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.yugabyte.com/ben_jiro/32/242_2.png) [@Ben\_Jiro](https://forum.yugabyte.com/u/Ben_Jiro)\
**Post date:** [August 3, 2020, 6:25pm UTC](https://forum.yugabyte.com/t/fault-tolerance/784/3 "2020-08-03T18:25:56Z")

</div>

Indeed. A replication of 3.

---

<div class="post-metadata">

**Author:** ![dorian\_yugabyte](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.yugabyte.com/dorian_yugabyte/32/206_2.png) [@dorian\_yugabyte](https://forum.yugabyte.com/u/dorian_yugabyte)\
**Post date:** [August 5, 2020, 6:52am UTC](https://forum.yugabyte.com/t/fault-tolerance/784/4 "2020-08-05T06:52:36Z")

</div>

Assuming RF=3 in all scenarios.  
Assuming free disk space in all scenarios.

1. **If a node fails** , the tablets which leaders resided on the failed node won’t accept writes until new leaders have been chosen on the other nodes after [`--leader_failure_max_missed_heartbeat_periods`](https://docs.yugabyte.com/latest/reference/configuration/yb-tserver/#leader-failure-max-missed-heartbeat-periods) (default 3 seconds)

2. **If a node is down for** [`--follower_unavailable_considered_failed_sec`](https://docs.yugabyte.com/latest/reference/configuration/yb-tserver/#follower-unavailable-considered-failed-sec) (default 15 minutes), the node is considered dead and data will start replicating to the other machines.

3. **If we lose 2 peers of a tablet** , then we have read only of those tablets ONLY when the peer that is available already was the leader (it can’t pick new leaders because there’s no quorum) OR when using follower reads. If those come back online, they will start replicating. If we lost them forever(15+minutes), then we must do [manual recovery](https://github.com/yugabyte/yugabyte-db/pull/4852) of those tablets.

4. **If we lose 3 peers of a tablet** , that tablet is unavailable for read/writes. No replicas exist. You must resurrect at least one of those nodes. If 1 tablet resurrection, follow **if we lose 2 peers of a tablet**.

5. Generally if [RF is `n` , YugabyteDB can survive `(n - 1) / 2` failures](https://docs.yugabyte.com/latest/deploy/checklist/#replication) without compromising correctness or availability of data.

6. Replication happens per-tablet. Example: assuming a transaction that needs data in 2 tablets… If both have leaders, the transaction can continue.

> [@](#):
>
> 3 nodes, 1 goes down.

See **1**.

There is no reason to have 2 replicas of the same tablet in a server (3 replicas distributed on 2 servers) so no replication needed.

> [@](#):
>
> 3 nodes, 2 go down

See **3** above.

> [@](#):
>
> 4 nodes, 1 goes down

See **1** , **2**.

> [@](#):
>
> 4 nodes, 2 go down

See **1** , **2** , **3**.

> [@](#):
>
> 5 nodes, 1 goes down

See **1** , **2**.

> [@](#):
>
> 5 nodes, 2 go down

See **1** , **2** , **3**.

> [@](#):
>
> 5 nodes, 3 go down

See **1** , **2** , **3** , **4**.

---

<div class="post-metadata">

**Author:** ![pcyb](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.yugabyte.com/pcyb/32/310_2.png) [@pcyb](https://forum.yugabyte.com/u/pcyb)\
**Post date:** [February 14, 2021, 4:13pm UTC](https://forum.yugabyte.com/t/fault-tolerance/784/5 "2021-02-14T16:13:45Z")

</div>

> [@dorian\_yugabyte](#):
>
> > [@](#):
> >
> > 4 nodes, 1 goes down
> 
> See **1** , **2**.

Reads too wont be served until the tablet leader is chosen for those tablets that had the leaders on the 1 node that went down, isnt it?

---

<div class="post-metadata">

**Author:** ![dorian\_yugabyte](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.yugabyte.com/dorian_yugabyte/32/206_2.png) [@dorian\_yugabyte](https://forum.yugabyte.com/u/dorian_yugabyte)\
**Post date:** [February 15, 2021, 2:13pm UTC](https://forum.yugabyte.com/t/fault-tolerance/784/6 "2021-02-15T14:13:06Z")

</div>

> [@pcyb](#):
>
> Reads too wont be served until the tablet leader is chosen for those tablets that had the leaders on the 1 node that went down, isnt it?

Yes you won’t be able to read from leaders since they are down.  
In YCQL (soon in YSQL) you can also read from tablet peers.
