# Using Yugabytes with Lakefs 1.7x

**URL:** <https://forum.yugabyte.com/t/using-yugabytes-with-lakefs-1-7x/5004>\
**Category:** General\
**Created:** [April 30, 2026, 5:29am UTC](https://forum.yugabyte.com/t/using-yugabytes-with-lakefs-1-7x/5004 "2026-04-30T05:29:02Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![wazza7777](https://avatars.discourse-cdn.com/v4/letter/w/9f8e36/32.png) [@wazza7777](https://forum.yugabyte.com/u/wazza7777)\
**Post date:** [April 30, 2026, 5:29am UTC](https://forum.yugabyte.com/t/using-yugabytes-with-lakefs-1-7x/5004/1 "2026-04-30T05:29:02Z")

</div>

Hi

I have a SQL & Postgres DBA background, and our client is looking at using latest version of yugabytes for a back end for lakefs. I’d welcome expert thoughts on how feasible this is please?

I was wondering about is whether the distributed nature of yugabytes may introduce potential latencies in writes that might cause lakefs to have a bit of a tantrum. I’m not sure whether there is a way to buffer writes that means all transactions as 100% ACID in nature.

I’ve done some reading on both lakefs and yugabytes but a lot of this may come down to how a yb cluster behaves. I have a docker based set up of a single node of lakefs and single node of yb currently running, to have a play. In Prod we would look at installing yb on linux VMs ( not dockers ).

The plan seems to be having 2 yb nodes in one datacentre, 2 nodes in another with a fast link between them.

I’m also realistic in that not every technical solution may be optimal for a proposed use.

We would be counting on at least 100,000 writes per day to lakefs and the yb database.

Thoughts welcome.

---

<div class="post-metadata">

**Author:** ![dorian\_yugabyte](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.yugabyte.com/dorian_yugabyte/32/206_2.png) [@dorian\_yugabyte](https://forum.yugabyte.com/u/dorian_yugabyte)\
**Post date:** [April 30, 2026, 6:22am UTC](https://forum.yugabyte.com/t/using-yugabytes-with-lakefs-1-7x/5004/2 "2026-04-30T06:22:33Z")

</div>

Hi @wazza7777

YugabyteDB is a perfect store for metadata storage of a distributed filesystem.

> [@wazza7777](#):
>
> I was wondering about is whether the distributed nature of yugabytes may introduce potential latencies in writes that might cause lakefs to have a bit of a tantrum.

The latency may come from multi datacenter deployment, but lakefs shouldn’t have a tantrum, I’d consider that a bug.

> [@wazza7777](#):
>
> We would be counting on at least 100,000 writes per day to lakefs and the yb database.

Very low usage, should be fine in all cases.

> [@wazza7777](#):
>
> The plan seems to be having 2 yb nodes in one datacentre, 2 nodes in another with a fast link between them.

What is the exact reason behind this type of deployment? Do you need ability to lose a datacenter and continue to function? Or have 1 DC as a disaster recovery only?

Minimum is 3 nodes, see [Deployment checklist for YugabyteDB clusters | YugabyteDB Docs](https://docs.yugabyte.com/stable/deploy/checklist/#replication) .

For 2 datacenters, you can have only asynchronous replication between 2 clusters, so 6 nodes minimum: [xCluster deployments | YugabyteDB Docs](https://docs.yugabyte.com/stable/deploy/multi-dc/async-replication/)

For synchronous replication of multi datacenter, you need 3 datacenters, minimum 3 nodes: [Multi-DC deployments | YugabyteDB Docs](https://docs.yugabyte.com/stable/deploy/multi-dc/)

---

<div class="post-metadata">

**Author:** ![wazza7777](https://avatars.discourse-cdn.com/v4/letter/w/9f8e36/32.png) [@wazza7777](https://forum.yugabyte.com/u/wazza7777)\
**Post date:** [April 30, 2026, 10:30pm UTC](https://forum.yugabyte.com/t/using-yugabytes-with-lakefs-1-7x/5004/3 "2026-04-30T22:30:32Z")

</div>

Thanks for your reply.

While I read up on your responses and look through the links provided, the design aim is to have the database synchronized across at least 2 ( or more ) datacentres with synchronous replication between datacentres, such that we can write from Lakefs into the database at any datacentre, and the data made immediately available in all datacentres. The idea is to have a single resilient database for LakeFS to use under all conditions.

The system needs to survive the loss of at least one datacentre, ideally 2, so we can fall back and run on a single datacentre and maintain BAU in that scenario without manual intervention ( if possible ).

It seems initially we would need a 3 node cluster at each of our 3 datacentres to survive loss of 1 datacentre, and we would need 5 datacentres set up each with a 3 node cluster to survive loss of 2 datacentres, is that correct?

What sort of latency penalty is there the more datacentres you run in please? Is it possible to calculate it?

---

<div class="post-metadata">

**Author:** ![dorian\_yugabyte](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.yugabyte.com/dorian_yugabyte/32/206_2.png) [@dorian\_yugabyte](https://forum.yugabyte.com/u/dorian_yugabyte)\
**Post date:** [May 1, 2026, 4:35am UTC](https://forum.yugabyte.com/t/using-yugabytes-with-lakefs-1-7x/5004/4 "2026-05-01T04:35:46Z")

</div>

> [@wazza7777](#):
>
> with synchronous replication between datacentres

You need minimum 3 DCs.

> [@wazza7777](#):
>
> ideally 2

You need minimum 5 DCs.

> [@wazza7777](#):
>
> It seems initially we would need a 3 node cluster at each of our 3 datacentres to survive loss of 1 datacentre

Minimum is 3 total nodes, one at each DC, not 9.

> [@wazza7777](#):
>
> and we would need 5 datacentres set up each with a 3 node cluster to survive loss of 2 datacentres, is that correct?

5 DCs, each with 1 node. The 3-nodes-for-each-DC-minimum is when you want async-replication between 2 separate clusters.

While for synchronous replication, it’s only in-cluster.

> [@wazza7777](#):
>
> What sort of latency penalty is there the more datacentres you run in please? Is it possible to calculate it?

Writes need to get ack from majority of replicas before returning to client. 3DCs has 2 majority, 5DCs has 3 majority.

---

<div class="post-metadata">

**Author:** ![scott](https://avatars.discourse-cdn.com/v4/letter/s/3e96dc/32.png) [@scott](https://forum.yugabyte.com/u/scott)\
**Post date:** [June 3, 2026, 10:30am UTC](https://forum.yugabyte.com/t/using-yugabytes-with-lakefs-1-7x/5004/5 "2026-06-03T10:30:44Z")

</div>

The latency impact will largely depend on the network round-trip time between datacentres, since synchronous replication requires acknowledgements from a majority of replicas before a write is considered successful. As the number of datacentres increases, write latency can increase because more geographically distributed nodes participate in the consensus process. This is a common consideration in distributed database architectures and consensus-based systems such as those described in the **[Raft Consensus Algorithm](https://raft.github.io/)**. If you’re evaluating multi-datacentre database designs, it’s also worth reviewing concepts around **[Distributed Databases](https://cloudfoundation.com/blog/distributed-database-management-system/)** and **[CAP Theorem](https://www.ibm.com/think/topics/cap-theorem)**, as they provide useful context for understanding the trade-offs between consistency, availability, and performance in setups like the one you’re describing.
