Gronex
Log in

Distributed Systems

23 guides

Build an intuition for how distributed systems behave when networks, machines and dependencies fail.

IdempotencyAn operation is idempotent when performing it more than once has the same effect as performing it once — the property that makes retrying safe.Distributed systemsCAP theoremCAP says that when a network partition occurs, a replicated data store must give up either linearizable reads and writes or the ability to answer every request — it does not say you choose two properties out of three.Distributed systemsDelivery semanticsExactly-once delivery is unachievable over a network that can lose messages; exactly-once *effects* are achievable, by combining at-least-once delivery with deduplication or idempotent processing.Distributed systemsMessage orderingBrokers give total order only inside a partition, so ordering is something you buy with a partition key — and pay for in parallelism.Distributed systemsDistributed lockingA distributed lock is a lease, not a mutex: it can expire while its holder still believes it holds the lock, which is why correctness requires fencing at the resource rather than trust at the client.Distributed systemsLeader electionLeader election gives one node the right to act for a bounded term — and because a leader cannot tell a partition from a slow network, safe designs assume there may briefly be two.Distributed systemsConsistent hashingConsistent hashing maps keys and nodes onto the same ring so that adding or removing a node relocates roughly 1/N of the keys instead of nearly all of them.Distributed systemsQuorum reads and writesA quorum system replicates to N nodes and requires W acknowledgements to write and R responses to read; when R + W > N the read set and the write set must overlap, so a read sees at least one copy of the latest completed write.Distributed systemsVector clocks and causalityVector clocks track causality rather than time: they can tell you that one version descends from another, or that two versions are concurrent and must be merged — something wall-clock timestamps cannot do.Distributed systemsSplit brainSplit brain is the state where two nodes both believe they are the authoritative writer, which happens because a node cannot distinguish “the others are gone” from “I am cut off”.Distributed systemsService discoveryService discovery answers “where is an instance of X right now”, and its hard part is not registration but timely, safe removal of instances that have stopped being healthy.Distributed systemsDead letter queueA dead-letter queue parks messages a consumer cannot process so the rest of the stream keeps flowing — which only helps if something distinguishes permanent failures from transient ones and someone actually drains it.Distributed systemsChange data captureCDC turns a database’s own change log into a stream of events, so other systems can follow every row change without the application publishing them.Distributed systemsEvent sourcingEvent sourcing stores the sequence of state changes as the source of truth, and derives current state by replaying them.Distributed systemsCQRSCQRS uses different models for writing and reading, so each can be shaped for its own job instead of compromising on one.Distributed systemsConsensus and RaftConsensus gets a group of nodes to agree on an ordered sequence of operations despite failures, which is what makes replicated state safe.Distributed systemsGossip protocolsEach node periodically exchanges state with a few random peers, so information spreads exponentially without any node knowing the whole cluster.Distributed systemsMonolith vs MicroservicesA modular monolith is the default for most teams; microservices earn their cost when independent scaling, ownership, or fault isolation is a measured need.Distributed systemsSynchronous communication vs Asynchronous communicationSynchronous calls are simplest for immediate answers; asynchronous messaging wins for decoupling and work that can finish later. Compare failure, ordering, and user experience.Distributed systemsOrchestration vs ChoreographyOrchestration centralises workflow decisions; choreography lets services react to events. Choose based on visibility, coupling, ownership, and failure recovery.Distributed systemsMessage queue vs Event streamQueues distribute work to consumers; event streams retain an ordered log for independent readers and replay. Compare delivery, retention, ordering, and recovery.Distributed systemsKafka vs RabbitMQKafka is a retained partitioned log for high-volume streams; RabbitMQ is a flexible broker for routed work queues. Compare delivery, ordering, replay, and operations.Distributed systemsStateless services vs Stateful servicesStateless services scale and recover easily; stateful services are necessary when ownership, ordering, or local durable state matters. Compare routing and failure recovery.Distributed systems