Posts

Showing posts with the label distributed systems

Why I Hate Microservices Part 3: The Identity Crisis ๐Ÿ˜ต

Image
       Imagine going out to buy bread. And you know that there are two types of bakeries. One that gives you the bread right away because they know exactly when it's going to be done. The other kind takes the order from you and tells you to go home and it will deliver the bread to you once it's ready because it's unclear for them when it's going to be done. Then it's clear for you what to expect when dealing with both sorts of bakeries. Now, imagine if there were a third type of bakeries. One that tells you that it's going to give it to you right away while it's actually unclear for them when the bread is going to done. So, you expect the bakery to behave in a certain way but it will actually behave in another. Well, imagine working in such a bakery. Eventual ๐Ÿ“จ VS Transactional⚡ In my humble opinion, the problem that I'm discussing in this article is probably the most critical one when designing a system. It's a problem about the identity of your so...

Why I Hate Microservices Part 2: The Who's Telling the Truth Problem ๐Ÿคท

Image
      When designing a system, you should not limit yourself with using only one type of databases. Some business needs might require having a more flexible sort of databases like NoSQL / Document databases.  One of these databases is Elasticsearch ,   which is a topic that I have talked about in a previous post that you should definitely check out. However, if you use more than one database for your system, making sure that all your databases are in sync is crucial . How It Started ๐Ÿšฉ As I said in my previous post . The other sort of problems that we have faced while working on the microservices project is data inconsistency. To explain this problem I have to tell you why the system needed an Elasticsearch index. In microservices projects you usually have more than one database, one for each service/domain. The client-side needed an aggregated document of the data from all the databases. Which in microservices is called The Aggregate Root ,   and that is ...

Why I Hate Microservices Part 1: The Russian Dolls Problem ๐Ÿช†๐Ÿช†๐Ÿช†

Image
     People always want the newest things. The more trendy something is, the more it appeals to most people. Same applies to technology. Adding the word microservices to the CV can help it shine a little bit. These are some of the thoughts I was having while I was 2 hours deep into debugging an issue in a project that involved .Net, Kafka, SignalR, React js, Elasticsearch, ECS, MediatR and more. Beware of Microservices  ⚠️ What I'm about to share is a cautionary tale for people who are on the verge of shifting their solution into a microservices architecture. Before you start implementing such an architecture, please read carefully the decisions that others have made along with their consequences so that you can see for yourself the outcome of each design decision instead of falling for the same mistakes. Before anything : The case study that I’m about to discuss might have had its problems. However, the people who made the design decisions are very good software e...

Partitioning in Kafka: A Guide to Publishing Batched Data ๐Ÿ“ƒ๐Ÿ“ฉ

Image
 There are two keywords that you must understand when publishing a message: latency and throughput.  Throughput is a measure of how much data can be processed in a given amount of time. It's usually measured in bits per second (bit/s), or data packets per second. High throughput means the system can process a large amount of data quickly, which is often desirable in high-load scenarios. For example, if a Kafka producer can send 1000 messages per second to a broker, the throughput is 1000 messages per second. Latency, on the other hand, is a measure of time delay experienced in a system, the time it takes for a bit of data to travel from one point to another in a network. It is usually measured in milliseconds. Low latency means that data can be transferred quickly from source to destination. For example, if a message takes 10 milliseconds from the time it's sent by a Kafka producer until it's received by a broker, the latency is 10 milliseconds. In all systems, there's ...

Introduction to Consumer Groups: Learn How to Horizontally Scale Kafka Consumers ๐Ÿ“ฅ

Image
 People say that Kafka is a dumb broker. It just holds data under some defined topics and forward messages from producers to consumers. It doesn't do much processing on the messages, doesn't route them based on content, and doesn't transform them. But that actually isn't entirely true. Kafka does much more than just storing messages. It keeps count of which consumer group consumed which message by tracking its offset. The consumer group is the group Id you give your consumer once you initialize it. Having more than one consumer helps you avoid consumer failures and scale better. If a message is consumed by one consumer in a group, no other consumers in the same group will receive it. However, all other consumers with other group Ids will receive the message. It's important to note that the offset of each message is basically the id of the message related to the group. So, if "message A" was published on a brand-new group its offset will be 0. But if anoth...

Confluent Schema Registry: Learn Efficient Kafka Avro Serialization ๐Ÿ“œ

Image
When we talk about Kafka it's important to mention Avro serialization. Like I mentioned in my Kafka introductory post : as a message broker, Kafka is quite known for being complex as opposed to some other brokers such as SQS (which is literally called Simple Queue Service) or RabbitMQ which calls its producers Basic Produce and consumers Basic Consume . And although it requires a little more effort than the other brokers, it offers much more for services which need to trim down every inch of latency they can, due to the criticality of the business of the sheer volume of data it requires to stream. And Confluent Kafka can help with that by transporting Avro serialized messages. But in order to do that, Kafka relies on Confluent   Schema Registry . Because after it gets serialized to bytes, it needs to look for the schema that can help these bytes go back to the model that it initially was. Confluent Schema Registry is basically a store that keeps track of the schemas of the mess...

Help Your Messages Find Peace: A Guide to Dead Letter Queues in RabbitMQ ๐Ÿ’€✉️

Image
    Exceptions can happen anytime and if you don't know how to handle them, then you have got a problem. Or maybe you just haven't checked out my post about exception handling. NET Global Exception Handling: 3 Techniques Beyond Try/Catch Blocks . Anyway, because of that you need to plan what to do if an exception happens during processing a message that you received from a  RabbitMQ queue. You have two options: Negatively Acknowledge (Nack) Using BasicNack:  This method is more flexible. It allows you to negatively acknowledge one or more messages. It takes three parameters: the delivery tag of the message to nack, a boolean indicating whether to nack multiple messages, and a boolean indicating whether to requeue the message.  If the multiple parameter is true, all messages up to and including the one with the specified delivery tag are nacked. If the requeue parameter is true, the nacked messages will be requeued. If it's false, the messages will be dis...

RabbitMQ’s Lost & Found: 2 Different Techniques to Rescue Unrouted Messages ๐Ÿ“ฌ

Image
  RabbitMQ is often compared to as a post office . And this post office has multiple postmen (Exchanges) and each exchange is responsible for a number of mailboxes (Queues) and each postman has his own style of delivery (Exchange Types). Well, the postmen of RabbitMQ are by default careless.  So, if they get an address that they can't find, they just throw away the message . So, you have to set your ground rules and show them what to do if they get a non-existing address. In my last RabbitMQ post  Publish and Consume Messages Using RabbitMQ as a Message Broker in 4 Simple Steps ๐Ÿฐ  I showcased this problem when I deliberately sent a message to a non-existing queue name and as you can see in the post there was no exception thrown or anything. So basically, the message was a victim of fire and forget . So, let's use the same setup as the previous post and let's try two techniques that could help us avoid this problem. Our aim is to at least log any message that did not...