Posts

Showing posts with the label docker compose

JDBC Kafka Source Connector: Stream Database Changes or ENTIRE Tables 🚂📑

Image
  Confluent has built some useful tools that rely on Kafka . One of these tools is Kafka Connect . And as I have shown you before  one of the ways to use Kafka Connect as a consumer of Kafka Topics, I will show you how to use Kafka Connect as a producer to Kafka topics Kafka Connect offers a very powerful feature which is connecting to databases using certain connectors as JDBC (Java Database Connectivity).  The Kafka Connect JDBC source connector allows you to import data from any relational database with a JDBC driver into an Apache Kafka topic. This connector supports a wide variety of databases. Data is loaded by periodically executing a SQL query and creating an output record for each row in the result set. It enables you to pull data (source) from a database into Kafka , and also push data (sink) from a Kafka topic to a database. In this tutorial we'll be focusing on pulling data from a SQL database. I will be reusing the same setup as my p...

Confluent Kafka Connect: Store Kafka Messages in Elasticsearch ✉️📂

Image
 Suppose you need to store all the messages published on a certain topic. Normally you would create a consumer that listens to that topic, receives the message then stores it in the dataset that you want. What if I told you that Confluent has created a shortcut. Instead of doing this, you can just create a Kafka connector using Confluent Kafka Connect and this connector will automatically store any published message in your database. Kafka Connect is a tool for scalably and reliably streaming data between Apache Kafka and other data systems. It makes it simple to quickly define connectors that move large collections of data into and out of Kafka . Kafka Connect can ingest entire databases or collect metrics from all your application servers into Kafka topics, making the data available for stream processing with low latency. It can also deliver data from Kafka topics into secondary indexes like Elasticsearch or batch systems such as Hadoop for offline analysis. In this tutorial we...

Confluent Schema Registry: Learn Efficient Kafka Avro Serialization 📜

Image
When we talk about Kafka it's important to mention Avro serialization. Like I mentioned in my Kafka introductory post : as a message broker, Kafka is quite known for being complex as opposed to some other brokers such as SQS (which is literally called Simple Queue Service) or RabbitMQ which calls its producers Basic Produce and consumers Basic Consume . And although it requires a little more effort than the other brokers, it offers much more for services which need to trim down every inch of latency they can, due to the criticality of the business of the sheer volume of data it requires to stream. And Confluent Kafka can help with that by transporting Avro serialized messages. But in order to do that, Kafka relies on Confluent   Schema Registry . Because after it gets serialized to bytes, it needs to look for the schema that can help these bytes go back to the model that it initially was. Confluent Schema Registry is basically a store that keeps track of the schemas of the mess...

Breaking Down Kafka: A Step-by-Step Guide to Publish & Consume Messages ✉️

Image
 Message brokers come in all colors. Each broker has its own edge. There are some brokers which are aimed to be simple and direct such as Amazon SQS (which is literally called Simple Queue Service). Some of them can be used in a simple fashion, but can also be used to implement complex patterns like RabbitMQ which is a topic I talked about in my previous posts . And then there is a broker that was designed to: handle heavy-duty streaming, maximize efficiency, allow you to scale up, serialize messages to bytes, partition messages and much more. That broker is indeed Apache Kafka . Due to its wide array of features, Kafka can be overwhelming. And sometimes it feels too exhausting to start comprehending its principles and what the broker can provide. So, why not strip down the broker from all of its extra shining features and just start from its core features like publishing and consuming simple JSON serialized string messages. What we're going to explore isn't just Apache Kafk...

Introduction to Beats: Collect Data from Anywhere & Level Up your Elastic Stack 📡

Image
    Logstash did a very impressive job transforming our logs into documents that allowed us to understand and visualize how our applications are behaving. But, because of that Logstash could require quite the memory and CPU to run. So, it's not the most efficient move to make Logstash collect the data from various sources. Elastic offered a solution to this concern and introduced Beats . Beats  is a lightweight shipper for forwarding and centralizing log data. It's installed as an agent on your servers to capture all sorts of operational data like logs or network packet data. Beats is great for gathering data and works efficiently with a large number of files. It can also handle back pressure (when Logstash is busy) and ensures that no data is lost during such periods. And to be clear Logstash can do most of what Beats . So, why use Beats instead? 1. Lightweight Data Shipping : Beats is designed to be lightweight and requires fewer resources than Logstash . This makes it ...

Introduction to Kibana: Explore, Visualize and Analyze Elasticsearch Data 📊

Image
   In my previous post  Introduction to Elasticsearch: Create, Update, Delete and Search Documents 🔍  I showed you how to set up an Elasticsearch index and manage its data. But as you could see in that post the data looked very ugly on the cmd window. That's because this is not Elasticsearch 's job. Data representation and visualization is key to understanding your data and even predicting its trend. And as I previously said, the power in Elasticsearch lies in its integration with other powerful tools such as our guest of honor: Kibana . Kibana , developed by Elastic , is a powerful open-source data visualization and exploration platform. It seamlessly integrates with Elasticsearch, making it an essential component of the Elastic Stack. Whether you’re a data analyst, developer, or business user, Kibana empowers you to unlock valuable insights from your data. With its intuitive interface, you can create interactive dashboards, explore logs, analyze metrics, and visu...