DebeziumSoftware
recommended
For the first stage, we chose Debezium as a Change Data Capture (CDC) provider. Debezium is an open source distributed change data capture platform built on top of Kafka Connect. Debezium comes with a well-documented first class Postgres CDC connector that’s battle-tested.
Robinhood choosing Debezium for change data capture into its data lake
robinhood.com ↗·2022-02-18Wrong?
Apache HudiSoftware
recommended
For the second stage, we are using Apache Hudi to incrementally ingest changelogs from Kafka to create data-lake tables.
Robinhood choosing Hudi for incremental data lake ingestion
robinhood.com ↗·2022-02-18Wrong?
KubernetesSoftware
recommended
In this post, we’ll talk a little about why we’re embracing Kubernetes to tackle these challenges, share some stories from our experience onboarding applications onto Kubernetes, and discuss the platform we built to manage and standardize our Kubernetes-powered applications, called the Archetype Framework.
Robinhood moves its microservice deployments onto Kubernetes
robinhood.com ↗·2019-11-13Wrong?
Apache SparkSoftware
uses
We run production batch processing pipelines predominantly using Apache Spark. Our dashboards are powered by Trino distributed SQL query engine.
Robinhood's data lake compute stack
robinhood.com ↗·2022-02-18Wrong?
TrinoSoftware
uses
We run production batch processing pipelines predominantly using Apache Spark. Our dashboards are powered by Trino distributed SQL query engine.
Robinhood's data lake query engine
robinhood.com ↗·2022-02-18Wrong?
KafkaSoftware
uses
We use Kafka as our databus to stream data, which usually comes from external feeds, or from internal Faust apps.
Robinhood's data lake ingestion layer
robinhood.com ↗·2019-09-09Wrong?
PostgreSQLSoftware
uses
We use PostgreSQL and AWS Aurora as our relational databases, Elasticsearch as a document storage and indexing solution, AWS S3 for object storage, and InfluxDB as a time series database.
Robinhood's data stores ahead of building its data lake
robinhood.com ↗·2019-09-09Wrong?
SaltStackSoftware
mixed
Historically, we have used a combination of Terraform and SaltStack to manage our AWS infrastructure. While this combination of technologies has carried us quite far
Robinhood on the Terraform and SaltStack setup it was moving away from
robinhood.com ↗·2019-11-13Wrong?
RedshiftSoftware
disliked
We initially used Elasticsearch and Redshift as our analytics and data warehousing solutions. After a couple of years of maintaining these systems, we determined that scaling these systems to support our increasing data workload was neither efficient nor economical.
Why Robinhood built a data lake instead of scaling Elasticsearch and Redshift
robinhood.com ↗·2019-09-09Wrong?
ElasticsearchSoftware
disliked
We initially used Elasticsearch and Redshift as our analytics and data warehousing solutions. After a couple of years of maintaining these systems, we determined that scaling these systems to support our increasing data workload was neither efficient nor economical.
Why Robinhood built a data lake instead of scaling Elasticsearch and Redshift
robinhood.com ↗·2019-09-09Wrong?