InbytearraybyVinoth Chandar·Nov 28, 2023Doing range gets on cloud storage for fun and profitThe thing that stands between good and great cloud read performanceA response icon1A response icon1
InbytearraybyVinoth Chandar·Apr 20, 2022Corrections in data lakehouse table format comparisonsA live document to serve as a point of reference for corrections for inaccuracies for different comparative studies of Hudi, Delta Lake, or…A response icon3A response icon3
Inapache-hudi-blogsbyVinoth Chandar·Sep 2, 2021Reliable ingestion from AWS S3 using HudiIn this post we will talk about a new deltastreamer source which reliably and efficiently processes new data files as they arrive in AWS S3
Inapache-hudi-blogsbyVinoth Chandar·Jul 27, 2021Apache Hudi — The Streaming Data Lake PlatformThis blog is a repost of the original blog here
Inapache-hudi-blogsbyVinoth Chandar·Mar 15, 2021Streaming Responsibly Into the Data LakeHow Apache Hudi maintains optimum sized files
Inapache-hudi-blogsbyVinoth Chandar·Jan 28, 2021Optimize Data Lake layout using Clustering in Apache HudiThis blog is a repost of this Hudi blog on medium.
Inapache-hudi-blogsbyVinoth Chandar·Dec 19, 2020Employing the right indexes for fast updates, deletes in Apache HudiThis blog is a repost of this Hudi blog on medium.
InbytearraybyVinoth Chandar·Apr 27, 2020Apache Hudi (Incubating) Support on Apache ZeppelinReposted translation of the original article : https://mp.weixin.qq.com/s/_mNwL5uXSDYyqtLDPx0iDA
Vinoth Chandar·Jun 8, 2019Embrace the Data Lake ArchitectureOften times, data engineers build data pipelines to extract data from external sources, transform them and enable other parts of the…
InbytearraybyVinoth Chandar·Sep 16, 2018Setting up Hadoop/YARN/Spark/Hive on Mac OSX[Reposted from my blogger]A response icon1A response icon1