DataLane

All stacks · Compute & processing

Hadoop

HDFS, YARN, Hive, and the on-prem stack still running in many banks.

Hadoop cover

Related reading

About Hadoop

Hadoop is not dead in the places that matter for interviews: banks, telcos, and on-prem estates still run HDFS, YARN, Hive, and Sqoop. The cloud lakehouse is the migration story, not the starting point for those teams.

Learn Hive tables, partitioning, and why small files kill NameNodes — then learn Iceberg on object storage as the exit. Skipping Hadoop entirely is how you fail a legacy-platform screen.

What you'll learn here

  • HDFS, YARN, and what a “cluster” meant before object storage
  • Hive Metastore as the catalog everyone still talks to
  • Small-file problems and why they follow you to S3
  • Migration patterns: Hive → Iceberg / Delta without a big-bang cutover

Frequently asked questions

Should a new platform start on Hadoop?

No. New platforms start on object storage plus Iceberg or a warehouse. Hadoop skills still matter because migrations and interviews assume you can read a Hive table and a YARN queue.

Hive or Spark SQL on a lake?

Spark (or Trino) on Iceberg/Delta is the modern path. Hive-on-Tez is maintenance. If the metastore is Hive, you can still query it from engines that are not Hive.

What is the usual Hadoop-to-cloud failure?

Lift-and-shift of every small file and every cron Sqoop job into S3 + EMR with no table format. You moved the NameNode problem into LIST requests and a bigger bill.

New Hadoop posts, straight to your inbox

One email a week with our latest tutorials. No spam.

Newsletter signup is not live yet. Use the contact form if you want to be notified.

↑↓ navigate openesc close