DataLane

All stacks · Platform & IaC

Kubernetes

Airflow, Spark, Flink, and Kafka operators in production clusters.

Kubernetes cover

About Kubernetes

Kubernetes is where many platforms run Airflow, Flink, Kafka, and Spark operators. You do not need to be a cluster admin, but you do need requests/limits, liveness, and why a DAG that forks 200 pods without quotas takes the node pool down.

What you'll learn here

  • Operators vs raw Deployments for Airflow, Spark, Flink, Kafka
  • Requests, limits, and HPA — the difference between busy and evicted
  • Jobs vs CronJobs vs an orchestrator you already paid for
  • When MWAA / Composer / a vendor is the right way to not run K8s

Frequently asked questions

Do I have to learn Kubernetes to be a data engineer?

You have to read a pod spec and a crash loop. You do not have to pass CKA. If your company runs Spark-on-K8s, learn the operator; if they run Databricks, learn Databricks.

Spark on Kubernetes or Databricks?

Databricks if you want a product and a support contract. Spark-on-K8s if you already staff a platform team and the bill of Databricks is the problem. The APIs look similar; the on-call does not.

Why did Airflow on K8s page us?

Workers without limits, a metadata DB on a tiny disk, and executors that spawn unbounded pods. Treat the scheduler like a production app, not a notebook sidecar.

New Kubernetes posts, straight to your inbox

One email a week with our latest tutorials. No spam.

Newsletter signup is not live yet. Use the contact form if you want to be notified.

↑↓ navigate openesc close