Sitemap
A human index of the site. Machines can usearticles.json,RSS, orllms.txt.
Pages
- Home
- Start here
- All posts
- All stacks
- SQL Playground
- Cheat sheets
- Compare tools
- Glossary
- Interview prep
- Interview by company
- Certification practice tests
- Resources
- Saved posts
- Advertise
- Write for us
- About
- Contact
- Privacy
- Terms
- Disclaimer
- Affiliate disclosure
Feeds & catalogs
- RSS feed
- Article catalog (JSON)
- llms.txt
- llms-full.txt
- LLM sitemap
- OpenAPI
- Agent access rules
- Search index
- security.txt
- API catalog
Stacks
- Airflow
- Dagster
- Prefect
- Apache NiFi
- Apache Spark
- Apache Flink
- Apache Beam
- Hadoop
- Kafka
- Apache Pulsar
- Debezium
- dbt
- SQLMesh
- Snowflake
- Databricks
- BigQuery
- Amazon Redshift
- ClickHouse
- DuckDB
- Microsoft Fabric
- Trino
- Iceberg
- Apache Hudi
- Airbyte
- Fivetran
- Data Modeling
- Data Quality
- AWS
- Azure
- GCP
- Terraform
- Kubernetes
- Docker
- Python
- SQL
- Polars
- Scala
- PostgreSQL
- MongoDB
- Cassandra
- Redis
- Elasticsearch
- BI & semantic layer
- AI & GenAI
- MLOps
- Productivity
- Career
- Tools
Cheat sheets (69)
- Airflow Best Practices
- Airflow CLI & Concepts
- Airflow Interview Questions
- AWS Data Engineer Interview Questions
- AWS for data engineers
- Azure & Fabric
- Azure Data Engineer Interview Questions
- Behavioral Interview for Data Engineers
- BigQuery Interview Questions
- BigQuery SQL
- CDC Patterns
- Dagster
- Data modeling
- Data Modeling Interview Questions
- Data Quality Interview Questions
- Data System Design Interview
- Data Warehouse Interview Questions
- Databricks CLI & Unity Catalog
- Databricks Interview Questions
- dbt Best Practices
- dbt CLI
- dbt Interview Questions
- dbt Tests & Contracts
- DevOps for Data Interview Questions
- Docker for Data Engineers
- DuckDB SQL
- ETL & ELT Interview Questions
- GCP Data Engineer Interview Questions
- GCP for Data Engineers
- Git for Data Engineers
- Great Expectations & Soda
- IAM for Data Teams
- Iceberg & Lakehouse Interview Questions
- Indexes & Constraints
- JSON in SQL (Snowflake, BigQuery, Postgres, DuckDB)
- Kafka CLI
- Kafka Interview Questions
- Medallion Architecture
- MERGE & Upsert Patterns
- Microsoft Fabric
- Open Table Formats (Iceberg, Delta, Hudi)
- pandas
- pandas Interview Questions
- polars
- Postgres for Data Engineers
- Prefect
- PySpark
- PySpark Coding Questions
- PySpark Tuning Flags
- pytest for Data Pipelines
- Python Interview Questions
- Python Regex
- Python Stdlib for Data Engineers
- Regex in SQL
- S3 Patterns for Data Lakes
- Senior Data Engineer Interview Questions
- Snowflake Admin Commands
- Snowflake Best Practices
- Snowflake Cortex and Gen AI
- Snowflake SQL
- Spark Interview Questions
- SQL Date & Time Functions Across Warehouses
- SQL Query Optimization Questions
- SQL Window Functions
- Streaming Architectures
- Streaming Interview Questions
- Terraform for Data Infrastructure
- Top Snowflake Interview Questions
- Top SQL Interview Questions for Data Engineers
Articles (127)
- A dbt Testing Strategy That Catches Regressions, Not 4,000 Warnings · dbt
- ADF Pipeline Patterns That Scale: Parameterized Datasets, Metadata-Driven Ingestion, and Integration Runtime Sizing · Azure
- Airbyte in Production: A Successful Sync of Zero Rows Is Still an Outage · Airbyte
- Airflow 3 for Pipeline Authors: What Actually Changed and What to Fix First · Airflow
- Airflow Datasets and Assets: Replacing Cron Guesswork With Data-Aware Scheduling · Airflow
- Airflow Deferrable Operators: Stop Burning Worker Slots on Sensors · Airflow
- Airflow Dynamic Task Mapping: Fan-Out That Works, and Fan-Out That Melts the Scheduler · Airflow
- Airflow vs Dagster vs Prefect in 2026: Which Orchestrator Should You Pick? · Tools
- AWS Glue vs EMR for Spark: The Cost Model, the Cold Starts, and the Migration Point · AWS
- Azure Data Engineering in 2026: Data Factory, Synapse, and Where Fabric Fits · Azure
- Batch LLM Enrichment: Idempotency, Cost Per Row, and Knowing When the Output Is Wrong · AI & GenAI
- Beam Runners and Dataflow Cost: Portability vs Workers That Never Scale to Zero · Apache Beam
- BigQuery Partitioning and Clustering: Stop Paying for Full Table Scans · BigQuery
- Building Your First Data Pipeline with Python and Airflow · Airflow
- Cassandra: A New Query Is a New Table, Not a New Index Hope · Cassandra
- ClickHouse MergeTree: ORDER BY Is the Index, Tiny Inserts Are the Outage · ClickHouse
- Cortex Analyst and Cortex Search: Shipping Chat With Your Data That Answers Correctly · AI & GenAI
- CTE vs Subquery: When the WITH Clause Is a Materialization Barrier · SQL
- Cutting a BigQuery Bill: Partitioning, Byte Limits, Materialized Views, and On-Demand vs Editions · BigQuery
- Cutting a Databricks Bill: Job Clusters, Spot, Photon, and Finding the Top Offenders in System Tables · Databricks
- Cutting Memory in Python Data Jobs: Dtypes, Chunking, Arrow, and Streaming · Python
- Dagster Assets: Partitions, Checks, and Secrets That Do Not Live in Ops · Dagster
- Data Contracts for Pipeline Teams: Fail the Producer, Not Just the Test · Python
- Data Quality Checks in Python: Catch Bad Data Before Your Users Do · Python
- Data Quality: Fail the Job, Open a Ticket, or Ignore Red Forever · Data Quality
- Databricks and Delta Lake: A Practical Introduction to the Lakehouse · Databricks
- Dataflow vs Dataproc: The Beam Model, the Cost Profiles, and Which One Your Team Can Actually Run · GCP
- dbt Incremental Models in Production: unique_key, Merge, and Late Arrivals · dbt
- dbt Macros and Jinja: The Patterns That Earn Their Keep, and the Clever Ones That Do Not · dbt
- dbt Performance at Scale: Threads, Materializations, and Per-Model Cost Attribution · dbt
- dbt Project Structure: A Layout That Survives Two Years and Forty Models · dbt
- dbt Snapshots for SCD Type 2: Check vs Timestamp, and When to Hand-Roll · dbt
- dbt Tutorial: Build Your First Transformation Project the Right Way · dbt
- Debezium on Postgres: Replication Slots, Snapshots, and Applying Changes Idempotently · Debezium
- Declarative Pipelines in Databricks: Expectations, Streaming Tables, and When the Abstraction Fights You · Databricks
- Delta Lake Internals: The Transaction Log, Checkpoints, and Why Small Files Happen · Databricks
- Delta Lake vs Apache Iceberg: Default on Databricks vs Shared Lake · Databricks
- Docker Images: Pin Digests or Watch :latest Ship a Different Binary · Docker
- DuckDB in Production: The Jobs It Genuinely Wins and the Single-Node Limits You Will Hit · DuckDB
- DuckDB vs Spark on One Node: Where the Crossover Actually Happens · DuckDB
- DuckDB: The Fastest Way to Build Local Data Pipelines in 2026 · DuckDB
- Elasticsearch: Sync Documents, Stop Treating Aggregations as Facts · Elasticsearch
- Fabric OneLake and Capacity: Shortcuts First, Then Stop Starving Spark · Microsoft Fabric
- Fivetran MAR: The Invoice That Doubled, and the Credits You Still Pay · Fivetran
- Flink Event Time and Watermarks: Checkpoints, Savepoints, and the Sink Key You Still Need · Apache Flink
- Hadoop Hive to Iceberg: Why Lift-and-Shift to S3 Fails and How Banks Actually Exit · Hadoop
- How BigQuery Pruning Actually Works: Partition Metadata, Block Statistics, and the Shapes That Defeat Both · BigQuery
- How Teams Cut Warehouse Costs by 60% with Query Optimization (Sponsored Example) · Tools
- Hudi COW vs MOR: Incremental Queries Need Compaction, Not Another Format · Apache Hudi
- Iceberg vs Delta Lake in 2026: Metadata Design, Catalog Options, and How to Actually Choose · Iceberg
- Joins and the Fan-Out Bug: How Revenue Doubles Without an Error · SQL
- Kafka Connect in Production: Converters, Dead Letter Queues, and Worker Sizing · Kafka
- Kafka Consumer Group Rebalancing: Why Your Consumers Stall and How to Stop It · Kafka
- Kafka Exactly-Once Semantics: What Transactions Actually Guarantee · Kafka
- Kafka Fundamentals: Topics, Partitions, and Consumer Groups Explained · Kafka
- Kafka Schema Registry: Evolving Event Schemas Without Breaking Consumers · Kafka
- Kafka vs Amazon Kinesis: Control vs Less Ops · Kafka
- Kinesis Data Streams vs Firehose: Shards, On-Demand, and Landing Streams in S3 · AWS
- Kubernetes Operators for Data Jobs: Unbounded Pods, Then the Cluster Dies · Kubernetes
- Making Athena Fast and Cheap: Partition Projection, Parquet Layout, CTAS, and Workgroup Byte Limits · AWS
- MCP for Data Engineers: A Tool Bus, Not a Warehouse Login · AI & GenAI
- Migrating from Synapse to Microsoft Fabric: What Maps Cleanly, What Does Not, and How to Sequence It · Azure
- MLOps for Data Engineers: Feature Stores, Training Pipelines, and Where You Fit · MLOps
- MongoDB to Silver: Flatten Nested Arrays Before They Double Revenue · MongoDB
- NiFi Backpressure and Provenance: Unbounded Queues, Disk Fill, and When Kafka Connect Wins · Apache NiFi
- Pandas to Polars: What Translates, What Does Not, and What Breaks Quietly · Python
- Pivot and Unpivot in SQL: Patterns That Survive a Changing Column Set · SQL
- Polars LazyFrames: The Speedup Is the Plan, Not the Syntax · Polars
- Postgres: EXPLAIN ANALYZE, VACUUM, and Why BI Must Leave the Primary · PostgreSQL
- Prefect Deployments: Flows, Work Pools, and the Green Run With Zero Rows · Prefect
- Pulsar Multi-Tenant Messaging: Brokers, BookKeeper, and When Not to Leave Kafka · Apache Pulsar
- PySpark for Data Engineers: From Zero to Your First Production Job · Apache Spark
- RAG Over Warehouse Data: Text-to-SQL, Curated Marts, and the Guardrails Between Them · AI & GenAI
- RAG Pipelines for Data Engineers: You Already Know How to Build This · AI & GenAI
- Reading the Snowflake Query Profile: Where the Credits Actually Go · Snowflake
- Recursive CTEs in Production: Org Charts, Bill of Materials, and Cycle Protection · SQL
- Redis Features: Cache the Online Path, Never the Source of Truth · Redis
- Redshift Serverless vs Snowflake in 2026: Architecture, Concurrency, Cost Model, and Migration Friction · AWS
- Redshift Sort Keys, WLM, and Spectrum: RA3 Tuning and the Concurrency Scaling Bill · Amazon Redshift
- Scala Datasets: Read the Hot Path Even If You Ship PySpark · Scala
- Semantic Layer: One Revenue Definition, Not Three Tableau Workbooks · BI & semantic layer
- Six Ways to Deduplicate a Table: What Each Costs and Which Row Survives · SQL
- Slim CI in dbt: state:modified+ and --defer on a 1,000-Model Project · dbt
- Slowly Changing Dimensions: Types 1 Through 6, and the Two You Will Actually Use · Data Modeling
- Snowflake Clustering Keys: When They Pay for Themselves and When They Burn Credits · Snowflake
- Snowflake Cortex AI for Data Engineers: SQL Functions, Not Another Chatbot · Snowflake
- Snowflake Cost Optimization: 7 Techniques That Actually Cut the Bill · Snowflake
- Snowflake Data Sharing: Secure Shares, Reader Accounts, and the Marketplace Without Shipping Copies · Snowflake
- Snowflake Dynamic Tables: TARGET_LAG, Refresh Warehouses, and When to Keep dbt · Snowflake
- Snowflake Masking and Row Access Policies: Tag-Based Governance That Scales · Snowflake
- Snowflake Micro-Partitions and Pruning: Why Your Filter Still Scans the Table · Snowflake
- Snowflake RBAC Design: A Role Hierarchy That Survives an Audit and a Reorg · Snowflake
- Snowflake Search Optimization Service: Point Lookups Without the Full Scan · Snowflake
- Snowflake Streams and Tasks: Incremental Apply, Schedules, and When Airflow Still Owns the Graph · Snowflake
- Snowflake Time Travel and Cloning: What They Really Cost and How to Migrate Safely · Snowflake
- Snowflake Time Travel: How It Works, What It Costs, and When It Is Not a Backup · Snowflake
- Snowflake vs BigQuery: Credits vs Bytes Scanned · Tools
- Snowflake vs Databricks in 2026: Pick the Workload, Not the Logo · Tools
- Snowflake Warehouse Sizing: The Spill and Queue Signals That Tell You Up or Out · Snowflake
- Snowflake Zero-Copy Clones: Storage After Writes, Time Travel, and CI Schemas · Snowflake
- Snowflake-Managed Iceberg Tables: Open Storage without Losing Warehouse SQL · Snowflake
- Snowpark vs SQL: When Python Wins and When It Is Just a Slower Way to Write SQL · Snowflake
- Snowpipe vs Snowpipe Streaming: Latency, Cost Per Row, and Which One You Need · Snowflake
- Spark Data Skew: Diagnosing It Properly, Then Fixing It With Salting, AQE, and Broadcast · Apache Spark
- Spark Executor Memory: Spill, GC, and the OOM Errors That Are Really Partition-Sizing Errors · Apache Spark
- Spark Shuffle Explained: What It Costs, How to Read It in the UI, and the Rewrites That Remove It · Apache Spark
- SQL Anti-Patterns That Quietly Multiply Your Warehouse Bill · SQL
- SQL Window Functions: The 5 Patterns Every Data Engineer Uses Weekly · SQL
- SQLMesh Plans and Virtual Environments: What dbt State Cannot Isolate, and When to Stay · SQLMesh
- Star Schema vs One Big Table: What Actually Costs Money on a Columnar Warehouse · Data Modeling
- Step Functions vs Airflow for Data Orchestration: Cost, Observability, and the Point Where You Outgrow One · AWS
- Structured Streaming in Production: Triggers, Checkpoints, Watermarks, and Exactly-Once Sinks · Apache Spark
- Terraform for Data Platforms: Modules for Warehouses, Not Click-Ops Grants · Terraform
- Testing Airflow Locally: A Setup That Catches Bugs Before Production Does · Airflow
- Testing Data Pipelines in Python: Fixtures, Warehouse Fakes, and Property-Based Transforms · Python
- Text-to-SQL in Production: Why the Demo Works and the Deployment Fails · AI & GenAI
- The 5 Best Data Engineering Courses in 2026 (Honest Review) · Career
- The AI-Assisted Data Engineer: A Practical Daily Workflow · Productivity
- The AWS Data Engineering Stack: Which Service Does What (and What to Skip) · AWS
- The Snowflake Cost Optimization Playbook: Cut the Bill Without Anyone Noticing · Snowflake
- Trino Federation: Query the Lake Without Melting Prod Postgres · Trino
- Unity Catalog in Practice: Metastore Layout, Grants That Scale, and Leaving hive_metastore · Databricks
- Using LLMs Inside Data Pipelines: 4 Patterns That Actually Work in Production · AI & GenAI
- Vector Databases Compared: pgvector, Pinecone, Chroma, and When You Need None · AI & GenAI
- Vector Search in the Warehouse: When You Do Not Need a Vector Database · AI & GenAI
- Where Lambda Belongs in a Data Pipeline and Where It Quietly Becomes a Distributed System You Cannot Debug · AWS
- Window Frames Deep Dive: RANGE vs ROWS and the Patterns That Replace Self-Joins · SQL