DataLane
← All cheat sheets

AWS for data engineers cheat sheet

S3, Glue, Athena, Kinesis, and what to skip on day one.

Cloud PlatformsIntermediate3 sections

Storage and query

s3://bucket/bronze/dt=2026-08-01/
Hive-style prefixes so Athena/Glue can prune.
CREATE TABLE t ... PARTITIONED BY (dt) LOCATION s3://...
Glue catalog + Athena SQL on the lake.
MSCK REPAIR TABLE / Glue crawler
Register new partitions; prefer explicit ADD PARTITION in prod.

Ingest and transform

Kinesis Data Streams / Firehose → S3
Streams when you need sub-minute; Firehose for managed batching.
Glue job (Spark) or Glue Python shell
Serverless Spark for batch; skip EMR until you outgrow Glue.
Lambda for <15 min transforms
Great for small files; terrible as a warehouse.

Orchestration

MWAA or Step Functions
MWAA if the team already writes Airflow; Step Functions for AWS-native graphs.

From DataLane — tutorials at/blog, practice SQL live in theplayground.

↑↓ navigate openesc close