Spark locally on a big machine#
A single machine with many cores and a lot of memory runs most Spark
jobs without a cluster. The spark preset sets Spark up for that: local
mode on the runtime's cores (spark.master local[<cores>]), the driver
given 60% of the runtime's memory, and its scratch space (shuffles,
spills) on the drive rather than in the container. In this recipe you make
a 16-core, 64 GB environment, run a job over 200 million rows it
synthesises, and watch it in Spark's UI.
Before you begin#
- A machine with at least 16 free cores and 64 GB of free memory in your
workspace, with a Data location (scratch space goes to the drive).
Below,
<big-machine>stands for its name. - The editor or admin role, and
astrasigned in (Install the CLI).
1. Make the environment#
- Hesperus → Environments → New environment.
- Kind: Shell. Name:
spark. - Under Machine, choose
<big-machine>. - Under Environment, choose Spark. CPU cores
16, Memory64GB, GPUs0. - Press Create and start.
With ASTRAEUS_TOKEN and API as in API:
$ curl -fsS -X POST "$API/notebook-runtimes" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
-H 'content-type: application/json' \
-d '{"metadata": {"name": "spark"}, "spec": {"kind": "shell", "image": "spark",
"resources": {"cpu_cores": 16, "memory_bytes": 68719476736},
"placement": {"node": "<big-machine>"}, "apps": [{"name": "Spark UI", "port": 4040}]}}' | jq -r .status.state
Pending
The first start pulls the Spark image (803 MB).
2. Check Spark's settings#
In the environment, pyspark, spark-submit and spark-shell are on the
PATH, already set up:
$ echo $PYSPARK_SUBMIT_ARGS
--master local[16] --driver-memory 39321m pyspark-shell
$ cat $SPARK_CONF_DIR/spark-defaults.conf
spark.master local[16]
spark.driver.memory 39321m
spark.local.dir /content/.hesperus/env/spark/spark-tmp
39321 MiB is 60% of the 64 GiB asked; the rest is for Python and the JVM's own memory.
3. Run a job#
$ mkdir -p /content/projects && cd /content/projects
$ cat > buckets.py <<'EOF'
from pyspark.sql import SparkSession, functions as F
spark = SparkSession.builder.appName("buckets").getOrCreate()
print("master:", spark.sparkContext.master)
df = spark.range(0, 200_000_000).withColumn("bucket", F.col("id") % 10).withColumn("x", F.rand(seed=1))
out = df.groupBy("bucket").agg(F.count("*").alias("n"), F.round(F.avg("x"), 2).alias("mean_x")).orderBy("bucket")
out.show()
spark.stop()
EOF
$ spark-submit buckets.py 2> spark.log
master: local[16]
+------+--------+------+
|bucket| n|mean_x|
+------+--------+------+
| 0|20000000| 0.5|
| 1|20000000| 0.5|
…
| 9|20000000| 0.5|
+------+--------+------+
Spark's own log goes to spark.log. Shuffle files went to
/content/.hesperus/env/spark/spark-tmp, on the drive.
4. Watch it in Spark's UI#
Spark's UI is served on port 4040 while an application runs. Start an interactive session and keep it open:
While it runs, open the UI:
On the environment's page, under Apps, press Open on Spark UI (or Open port… → Spark UI · 4040). The Jobs and Stages tabs show the 16 tasks running at once.
When the session ends, the UI does too: the app's address answers again when the next application starts.
Notes#
- Your data: when you make the environment, mount a drive of the
workspace beside its own (Drives → Attach a drive, or
--drive <drive>:/data:ro) and read it withspark.read.parquet("/data/…"). An environment's drives cannot be changed afterwards: delete it and make it again with the same name (its home drive stays). - Idle: an open SSH session counts as activity; the environment stops
after 60 minutes with nobody connected. A job started with
nohupkeeps running only while the environment does: raise the idle stop (--idle, up to 1440 minutes) for long jobs, or make it a run. - In a notebook: a notebook's runtime with the Spark preset has the
same settings;
from pyspark.sql import SparkSessionworks in a cell.