Skip to content

Spark locally on a big machine#

A single machine with many cores and a lot of memory runs most Spark jobs without a cluster. The spark preset sets Spark up for that: local mode on the runtime's cores (spark.master local[<cores>]), the driver given 60% of the runtime's memory, and its scratch space (shuffles, spills) on the drive rather than in the container. In this recipe you make a 16-core, 64 GB environment, run a job over 200 million rows it synthesises, and watch it in Spark's UI.

Before you begin#

  • A machine with at least 16 free cores and 64 GB of free memory in your workspace, with a Data location (scratch space goes to the drive). Below, <big-machine> stands for its name.
  • The editor or admin role, and astra signed in (Install the CLI).

1. Make the environment#

  1. Hesperus → Environments → New environment.
  2. Kind: Shell. Name: spark.
  3. Under Machine, choose <big-machine>.
  4. Under Environment, choose Spark. CPU cores 16, Memory 64 GB, GPUs 0.
  5. Press Create and start.

$ astra env create spark --preset spark --cpu 16 --memory 64G --machine <big-machine> --app "Spark UI=4040"
spark is pending: `astra ssh spark` connects once it is ready (and waits for it)

With ASTRAEUS_TOKEN and API as in API:

$ curl -fsS -X POST "$API/notebook-runtimes" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' \
    -d '{"metadata": {"name": "spark"}, "spec": {"kind": "shell", "image": "spark",
         "resources": {"cpu_cores": 16, "memory_bytes": 68719476736},
         "placement": {"node": "<big-machine>"}, "apps": [{"name": "Spark UI", "port": 4040}]}}' | jq -r .status.state
Pending

The first start pulls the Spark image (803 MB).

2. Check Spark's settings#

$ astra ssh spark

In the environment, pyspark, spark-submit and spark-shell are on the PATH, already set up:

$ echo $PYSPARK_SUBMIT_ARGS
--master local[16] --driver-memory 39321m pyspark-shell
$ cat $SPARK_CONF_DIR/spark-defaults.conf
spark.master local[16]
spark.driver.memory 39321m
spark.local.dir /content/.hesperus/env/spark/spark-tmp

39321 MiB is 60% of the 64 GiB asked; the rest is for Python and the JVM's own memory.

3. Run a job#

$ mkdir -p /content/projects && cd /content/projects
$ cat > buckets.py <<'EOF'
from pyspark.sql import SparkSession, functions as F

spark = SparkSession.builder.appName("buckets").getOrCreate()
print("master:", spark.sparkContext.master)
df = spark.range(0, 200_000_000).withColumn("bucket", F.col("id") % 10).withColumn("x", F.rand(seed=1))
out = df.groupBy("bucket").agg(F.count("*").alias("n"), F.round(F.avg("x"), 2).alias("mean_x")).orderBy("bucket")
out.show()
spark.stop()
EOF
$ spark-submit buckets.py 2> spark.log
master: local[16]
+------+--------+------+
|bucket|       n|mean_x|
+------+--------+------+
|     0|20000000|   0.5|
|     1|20000000|   0.5|
…
|     9|20000000|   0.5|
+------+--------+------+

Spark's own log goes to spark.log. Shuffle files went to /content/.hesperus/env/spark/spark-tmp, on the drive.

4. Watch it in Spark's UI#

Spark's UI is served on port 4040 while an application runs. Start an interactive session and keep it open:

$ pyspark
>>> spark.range(0, 2_000_000_000).selectExpr("sum(id)").show()

While it runs, open the UI:

On the environment's page, under Apps, press Open on Spark UI (or Open port… → Spark UI · 4040). The Jobs and Stages tabs show the 16 tasks running at once.

$ astra env port spark 4040
Opened https://r-4f1c9a0b2d7e83a65c10-4040.<runtime domain>

$ curl -fsS -X POST "$CONSOLE/clusters/<cluster>/runtime-links" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d '{"runtime": "spark", "port": 4040}' | jq -r .url

Open the URL within 60 seconds; it works once.

When the session ends, the UI does too: the app's address answers again when the next application starts.

Notes#

  • Your data: when you make the environment, mount a drive of the workspace beside its own (Drives → Attach a drive, or --drive <drive>:/data:ro) and read it with spark.read.parquet("/data/…"). An environment's drives cannot be changed afterwards: delete it and make it again with the same name (its home drive stays).
  • Idle: an open SSH session counts as activity; the environment stops after 60 minutes with nobody connected. A job started with nohup keeps running only while the environment does: raise the idle stop (--idle, up to 1440 minutes) for long jobs, or make it a run.
  • In a notebook: a notebook's runtime with the Spark preset has the same settings; from pyspark.sql import SparkSession works in a cell.