Showing posts with label scala. Show all posts
Showing posts with label scala. Show all posts

Thursday, 10 September 2026

Databricks Apache Spark-Oriented Origins

It is enlightening to reflect that Databricks grew out of the AMPLab project (AMPLab was an acronym for Algorithms, Machines and People Lab) at UC Berkeley which worked on a variety of big data projects.

AMPLab invented Apache Spark, Apache Mesos (for compute cluster management, retired in August 2025, having lost mindshare to Kubernetes which was backed by Google and CNCF) and Alluxio. Of these projects, Spark is currently the most impactful product to come out of the AMPLab project.

It was founded in 2013 and offers a cloud-based platform for data analytics and AI. It operates in Azure, AWS and GCP (the "big three" providers).

Databricks the data "lakehouse" architecture, combining aspects of data lakes and data warehouses for managing structured and unstructured data. The company develops the Delta Lake open source project adding ACID transactions to data lakes. A paper describing the challenge of adding ACID transactional capability to data lakes has been published (Armbrust et al).

To recap: ACID is an acronym which stands for atomicity, consistency, isolation and durability. These properties are used to characterize reliable database transactions and was first expounded in an influential 1983 paper "Principles of Transaction-Oriented Database Recovery" (Theo Harder and Andreas Reuter).

Tuesday, 8 September 2026

Scala vs Java - The Duel (Both JVM languages)

Scala and Java are highly compatible as both are designed to run beautifully on the JVM (the JVM spec should be read/re-read periodically for all serious Java enthusiasts).

Scala uses lambda expressions more aggressively than Java. Both use invokedynamic under the hood.

Scala dominates now in distributed systems (Spark, Akka) particularly for data engineering and ML pipelines, as well as high performance back-ends.

The following is a good website to keep up to speed on Java: Dev.java: The Destination for Java Developers.

Monday, 7 September 2026

Why Spark Works - Secrets of its Parallel Processing Powers

A Spark progam consists of a driver program that runs the user's main function and runs several parallel operations on a cluster.   It is a parallel processing engine for data. The data structure that enables this is the RDD abstraction.

RDDs are resilient distributed datasets (perhaps a better acronym could be RDDS?) - a collection of elements partitioned across nodes of a cluster that can be operated on in parallel.  

Clearly, this definition speaks to the distributed dataset aspect, but what about the resilient aspect?  Can we say it is implicit in the "can be operated on in parallel" dimension? 

We need to probe what that actually means, to uncover the "secret" of RDDs. Resilient means data remains correct and queryable in the event that nodes go down (which can happen in the physical world).

RDDs are born as files in the Hadoop system (or any Hadoop-supported file system) or an existing Scale collection. RDDs can be persisted in memory for efficient processing and they are resilient against node failures (this bit needs to be understood better - how is this achieved - redundancy of storage??).

(Footnote - once you probe deeper you will start to see ideas percolating from older frameworks like MPI in C++).



Thursday, 23 April 2026

Scala, Scala, Everywhere

For legacy observations on Scala, check out JVM stuff.  Here we build a fresh relationship with Scala.

Scala is a strongly statically typed language supporting OOP and functional programming.  Strong static typing means it avoids implicit type conversions when calling functions and other scenarios.

A good starting point for learning Scala is scala.dev here.

Apache Spark (and its roots in Scala)

Apache Spark is a foundational layer underlying many data platforms. 

It is written both in Java and Scala. Read the source code here.

A good starting point is SparkSession.scala.

One of Spark's "selling points" is "Exploratory Data Analysis (EDA) on petabyte-scale data without having to resort to downsampling" (see detailed post on downsampling). 

A petabyte (PB) holds 1000 terabytes (one thousand million million bytes).