1. Zaharia, M., Chowdhury, M., Franklin, M. J., Shenker, S., & Stoica, I. (2012). Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing. In Proceedings of the 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI '12).
Section 3, "Resilient Distributed Datasets (RDDs)": "The most important property of RDDs is that they are fault-tolerant. RDDs track the graph of transformations that were used to build them (their lineage) to recompute lost data. [...] if a partition of an RDD is lost, the RDD has enough information about how it was derived from other datasets to recompute just that partition." This directly explains the mechanism that prevents a single failure from crashing the application.
DOI: https://doi.org/10.5555/2228298.2228301
2. Official Apache Spark Documentation. Application Scheduling - Fault Tolerance.
Section: "Fault Tolerance for Tasks": The documentation states, "The driver will also re-execute tasks if they fail. It will retry a configurable number of times (spark.task.maxFailures) before giving up on the job." This confirms that task failures lead to retries, not immediate application failure.
3. Karau, H., Konwinski, A., Wendell, P., & Zaharia, M. (2015). Learning Spark: Lightning-Fast Big Data Analysis. O'Reilly Media.
Chapter 2, "Downloading Spark and Getting Started": In the section on Spark's components, the text describes the fault-tolerance mechanism: "If a worker running a task fails, Spark will rerun the task on another node." This is a foundational concept introduced early in official learning materials.
4. University of California, Berkeley. CS 186/286: Introduction to Database Systems / Advanced Topics in Database Systems.
Lecture Notes on Spark: Course materials frequently cover Spark's architecture. The lectures on RDDs emphasize their immutability and lineage as the basis for fault tolerance. They explain that unlike systems that rely on data replication for fault tolerance, Spark's lineage-based approach allows it to recover lost data partitions by re-executing the necessary transformations, thus handling task failures gracefully without halting the entire application.