{"id":1563,"date":"2026-09-14T12:54:32","date_gmt":"2026-09-14T12:54:32","guid":{"rendered":"https:\/\/prwatech.in\/blog\/?p=1563"},"modified":"2026-09-15T05:29:36","modified_gmt":"2026-09-15T05:29:36","slug":"introduction-to-apache-spark","status":"publish","type":"post","link":"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/","title":{"rendered":"Introduction to Apache Spark"},"content":{"rendered":"<h1>Introduction to Apache Spark<\/h1>\n<p><a href=\"https:\/\/prwatech.in\/apache-spark-training-institute-in-bangalore\/\"><strong><b>Apache Spark<\/b><\/strong><\/a>\u00a0is a fast, general purpose engine for processing large amounts of data across a cluster of machines. It offers high level APIs in Java, Scala, Python, and R, backed by an optimized execution engine, along with higher level tools for SQL and structured data processing, machine learning, graph processing, and stream processing, all built on the same core engine.<\/p>\n<h2>History of Apache Spark<\/h2>\n<p>Spark began life in 2009 as a research project inside UC Berkeley&#8217;s AMPLab. It was released as open source in 2010 under the BSD license, then donated to the Apache Software Foundation in 2013, becoming a top level Apache project in 2014. Development has continued steadily since, and Spark is now in its fourth major version line, still actively maintained by the Apache community.<\/p>\n<h2>Why Spark Was Needed<\/h2>\n<p>Before Spark, different jobs typically needed different specialized tools. Hadoop MapReduce handled batch processing. Apache Storm or S4 handled stream processing. Apache Impala or Tez handled interactive queries. Neo4j or Apache Giraph handled graph processing. No single engine covered all of these well, and none of them offered fast, in memory processing with sub second response times.<\/p>\n<p>Spark&#8217;s pitch was to be that one engine, real time streaming, interactive queries, graph processing, and batch processing, all handled through in memory processing that runs meaningfully faster than reading and writing to disk at every step. That combination, plus a consistent interface across all of it, is what set Spark apart from Hadoop and from single purpose tools like Storm.<\/p>\n<h2 style=\"text-align: left;\">Components of Apache Spark<\/h2>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-1564\" src=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/1-25.png\" alt=\"Components of Apache Spark\" width=\"850\" height=\"478\" srcset=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/1-25.png 523w, https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/1-25-300x169.png 300w\" sizes=\"auto, (max-width: 850px) 100vw, 850px\" \/><\/p>\n<ul>\n<li><b><\/b><strong><b>Spark Core, <\/b><\/strong>the foundation everything else runs on top of, providing the execution platform for every Spark application.<\/li>\n<li><b><\/b><strong><b>Spark SQL, <\/b><\/strong>which lets you run SQL and HiveQL queries against structured and semi structured data, often meaningfully faster than the same queries running unmodified on older systems.<\/li>\n<li><b><\/b><strong><b>Spark Streaming, <\/b><\/strong>which processes live data by breaking it into small micro batches that run on top of Spark Core.<\/li>\n<li><b><\/b><strong><b>MLlib, <\/b><\/strong>Spark&#8217;s machine learning library, built to take advantage of in memory processing for iterative algorithms that would otherwise be slow.<\/li>\n<li><b><\/b><strong><b>GraphX, <\/b><\/strong>a graph computation engine for processing graph structured data at scale.<\/li>\n<li><strong><b>SparkR, <\/b><\/strong>a lightweight front end that lets R users analyze large datasets and run jobs interactively from the R shell, combining R&#8217;s usability with Spark&#8217;s scalability.<\/li>\n<\/ul>\n<h2>Apache Spark architecture overview<\/h2>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-1565\" src=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/2-13.png\" alt=\"apache spark architecture overview\" width=\"850\" height=\"447\" srcset=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/2-13.png 699w, https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/2-13-300x158.png 300w\" sizes=\"auto, (max-width: 850px) 100vw, 850px\" \/><\/p>\n<p>Spark has a layered architecture where its components stay loosely coupled, which is what makes it straightforward to extend with additional libraries. The architecture rests on two core abstractions: the Resilient Distributed Dataset, or RDD, and the Directed Acyclic Graph or DAG, which represents the sequence of steps in a computation.<\/p>\n<h2><strong>How Apache Spark Works\u00a0<\/strong><\/h2>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-1570\" src=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/3-11.png\" alt=\"How does Apache Spark work\" width=\"850\" height=\"652\" srcset=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/3-11.png 617w, https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/3-11-300x230.png 300w\" sizes=\"auto, (max-width: 850px) 100vw, 850px\" \/><\/p>\n<p>Like Hadoop MapReduce, Spark distributes data across a cluster and processes it in parallel. It uses a master and worker style architecture, one central coordinator and many distributed workers. That central coordinator is called the <strong><b>driver<\/b><\/strong>.<\/p>\n<h2>Apache Spark Terminology<\/h2>\n<h3>Spark Context<\/h3>\n<p>The Spark Context is the entry point of a Spark application. It connects to the Spark execution environment and is used to create RDDs, accumulators, and broadcast variables, access Spark services, and run jobs. Its responsibilities include:<\/p>\n<ul>\n<li>Checking the current status of a Spark application<\/li>\n<li>Canceling a job or a stage<\/li>\n<li>Running a job synchronously or asynchronously<\/li>\n<li>Accessing or unpersisting a cached RDD<\/li>\n<li>Programmable dynamic allocation of resources<\/li>\n<\/ul>\n<h3>Spark Shell<\/h3>\n<p>The Spark Shell is a Spark application written in Scala that gives you a command line environment with auto completion, useful for getting familiar with Spark&#8217;s features before writing a standalone application of your own.<\/p>\n<h3>Spark Application<\/h3>\n<p>A Spark application is a self contained computation that runs user supplied code to produce a result. It can have processes running on its behalf even when it is not actively running a job.<\/p>\n<h3>Task, Job, and Stage<\/h3>\n<ul>\n<li><b><\/b><strong><b>Task, <\/b><\/strong>the smallest unit of work, sent to a single executor. Each partition of an RDD gets its own task.<\/li>\n<li><b><\/b><strong><b>Job, <\/b><\/strong>a parallel computation made up of multiple tasks, created in response to an action.<\/li>\n<li><b><\/b><strong><b>Stage, <\/b><\/strong>a set of tasks within a job that depend on each other. A job is split into stages because not every computation can happen in one continuous pass, some steps have to wait on others to finish first.<\/li>\n<\/ul>\n<h2>Spark Runtime Architecture<\/h2>\n<h3>The Driver<\/h3>\n<p>The driver runs the main method of your program. It is the process that creates RDDs, performs transformations and actions on them, and creates the Spark Context. Launching the Spark Shell is itself an example of starting a driver program, and the application ends once the driver terminates.<\/p>\n<p>The driver splits a Spark application into tasks and schedules them onto executors, through a task scheduler that lives inside the driver itself. Its two key jobs are converting your program into tasks, and scheduling those tasks on executors.<\/p>\n<p>At a high level, a Spark program takes some input data as an RDD, derives new RDDs from it through transformations, then runs an action to actually compute a result. This builds up a DAG of operations implicitly as you write your code, and when the driver runs, it converts that DAG into an actual physical execution plan.<\/p>\n<h3>The Cluster Manager<\/h3>\n<p>Spark relies on a cluster manager to launch executors, and in some setups, to launch the driver as well. The cluster manager is a pluggable component, jobs and actions within a Spark application are scheduled by Spark&#8217;s own scheduler, by default in FIFO order, though round robin scheduling is also available. Applications can grow or shrink their resource usage dynamically based on workload, freeing resources when idle and requesting more again when demand returns. Spark currently supports this dynamic allocation across standalone mode, YARN, and Kubernetes. Apache Mesos was a supported option in earlier Spark versions, but that support was deprecated in Spark 3.2 and fully removed in Spark 4.0.<\/p>\n<h3>Executors<\/h3>\n<p>Executors run the individual tasks that make up a Spark job. They launch once at the start of an application and stay alive for its entire lifetime, and if one executor fails, the application can generally continue without stopping entirely. Executors have two main jobs: running the tasks that make up the application and returning results to the driver, and providing in memory storage for any RDDs the user has chosen to cache.<\/p>\n<h2>Resilient Distributed Dataset, RDD<\/h2>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-1566\" src=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/4-10.png\" alt=\"Resilient Distributed\u00a0Dataset(RDD):\" width=\"850\" height=\"350\" srcset=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/4-10.png 693w, https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/4-10-300x123.png 300w\" sizes=\"auto, (max-width: 850px) 100vw, 850px\" \/><\/p>\n<p>RDDs are\u00a0the building blocks of any Spark application. RDDs Stands for:<\/p>\n<ul>\n<li><strong><em>Resilient:<\/em><\/strong>\u00a0Fault tolerant and is capable of rebuilding data on failure<\/li>\n<li><strong><em>Distributed:<\/em><\/strong>\u00a0Distributed data among the multiple nodes in a cluster<\/li>\n<li><strong><em>Dataset:<\/em><\/strong>\u00a0Collection of partitioned data with values<\/li>\n<\/ul>\n<h3>RDD Persistence :<\/h3>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-1567\" src=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/5-6.png\" alt=\"introduction to apache spark\" width=\"850\" height=\"433\" srcset=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/5-6.png 672w, https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/5-6-300x153.png 300w\" sizes=\"auto, (max-width: 850px) 100vw, 850px\" \/><\/p>\n<p>One of Spark&#8217;s most useful capabilities is persisting or caching, a dataset in memory across multiple operations. Once an RDD is persisted, each node keeps the partitions it computed in memory and reuses them for later actions on that same dataset or on datasets derived from it, which can make later operations dramatically faster, often by an order of magnitude, particularly for iterative algorithms and interactive analysis.<\/p>\n<p>You mark an RDD for persistence using persist or cache. The first action that computes it will keep it in memory going forward. Spark&#8217;s cache is itself fault tolerant, if a partition is lost, Spark automatically recomputes it using the same transformations that built it originally.<\/p>\n<p>A persisted RDD can use different storage strategies, called storage levels, depending on your priorities:<\/p>\n<ul>\n<li><b><\/b><strong><b>MEMORY_ONLY, <\/b><\/strong>the default, keeps deserialized objects in memory. Partitions that do not fit are recomputed on demand rather than cached.<\/li>\n<li><b><\/b><strong><b>MEMORY_AND_DISK, <\/b><\/strong>keeps deserialized objects in memory, but spills partitions that do not fit onto disk instead of recomputing them.<\/li>\n<li><b><\/b><strong><b>MEMORY_ONLY_SER and MEMORY_AND_DISK_SER, <\/b><\/strong>available in Java and Scala, store serialized objects instead, which uses less memory at the cost of more CPU work to read.<\/li>\n<li><b><\/b><strong><b>DISK_ONLY, <\/b><\/strong>stores partitions only on disk.<\/li>\n<li><b><\/b><strong><b>Replicated variants, <\/b><\/strong>such as MEMORY_ONLY_2, work the same as their base level but keep a copy of each partition on two separate nodes.<\/li>\n<li><b><\/b><strong><b>OFF_HEAP, <\/b><\/strong>an experimental option that stores serialized data outside the JVM heap, which requires off heap memory to be enabled.<\/li>\n<\/ul>\n<h2>Deploying Spark: Master URLs<\/h2>\n<p>When you launch a Spark application, the master URL tells Spark where and how to run:<\/p>\n<ul>\n<li><b><\/b><strong><b>local, <\/b><\/strong>runs Spark on one worker thread, with no real parallelism, generally used only for quick local testing.<\/li>\n<li><b><\/b><strong><b>local[K], <\/b><\/strong>runs Spark locally with K worker threads, typically set to match the number of cores on your machine.<\/li>\n<li><b><\/b><strong><b>local[*], <\/b><\/strong>runs Spark locally using as many worker threads as there are logical cores available.<\/li>\n<li><b><\/b><strong><b>spark colon slash slash HOST colon PORT, <\/b><\/strong>connects to a Spark standalone cluster at the given host and port.<\/li>\n<li><b><\/b><strong><b>yarn, <\/b><\/strong>runs on a YARN managed cluster, using whichever configuration is set for client or cluster deploy mode.<\/li>\n<li><b><\/b><strong><b>k8s colon slash slash HOST colon PORT, <\/b><\/strong>runs on a Kubernetes cluster, the current standard option for containerized Spark deployments.<\/li>\n<\/ul>\n<h2>Deployment Modes Compared<\/h2>\n<p>YARN can run Spark in two different modes, cluster and client, and standalone mode offers a third option that does not depend on YARN at all:<\/p>\n<table>\n<tbody>\n<tr>\n<td width=\"133\"><\/td>\n<td width=\"160\"><strong><b>YARN Cluster<\/b><\/strong><\/td>\n<td width=\"160\"><strong><b>YARN Client<\/b><\/strong><\/td>\n<td width=\"160\"><strong><b>Standalone<\/b><\/strong><\/td>\n<\/tr>\n<tr>\n<td width=\"133\">Driver runs in<\/td>\n<td width=\"160\">Application master<\/td>\n<td width=\"160\">Client machine<\/td>\n<td width=\"160\">Client machine<\/td>\n<\/tr>\n<tr>\n<td width=\"133\">Requests resources<\/td>\n<td width=\"160\">Application master<\/td>\n<td width=\"160\">Application master<\/td>\n<td width=\"160\">Client machine<\/td>\n<\/tr>\n<tr>\n<td width=\"133\">Starts executor processes<\/td>\n<td width=\"160\">YARN NodeManager<\/td>\n<td width=\"160\">YARN NodeManager<\/td>\n<td width=\"160\">Spark worker<\/td>\n<\/tr>\n<tr>\n<td width=\"133\">Long running services<\/td>\n<td width=\"160\">YARN ResourceManager, NodeManager<\/td>\n<td width=\"160\">YARN ResourceManager, NodeManager<\/td>\n<td width=\"160\">Spark master, worker<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-1568\" src=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/6-5.png\" alt=\"apache spark use cases\" width=\"850\" height=\"438\" srcset=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/6-5.png 666w, https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/6-5-300x155.png 300w\" sizes=\"auto, (max-width: 850px) 100vw, 850px\" \/><\/p>\n<h2>Real World Use Cases of Apache Spark<\/h2>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-1569\" src=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/8-4.png\" alt=\"apache spark tutorial\" width=\"850\" height=\"515\" srcset=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/8-4.png 672w, https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/8-4-300x182.png 300w\" sizes=\"auto, (max-width: 850px) 100vw, 850px\" \/><\/p>\n<ul>\n<li><b><\/b><strong><b>Large scale batch processing, <\/b><\/strong>such as transforming and aggregating logs or transaction data across a data lake.<\/li>\n<li><b><\/b><strong><b>Real time analytics, <\/b><\/strong>such as monitoring fraud signals or system metrics as events arrive, using Structured Streaming.<\/li>\n<li><b><\/b><strong><b>Machine learning at scale, <\/b><\/strong>training models against datasets too large to fit comfortably on a single machine, using MLlib.<\/li>\n<li><b><\/b><strong><b>ETL pipelines, <\/b><\/strong>reading from multiple sources, transforming the data, and loading it into a warehouse or lake using Spark SQL and DataFrames.<\/li>\n<li><b><\/b><strong><b>Graph analysis, <\/b><\/strong>such as analyzing social networks or recommendation graphs using GraphX.<\/li>\n<\/ul>\n<h2>Common Misconceptions to Watch For<\/h2>\n<ul>\n<li>Assuming Spark replaces Hadoop entirely. Spark commonly runs on top of Hadoop&#8217;s YARN resource manager and reads from HDFS, it replaces MapReduce as the processing engine, not the whole Hadoop ecosystem.<\/li>\n<li>Assuming Mesos is still a supported way to run Spark. It was removed in Spark 4.0. Standalone mode, YARN, and Kubernetes are the current options.<\/li>\n<li>Persisting every RDD by default. Caching costs memory. It pays off when a dataset is reused across multiple actions, but persisting data you only use once just adds overhead for no benefit.<\/li>\n<\/ul>\n<p>That covers Apache Spark&#8217;s architecture, terminology, deployment options, and real world use cases. To go further, explore <a href=\"https:\/\/prwatech.in\/apache-spark-training-institute-in-pune\/\"><strong><b>Prwatech&#8217;s Apache Spark training<\/b><\/strong>\u00a0program<\/a>, which includes placement assistance.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Introduction to Apache Spark Apache Spark\u00a0is a fast, general purpose engine for processing large amounts of data across a cluster of machines. It offers high level APIs in Java, Scala, Python, and R, backed by an optimized execution engine, along with higher level tools for SQL and structured data processing, machine learning, graph processing, and [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[21,1689],"tags":[774,878,877,873,285,875,876,874,871,879],"class_list":["post-1563","post","type-post","status-publish","format-standard","hentry","category-apache-spark","category-apache-spark-introduction","tag-apache-spark","tag-apache-spark-basics","tag-apache-spark-for-beginners","tag-apache-spark-intro","tag-introduction-to-apache-spark","tag-introduction-to-spark","tag-learn-apache-spark","tag-spark-basics","tag-spark-tutorial","tag-working-of-apache-spark"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v25.7 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Introduction to Apache Spark: Components, Architecture &amp;Uses<\/title>\n<meta name=\"description\" content=\"Learn what Apache Spark is, its components, architecture, RDDs, deployment modes, and real world use cases, with terminology explained by PrwaTech.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Introduction to Apache Spark: Components, Architecture &amp;Uses\" \/>\n<meta property=\"og:description\" content=\"Learn what Apache Spark is, its components, architecture, RDDs, deployment modes, and real world use cases, with terminology explained by PrwaTech.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/\" \/>\n<meta property=\"og:site_name\" content=\"Prwatech\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/prwatech.in\/\" \/>\n<meta property=\"article:published_time\" content=\"2026-09-14T12:54:32+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-09-15T05:29:36+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/1-25.png\" \/>\n\t<meta property=\"og:image:width\" content=\"523\" \/>\n\t<meta property=\"og:image:height\" content=\"294\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Prwatech\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@Eduprwatech\" \/>\n<meta name=\"twitter:site\" content=\"@Eduprwatech\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Prwatech\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"10 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/\",\"url\":\"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/\",\"name\":\"Introduction to Apache Spark: Components, Architecture &Uses\",\"isPartOf\":{\"@id\":\"https:\/\/prwatech.in\/blog\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/#primaryimage\"},\"image\":{\"@id\":\"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/#primaryimage\"},\"thumbnailUrl\":\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/1-25.png\",\"datePublished\":\"2026-09-14T12:54:32+00:00\",\"dateModified\":\"2026-09-15T05:29:36+00:00\",\"author\":{\"@id\":\"https:\/\/prwatech.in\/blog\/#\/schema\/person\/db90baff7744090b2288bbc98fea87f3\"},\"description\":\"Learn what Apache Spark is, its components, architecture, RDDs, deployment modes, and real world use cases, with terminology explained by PrwaTech.\",\"breadcrumb\":{\"@id\":\"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/#primaryimage\",\"url\":\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/1-25.png\",\"contentUrl\":\"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/1-25.png\",\"width\":523,\"height\":294},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/prwatech.in\/blog\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Introduction to Apache Spark\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/prwatech.in\/blog\/#website\",\"url\":\"https:\/\/prwatech.in\/blog\/\",\"name\":\"Prwatech\",\"description\":\"Share Ideas, Start Something Good.\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/prwatech.in\/blog\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Person\",\"@id\":\"https:\/\/prwatech.in\/blog\/#\/schema\/person\/db90baff7744090b2288bbc98fea87f3\",\"name\":\"Prwatech\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/prwatech.in\/blog\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/c00bafc1b04045f31eda917de39891456c44fa47c092b9bb6be0f860a3a30a2f?s=96&d=mm&r=g\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/c00bafc1b04045f31eda917de39891456c44fa47c092b9bb6be0f860a3a30a2f?s=96&d=mm&r=g\",\"caption\":\"Prwatech\"},\"url\":\"https:\/\/prwatech.in\/blog\/author\/prwatech123\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Introduction to Apache Spark: Components, Architecture &Uses","description":"Learn what Apache Spark is, its components, architecture, RDDs, deployment modes, and real world use cases, with terminology explained by PrwaTech.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/","og_locale":"en_US","og_type":"article","og_title":"Introduction to Apache Spark: Components, Architecture &Uses","og_description":"Learn what Apache Spark is, its components, architecture, RDDs, deployment modes, and real world use cases, with terminology explained by PrwaTech.","og_url":"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/","og_site_name":"Prwatech","article_publisher":"https:\/\/www.facebook.com\/prwatech.in\/","article_published_time":"2026-09-14T12:54:32+00:00","article_modified_time":"2026-09-15T05:29:36+00:00","og_image":[{"width":523,"height":294,"url":"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/1-25.png","type":"image\/png"}],"author":"Prwatech","twitter_card":"summary_large_image","twitter_creator":"@Eduprwatech","twitter_site":"@Eduprwatech","twitter_misc":{"Written by":"Prwatech","Est. reading time":"10 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/","url":"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/","name":"Introduction to Apache Spark: Components, Architecture &Uses","isPartOf":{"@id":"https:\/\/prwatech.in\/blog\/#website"},"primaryImageOfPage":{"@id":"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/#primaryimage"},"image":{"@id":"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/#primaryimage"},"thumbnailUrl":"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/1-25.png","datePublished":"2026-09-14T12:54:32+00:00","dateModified":"2026-09-15T05:29:36+00:00","author":{"@id":"https:\/\/prwatech.in\/blog\/#\/schema\/person\/db90baff7744090b2288bbc98fea87f3"},"description":"Learn what Apache Spark is, its components, architecture, RDDs, deployment modes, and real world use cases, with terminology explained by PrwaTech.","breadcrumb":{"@id":"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/#primaryimage","url":"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/1-25.png","contentUrl":"https:\/\/prwatech.in\/blog\/wp-content\/uploads\/2019\/04\/1-25.png","width":523,"height":294},{"@type":"BreadcrumbList","@id":"https:\/\/prwatech.in\/blog\/apache-spark\/apache-spark-introduction\/introduction-to-apache-spark\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/prwatech.in\/blog\/"},{"@type":"ListItem","position":2,"name":"Introduction to Apache Spark"}]},{"@type":"WebSite","@id":"https:\/\/prwatech.in\/blog\/#website","url":"https:\/\/prwatech.in\/blog\/","name":"Prwatech","description":"Share Ideas, Start Something Good.","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/prwatech.in\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Person","@id":"https:\/\/prwatech.in\/blog\/#\/schema\/person\/db90baff7744090b2288bbc98fea87f3","name":"Prwatech","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/prwatech.in\/blog\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/c00bafc1b04045f31eda917de39891456c44fa47c092b9bb6be0f860a3a30a2f?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/c00bafc1b04045f31eda917de39891456c44fa47c092b9bb6be0f860a3a30a2f?s=96&d=mm&r=g","caption":"Prwatech"},"url":"https:\/\/prwatech.in\/blog\/author\/prwatech123\/"}]}},"_links":{"self":[{"href":"https:\/\/prwatech.in\/blog\/wp-json\/wp\/v2\/posts\/1563","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/prwatech.in\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/prwatech.in\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/prwatech.in\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/prwatech.in\/blog\/wp-json\/wp\/v2\/comments?post=1563"}],"version-history":[{"count":20,"href":"https:\/\/prwatech.in\/blog\/wp-json\/wp\/v2\/posts\/1563\/revisions"}],"predecessor-version":[{"id":11765,"href":"https:\/\/prwatech.in\/blog\/wp-json\/wp\/v2\/posts\/1563\/revisions\/11765"}],"wp:attachment":[{"href":"https:\/\/prwatech.in\/blog\/wp-json\/wp\/v2\/media?parent=1563"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/prwatech.in\/blog\/wp-json\/wp\/v2\/categories?post=1563"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/prwatech.in\/blog\/wp-json\/wp\/v2\/tags?post=1563"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}