{"product_id":"9781785889585","title":"Learning Apache Spark 2","description":"\u003cp\u003e\u003cb\u003eLearn about the fastest-growing open source project in the world, and find out how it revolutionizes big data analytics\u003c\/b\u003e\u003c\/p\u003e\u003cp\u003e\u003cb\u003eAbout This Book\u003c\/b\u003e\u003c\/p\u003e\u003cul\u003e\n\u003cli\u003eExclusive guide that covers how to get up and running with fast data processing using Apache Spark\u003c\/li\u003e\n\u003cli\u003eExplore and exploit various possibilities with Apache Spark using real-world use cases in this book\u003c\/li\u003e\n\u003cli\u003eWant to perform efficient data processing at real time? This book will be your one-stop solution.\u003c\/li\u003e\n\u003c\/ul\u003e\u003cb\u003eWho This Book Is For\u003c\/b\u003e\u003cp\u003eThis guide appeals to big data engineers, analysts, architects, software engineers, even technical managers who need to perform efficient data processing on Hadoop at real time. Basic familiarity with Java or Scala will be helpful.\u003c\/p\u003e\u003cp\u003e\u003c\/p\u003e\u003cp\u003eThe assumption is that readers will be from a mixed background, but would be typically people with background in engineering\/data science with no prior Spark experience and want to understand how Spark can help them on their analytics journey.\u003c\/p\u003e\u003cp\u003e\u003cb\u003eWhat You Will Learn\u003c\/b\u003e\u003c\/p\u003e\u003cul\u003e\n\u003cli\u003eGet an overview of big data analytics and its importance for organizations and data professionals\u003c\/li\u003e\n\u003cli\u003eDelve into Spark to see how it is different from existing processing platforms\u003c\/li\u003e\n\u003cli\u003eUnderstand the intricacies of various file formats, and how to process them with Apache Spark.\u003c\/li\u003e\n\u003cli\u003eRealize how to deploy Spark with YARN, MESOS or a Stand-alone cluster manager.\u003c\/li\u003e\n\u003cli\u003eLearn the concepts of Spark SQL, SchemaRDD, Caching and working with Hive and Parquet file formats\u003c\/li\u003e\n\u003cli\u003eUnderstand the architecture of Spark MLLib while discussing some of the off-the-shelf algorithms that come with Spark.\u003c\/li\u003e\n\u003cli\u003eIntroduce yourself to the deployment and usage of SparkR.\u003c\/li\u003e\n\u003cli\u003eWalk through the importance of Graph computation and the graph processing systems available in the market\u003c\/li\u003e\n\u003cli\u003eCheck the real world example of Spark by building a recommendation engine with Spark using ALS.\u003c\/li\u003e\n\u003cli\u003eUse a Telco data set, to predict customer churn using Random Forests.\u003c\/li\u003e\n\u003c\/ul\u003e\u003cb\u003eIn Detail\u003c\/b\u003e\u003cp\u003eSpark juggernaut keeps on rolling and getting more and more momentum each day. Spark provides key capabilities in the form of Spark SQL, Spark Streaming, Spark ML and Graph X all accessible via Java, Scala, Python and R. Deploying the key capabilities is crucial whether it is on a Standalone framework or as a part of existing Hadoop installation and configuring with Yarn and Mesos.\u003c\/p\u003e\u003cp\u003e\u003c\/p\u003e\u003cp\u003eThe next part of the journey after installation is using key components, APIs, Clustering, machine learning APIs, data pipelines, parallel programming. It is important to understand why each framework component is key, how widely it is being used, its stability and pertinent use cases.\u003c\/p\u003e\u003cp\u003e\u003c\/p\u003e\u003cp\u003eOnce we understand the individual components, we will take a couple of real life advanced analytics examples such as 'Building a Recommendation system', 'Predicting customer churn' and so on.\u003c\/p\u003e\u003cp\u003e\u003c\/p\u003e\u003cp\u003eThe objective of these real life examples is to give the reader confidence of using Spark for real-world problems.\u003c\/p\u003e\u003cp\u003e\u003cb\u003eStyle and approach\u003c\/b\u003e\u003c\/p\u003e\u003cp\u003eWith the help of practical examples and real-world use cases, this guide will take you from scratch to building efficient data applications using Apache Spark.\u003c\/p\u003e\u003cp\u003e\u003c\/p\u003e\u003cp\u003eYou will learn all about this excellent data processing engine in a step-by-step manner, taking one aspect of it at a time.\u003c\/p\u003e\u003cp\u003e\u003c\/p\u003e\u003cp\u003eThis highly practical guide will include how to work with data pipelines, dataframes, clustering, SparkSQL, parallel programming, and such insightful topics with the help of real-world use cases.\u003c\/p\u003e\u003cp\u003e\u003cbr\u003e\u003c\/p\u003e","brand":"Packt Publishing","offers":[{"title":"Default Title","offer_id":47158315352304,"sku":"9781785889585","price":35.99,"currency_code":"USD","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0737\/7593\/9824\/files\/9781785889585_p0.jpg?v=1763729699","url":"https:\/\/shop-qa.barnesandnoble.com\/products\/9781785889585","provider":"Barnes \u0026 Noble (DEV)","version":"1.0","type":"link"}