{"product_id":"spark-9781119254010","title":"Spark","description":"\u003cb\u003eBook Synopsis\u003c\/b\u003e\u003cbr\u003e\u003cb\u003eProduction-targeted Spark guidance with real-world use cases\u003c\/b\u003e \u003cp\u003e\u003ci\u003eSpark: Big Data Cluster Computing in Production\u003c\/i\u003e goes beyond general Spark overviews to provide targeted guidance toward using lightning-fast big-data clustering in production. Written by an expert team well-known in the big data community, this book walks you through the challenges in moving from proof-of-concept or demo Spark applications to live Spark in production. Real use cases provide deep insight into common problems, limitations, challenges, and opportunities, while expert tips and tricks help you get the most out of Spark performance. Coverage includes Spark SQL, Tachyon, Kerberos, ML Lib, YARN, and Mesos, with clear, actionable guidance on resource scheduling, db connectors, streaming, security, and much more. \u003c\/p\u003e\u003cp\u003eSpark has become the tool of choice for many Big Data problems, with more active contributors than any other Apache Software project. General introductory books abound, but this book is the\u003cbr\u003e\u003cbr\u003e\u003cb\u003eTable of Contents\u003c\/b\u003e\u003cbr\u003eIntroduction xix \u003c\/p\u003e\u003cp\u003e\u003cb\u003eChapter 1 Finishing Your Spark Job 1\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003eInstallation of the Necessary Components 2\u003c\/p\u003e \u003cp\u003eNative Installation Using a Spark Standalone Cluster 3\u003c\/p\u003e \u003cp\u003eThe History of Distributed Computing That Led to Spark 3\u003c\/p\u003e \u003cp\u003eEnter the Cloud 4\u003c\/p\u003e \u003cp\u003eUnderstanding Resource Management 5\u003c\/p\u003e \u003cp\u003eUsing Various Formats for Storage 8\u003c\/p\u003e \u003cp\u003eText Files 10\u003c\/p\u003e \u003cp\u003eSequence Files 11\u003c\/p\u003e \u003cp\u003eAvro Files 11\u003c\/p\u003e \u003cp\u003eParquet Files 12\u003c\/p\u003e \u003cp\u003eMaking Sense of Monitoring and Instrumentation 13\u003c\/p\u003e \u003cp\u003eSpark UI 13\u003c\/p\u003e \u003cp\u003eSpark Standalone UI 15\u003c\/p\u003e \u003cp\u003eMetrics REST API 16\u003c\/p\u003e \u003cp\u003eMetrics System 16\u003c\/p\u003e \u003cp\u003eExternal Monitoring Tools 16\u003c\/p\u003e \u003cp\u003eSummary 17\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 2 Cluster Management 19\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003eBackground 21\u003c\/p\u003e \u003cp\u003eSpark Components 24\u003c\/p\u003e \u003cp\u003eDriver 25\u003c\/p\u003e \u003cp\u003eWorkers and Executors 26\u003c\/p\u003e \u003cp\u003eConfiguration 27\u003c\/p\u003e \u003cp\u003eSpark Standalone 30\u003c\/p\u003e \u003cp\u003eArchitecture 31\u003c\/p\u003e \u003cp\u003eSingle-Node Setup Scenario 31\u003c\/p\u003e \u003cp\u003eMulti-Node Setup 32\u003c\/p\u003e \u003cp\u003eYARN 33\u003c\/p\u003e \u003cp\u003eArchitecture 35\u003c\/p\u003e \u003cp\u003eDynamic Resource Allocation 37\u003c\/p\u003e \u003cp\u003eScenario 39\u003c\/p\u003e \u003cp\u003eMesos 40\u003c\/p\u003e \u003cp\u003eSetup 41\u003c\/p\u003e \u003cp\u003eArchitecture 42\u003c\/p\u003e \u003cp\u003eDynamic Resource Allocation 44\u003c\/p\u003e \u003cp\u003eBasic Setup Scenario 44\u003c\/p\u003e \u003cp\u003eComparison 46\u003c\/p\u003e \u003cp\u003eSummary 50\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 3 Performance Tuning 53\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003eSpark Execution Model 54\u003c\/p\u003e \u003cp\u003ePartitioning 56\u003c\/p\u003e \u003cp\u003eControlling Parallelism 56\u003c\/p\u003e \u003cp\u003ePartitioners 58\u003c\/p\u003e \u003cp\u003eShuffling Data 59\u003c\/p\u003e \u003cp\u003eShuffling and Data Partitioning 61\u003c\/p\u003e \u003cp\u003eOperators and Shuffl ing 63\u003c\/p\u003e \u003cp\u003eShuffling Is Not That Bad After All 67\u003c\/p\u003e \u003cp\u003eSerialization 67\u003c\/p\u003e \u003cp\u003eKryo Registrators 69\u003c\/p\u003e \u003cp\u003eSpark Cache 69\u003c\/p\u003e \u003cp\u003eSpark SQL Cache 73\u003c\/p\u003e \u003cp\u003eMemory Management 73\u003c\/p\u003e \u003cp\u003eGarbage Collection 74\u003c\/p\u003e \u003cp\u003eShared Variables 75\u003c\/p\u003e \u003cp\u003eBroadcast Variables 76\u003c\/p\u003e \u003cp\u003eAccumulators 78\u003c\/p\u003e \u003cp\u003eData Locality 81\u003c\/p\u003e \u003cp\u003eSummary 82\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 4 Security 83\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003eArchitecture 84\u003c\/p\u003e \u003cp\u003eSecurity Manager 84\u003c\/p\u003e \u003cp\u003eSetup Configurations 85\u003c\/p\u003e \u003cp\u003eACL 86\u003c\/p\u003e \u003cp\u003eConfiguration 86\u003c\/p\u003e \u003cp\u003eJob Submission 87\u003c\/p\u003e \u003cp\u003eWeb UI 88\u003c\/p\u003e \u003cp\u003eNetwork Security 95\u003c\/p\u003e \u003cp\u003eEncryption 96\u003c\/p\u003e \u003cp\u003eEvent logging 101\u003c\/p\u003e \u003cp\u003eKerberos 101\u003c\/p\u003e \u003cp\u003eApache Sentry 102\u003c\/p\u003e \u003cp\u003eSummary 102\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 5 Fault Tolerance or Job Execution 105\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003eLifecycle of a Spark Job 106\u003c\/p\u003e \u003cp\u003eSpark Master 107\u003c\/p\u003e \u003cp\u003eSpark Driver 109\u003c\/p\u003e \u003cp\u003eSpark Worker 111\u003c\/p\u003e \u003cp\u003eJob Lifecycle 112\u003c\/p\u003e \u003cp\u003eJob Scheduling 112\u003c\/p\u003e \u003cp\u003eScheduling within an Application 113\u003c\/p\u003e \u003cp\u003eScheduling with External Utilities 120\u003c\/p\u003e \u003cp\u003eFault Tolerance 122\u003c\/p\u003e \u003cp\u003eInternal and External Fault Tolerance 122\u003c\/p\u003e \u003cp\u003eService Level Agreements (SLAs) 123\u003c\/p\u003e \u003cp\u003eResilient Distributed Datasets (RDDs) 124\u003c\/p\u003e \u003cp\u003eBatch versus Streaming 130\u003c\/p\u003e \u003cp\u003eTesting Strategies 133\u003c\/p\u003e \u003cp\u003eRecommended Confi gurations 139\u003c\/p\u003e \u003cp\u003eSummary 142\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 6 Beyond Spark 145\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003eData Warehousing 146\u003c\/p\u003e \u003cp\u003eSpark SQL CLI 147\u003c\/p\u003e \u003cp\u003eThrift JDBC\/ODBC Server 147\u003c\/p\u003e \u003cp\u003eHive on Spark 148\u003c\/p\u003e \u003cp\u003eMachine Learning 150\u003c\/p\u003e \u003cp\u003eDataFrame 150\u003c\/p\u003e \u003cp\u003eMLlib and ML 153\u003c\/p\u003e \u003cp\u003eMahout on Spark 158\u003c\/p\u003e \u003cp\u003eHivemall on Spark 160\u003c\/p\u003e \u003cp\u003eExternal Frameworks 161\u003c\/p\u003e \u003cp\u003eSpark Package 161\u003c\/p\u003e \u003cp\u003eXGBoost 163\u003c\/p\u003e \u003cp\u003espark-jobserver 164\u003c\/p\u003e \u003cp\u003eFuture Works 166\u003c\/p\u003e \u003cp\u003eIntegration with the Parameter Server 167\u003c\/p\u003e \u003cp\u003eDeep Learning 175\u003c\/p\u003e \u003cp\u003eEnterprise Usage 182\u003c\/p\u003e \u003cp\u003eCollecting User Activity Log with Spark and Kafka 183\u003c\/p\u003e \u003cp\u003eReal-Time Recommendation with Spark 184\u003c\/p\u003e \u003cp\u003eReal-Time Categorization of Twitter Bots 186\u003c\/p\u003e \u003cp\u003eSummary 186\u003c\/p\u003e \u003cp\u003eIndex 189\u003c\/p\u003e","brand":"John Wiley \u0026 Sons Inc","offers":[{"title":"Default Title","offer_id":49407020171607,"sku":"9781119254010","price":37.99,"currency_code":"GBP","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0817\/1739\/5799\/files\/9781119254010.jpg?v=1730497899","url":"https:\/\/bookcurl.com\/products\/spark-9781119254010","provider":"Book Curl","version":"1.0","type":"link"}