Spark 2.x Cookbook 高清原版 pdf

所需积分/C币:50 2018-05-08 16:54:12 14.55MB PDF
收藏 收藏
举报

spark 2.0;spark;大数据;分布式计算框架;高清原版pdf
Apache Spark 2.X Cookbook Copyright C 2017 Packt Publishing All rights reserved. no part of this book may be reproduced stored in a retrieval system or transmitted in any form or by any means, without the prior written permission of the publisher, except in the case of brief quotations embedded in critical articles or reviews Every effort has been made in the preparation of this book to ensure the accuracy of the information presented. However the information contained in this book is sold without warranty either express or implied Neither the author nor packt Publishing, and its dealers and distributors will be held liable for any damages caused or alleged to be caused directly or indirectly by this book Packt Publishing has endeavored to provide trademark information about all of the companies and products mentioned in this book by the appropriate use of capitals However packt Publishing cannot guarantee the accuracy of this information First published: May 2017 Production reference: 1300517 Published by packt Publishing ltd Plac 35 Livery street Birmingham B3 2PB. UK ISBN978-1-78712-726-5 www.packtpub.com Credits Author Copy Editor Rishi yadav Gladson monteiro Reviewer Project Coordinator Prashant Verma Nidhi Joshi Commissioning editor Proofreader Amey varangaonkar Safis Editing Acquisition Editor Indexer Vinay argekar Pratik shirodkar Content Development Editor Graphics Jagruti Babara Tania dutta Technical editor Production coordinator Dinesh pawar Sharddha Falebhai About the author Rishi Yadav has 19 years of experience in designing and developing enterprise applications. He is an open source software expert and advises American companies on big data and public cloud trends. Rishi was honored as one of Silicon Valley's 40 under 40 in 2014. He earned his bachelor's degree from the prestigious Indian Institute of Technology, Delhi, in 1998 About 12 years ago, Rishi started InfoObjects, a company that helps data-driven businesses gain new insights into data InfoObjects combines the power of open source and big data to solve business challenges for its clients and has a special focus on Apache Spark. The company has been on the inc 5000 list of the fastest growing companies for 6 years in a row. InfoObjects has also been named the best place to work in the Bay Area in 2014 and 2015. Rishi is an open source contributor and active blogger This book is dedicated to my parents, Ganesh and Bhagwati Yadav; I would not be where I am without their unconditional support, trust, and providing me the freedom to choose a path of my own Special thanks go to my life partner, Anjali, for providing immense support and putting up with my long, arduous hours(yet again) Our 9-year-old son, Vedant, and niece, Kashmira, were the unrelenting force behind keeping me and the book on track Big thanks to InfoObjects CTO and my business partner, Sudhir Jangir, for providing valuable feedback and also contributing with recipes on enterprise security, a topic he is passionate about; to our Svp, Bart Hickenlooper, for taking the charge in leading the company to the next level to Tanmoy Chowdhury and Neeraj gupta for their valuable advice; to Yogesh Chandani, Animesh Chauhan, and Katie Nelson for running operations skillfully so that I could focus on this book; and to our internal review team(especially Rakesh Chandran) for ironing out the kinks. I would also like to thank marcel izumi for,as always, providing creative visuals. I cannot miss thanking our dog Sparky, for giving me company on my long nights out. Last but not least, special thanks to our valuable clients, partners, and employees, who have made info objects the best place to work at and, needless to say, an immensely successful organization about the reviewer Prashant Verma started his it career in 2011 as a Java developer at ericsson working in the telecom domain. After a couple of years of Java EE experience, he moved into the big data domain and has worked on almost all the popular big data technologies, such as Hadoop Spark, Flume, Mongo, and Cassandra. He has also played with Scala. Currently, he works with Qa Infotech as a lead data engineer working on solving e-learning problems using analytics and machine learning Prashant has also been working as a freelance consultant in his spare time I want to thank packt publishing for giving me the chance to review the book as well as my employer and my family for their patience while i was busy working on this book www.paCktpub.com Forsupportfilesanddownloadsrelatedtoyourbookpleasevisitwww.packtpub.Com Did you know that Packt offers eBook versions of every book published, with PDF and epubfilesavailableYoucanupgradetotheeboOkversionatwww.packtpub.comandasa print book customer, you are entitled to a discount on the eBook copy. Get in touch with us at service@packtpub com for more details Atwww.packtpub.comyoucanalsoreadacollectionoffreetechnicalarticlessignupfora range of free newsletters and receive exclusive discounts and offers on packt books and eBookS AMap https://www.packtpub.com/mapt Get the most in-demand software skills with Mapt. mapt gives you full access to all Packt books and video courses, as well as industry-leading tools to help you plan your personal development and advance your career Why subscribe? Fully searchable across every book published by Packt Copy and paste, print, and bookmark content On demand and accessible via a web browser Customer Feedback Thanks for purchasing this Packt book. At Packt, quality is at the heart of our editorial process. To help us improve, please leave us an honest review on this book's amazon page athttps://www.amazoncom/dp/1787127265 If you d like to join our team of regular reviewers you can e-mail us at customerreviews @packtpub com We award our regular reviewers with free e Books and videos in exchange for their valuable feedback. Help us be relentless in improving our roducts! Table of contents Preface Chapter 1: Getting Started with Apache Spark 6 Introduction 6 Leveraging Databricks Cloud 8 How to do it How it works 14 Cluster 15 Notebook 15 Table 15 Libra 15 Deploying Spark using Amazon EMR 15 What it represents is much bigger than what it looks 15 EMR's architecture 16 How to do it 16 How it works 23 EC2 instance types 24 T2-Free Tier Burstable(EBS only 25 M4-General purpose(EBS only 25 C4- Compute optimized 26 X1- Memory optimized 26 R4-Memory optimized 26 P2-General purpose GPU 13- Storage optimized 27 D2 -Storage optimized Installing Spark from binaries 27 Getting ready 28 How to do it 28 Building the spark source code with Maven 30 Getting ready 30 How to do it Launching spark on Amazon ec2 3 Getting ready 33 How to do it 34 See also 38 Deploying spark on a cluster in standalone mode 38 Getting ready 39 How to do it 39 How it works 41 See also 43 Deploying Spark on a cluster with Mesos 43 How to do it 43 Deploying Spark on a cluster with YARN 45 Getting read 45 How to do it 45 How it works 47 Understanding SparkContext and SparkSession 49 Spark Context 49 SparkSession 49 Understanding resilient distributed dataset-RDD 49 How to do it 50 Chapter 2: Developing Applications with Spark 54 Introduction 54 Exploring the Spark shell 55 How to do it 56 There' s more Developing a Spark applications in Eclipse with Maven 58 Getting ready 59 How to do it 59 Developing a spark applications in Eclipse with SBt 62 How to do it 62 Developing a Spark application in IntelliJ IDEA with Maven 64 How to do it 64 Developing a Spark application in IntelliJ IDEA with SBT 66 How to do it 66 Developing applications using the Zeppelin notebook 66 How to do it 66 Setting up Kerberos to do authentication 69 How to do it 70 There' s more Enabling Kerberos authentication for Spark 72 How to do it 72 There' s more 74 Securing data at rest 74 Securing data in transit 75 Chapter 3: Spark SQL 76 [i]

...展开详情
试读 127P Spark 2.x Cookbook 高清原版 pdf
立即下载 低至0.43元/次 身份认证VIP会员低至7折
一个资源只可评论一次,评论内容不能少于5个字
healthsunny 电子书很好 有目录
2019-01-08
回复
_a_0_ 电子书很好 有目录 方便阅读
2018-11-12
回复
aspirefhaha 超赞的资源,cookbook,随时可以看的资料
2018-07-11
回复
关注 私信 TA的资源
上传资源赚积分,得勋章
最新推荐
Spark 2.x Cookbook 高清原版 pdf 50积分/C币 立即下载
1/127
Spark 2.x Cookbook 高清原版 pdf第1页
Spark 2.x Cookbook 高清原版 pdf第2页
Spark 2.x Cookbook 高清原版 pdf第3页
Spark 2.x Cookbook 高清原版 pdf第4页
Spark 2.x Cookbook 高清原版 pdf第5页
Spark 2.x Cookbook 高清原版 pdf第6页
Spark 2.x Cookbook 高清原版 pdf第7页
Spark 2.x Cookbook 高清原版 pdf第8页
Spark 2.x Cookbook 高清原版 pdf第9页
Spark 2.x Cookbook 高清原版 pdf第10页
Spark 2.x Cookbook 高清原版 pdf第11页
Spark 2.x Cookbook 高清原版 pdf第12页
Spark 2.x Cookbook 高清原版 pdf第13页
Spark 2.x Cookbook 高清原版 pdf第14页
Spark 2.x Cookbook 高清原版 pdf第15页
Spark 2.x Cookbook 高清原版 pdf第16页
Spark 2.x Cookbook 高清原版 pdf第17页
Spark 2.x Cookbook 高清原版 pdf第18页
Spark 2.x Cookbook 高清原版 pdf第19页
Spark 2.x Cookbook 高清原版 pdf第20页

试读结束, 可继续阅读

50积分/C币 立即下载 >