{"product_id":"a-general-introduction-to-data-analytics-9781119296249","title":"A General Introduction to Data Analytics","description":"\u003cb\u003eBook Synopsis\u003c\/b\u003e\u003cbr\u003eA guide to the principles and methods of data analysis that does not require knowledge of statistics or programming A General Introduction to Data Analytics is an essential guide to understand and use data analytics. This book is written using easy-to-understand terms and does not require familiarity with statistics or programming. The authorsnoted experts in the fieldhighlight an explanation of the intuition behind the basic data analytics techniques. The text also contains exercises and illustrative examples.    Thought to be easily accessible to non-experts, the book provides motivation to the necessity of analyzing data. It explains how to visualize and summarize data, and how to find natural groups and frequent patterns in a dataset. The book also explores predictive tasks, be them classification or regression. Finally, the book discusses popular data analytic applications, like mining the web, information retrieval, social network analysis, working with text, and recommender syst\u003cbr\u003e\u003cbr\u003e\u003cb\u003eTable of Contents\u003c\/b\u003e\u003cbr\u003e\u003cp\u003ePreface xiii\u003c\/p\u003e \u003cp\u003eAcknowledgments xv\u003c\/p\u003e \u003cp\u003ePresentational Conventions xvii\u003c\/p\u003e \u003cp\u003eAbout the Companion Website xix\u003c\/p\u003e \u003cp\u003e\u003cb\u003ePart I Introductory Background 1\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e\u003cb\u003e1 What Can We Do With Data? 3\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e1.1 Big Data and Data Science 4\u003c\/p\u003e \u003cp\u003e1.2 Big Data Architectures 5\u003c\/p\u003e \u003cp\u003e1.3 Small Data 6\u003c\/p\u003e \u003cp\u003e1.4 What is Data? 7\u003c\/p\u003e \u003cp\u003e1.5 A Short Taxonomy of Data Analytics 9\u003c\/p\u003e \u003cp\u003e1.6 Examples of Data Use 10\u003c\/p\u003e \u003cp\u003e1.6.1 Breast Cancer in Wisconsin 11\u003c\/p\u003e \u003cp\u003e1.6.2 Polish Company Insolvency Data 11\u003c\/p\u003e \u003cp\u003e1.7 A Project on Data Analytics 12\u003c\/p\u003e \u003cp\u003e1.7.1 A Little History on Methodologies for Data Analytics 12\u003c\/p\u003e \u003cp\u003e1.7.2 The KDD Process 14\u003c\/p\u003e \u003cp\u003e1.7.3 The CRISP-DM Methodology 15\u003c\/p\u003e \u003cp\u003e1.8 How this Book is Organized 16\u003c\/p\u003e \u003cp\u003e1.9 Who Should Read this Book 18\u003c\/p\u003e \u003cp\u003e\u003cb\u003ePart II Getting Insights from Data 19\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e\u003cb\u003e2 Descriptive Statistics 21\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e2.1 Scale Types 22\u003c\/p\u003e \u003cp\u003e2.2 Descriptive Univariate Analysis 25\u003c\/p\u003e \u003cp\u003e2.2.1 Univariate Frequencies 25\u003c\/p\u003e \u003cp\u003e2.2.2 Univariate Data Visualization 27\u003c\/p\u003e \u003cp\u003e2.2.3 Univariate Statistics 32\u003c\/p\u003e \u003cp\u003e2.2.4 Common Univariate Probability Distributions 38\u003c\/p\u003e \u003cp\u003e2.3 Descriptive Bivariate Analysis 40\u003c\/p\u003e \u003cp\u003e2.3.1 Two Quantitative Attributes 41\u003c\/p\u003e \u003cp\u003e2.3.2 Two Qualitative Attributes, at Least one of them Nominal 45\u003c\/p\u003e \u003cp\u003e2.3.3 Two Ordinal Attributes 46\u003c\/p\u003e \u003cp\u003e2.4 Final Remarks 47\u003c\/p\u003e \u003cp\u003e2.5 Exercises 47\u003c\/p\u003e \u003cp\u003e\u003cb\u003e3 Descriptive Multivariate Analysis 49\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e3.1 Multivariate Frequencies 49\u003c\/p\u003e \u003cp\u003e3.2 Multivariate Data Visualization 50\u003c\/p\u003e \u003cp\u003e3.3 Multivariate Statistics 59\u003c\/p\u003e \u003cp\u003e3.3.1 Location Multivariate Statistics 59\u003c\/p\u003e \u003cp\u003e3.3.2 Dispersion Multivariate Statistics 60\u003c\/p\u003e \u003cp\u003e3.4 Infographics and Word Clouds 66\u003c\/p\u003e \u003cp\u003e3.4.1 Infographics 66\u003c\/p\u003e \u003cp\u003e3.4.2 Word Clouds 67\u003c\/p\u003e \u003cp\u003e3.5 Final Remarks 67\u003c\/p\u003e \u003cp\u003e3.6 Exercises 68\u003c\/p\u003e \u003cp\u003e\u003cb\u003e4 Data Quality and Preprocessing 71\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e4.1 Data Quality 71\u003c\/p\u003e \u003cp\u003e4.1.1 Missing Values 72\u003c\/p\u003e \u003cp\u003e4.1.2 Redundant Data 74\u003c\/p\u003e \u003cp\u003e4.1.3 Inconsistent Data 75\u003c\/p\u003e \u003cp\u003e4.1.4 Noisy Data 76\u003c\/p\u003e \u003cp\u003e4.1.5 Outliers 77\u003c\/p\u003e \u003cp\u003e4.2 Converting to a Diﬀerent Scale Type 77\u003c\/p\u003e \u003cp\u003e4.2.1 Converting Nominal to Relative 78\u003c\/p\u003e \u003cp\u003e4.2.2 Converting Ordinal to Relative or Absolute 81\u003c\/p\u003e \u003cp\u003e4.2.3 Converting Relative or Absolute to Ordinal or Nominal 82\u003c\/p\u003e \u003cp\u003e4.3 Converting to a Diﬀerent Scale 83\u003c\/p\u003e \u003cp\u003e4.4 Data Transformation 85\u003c\/p\u003e \u003cp\u003e4.5 Dimensionality Reduction 86\u003c\/p\u003e \u003cp\u003e4.5.1 Attribute Aggregation 88\u003c\/p\u003e \u003cp\u003e4.5.1.1 Principal Component Analysis 88\u003c\/p\u003e \u003cp\u003e4.5.1.2 Independent Component Analysis 91\u003c\/p\u003e \u003cp\u003e4.5.1.3 Multidimensional Scaling 91\u003c\/p\u003e \u003cp\u003e4.5.2 Attribute Selection 92\u003c\/p\u003e \u003cp\u003e4.5.2.1 Filters 92\u003c\/p\u003e \u003cp\u003e4.5.2.2 Wrappers 93\u003c\/p\u003e \u003cp\u003e4.5.2.3 Embedded 94\u003c\/p\u003e \u003cp\u003e4.5.2.4 Search Strategies 95\u003c\/p\u003e \u003cp\u003e4.6 Final Remarks 96\u003c\/p\u003e \u003cp\u003e4.7 Exercises 96\u003c\/p\u003e \u003cp\u003e\u003cb\u003e5 Clustering 99\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e5.1 Distance Measures 100\u003c\/p\u003e \u003cp\u003e5.1.1 Diﬀerences between Values of Common Attribute Types 101\u003c\/p\u003e \u003cp\u003e5.1.2 Distance Measures for Objects with Quantitative Attributes 103\u003c\/p\u003e \u003cp\u003e5.1.3 Distance Measures for Non-conventional Attributes 104\u003c\/p\u003e \u003cp\u003e5.2 Clustering Validation 107\u003c\/p\u003e \u003cp\u003e5.3 Clustering Techniques 108\u003c\/p\u003e \u003cp\u003e5.3.1 K-means 110\u003c\/p\u003e \u003cp\u003e5.3.1.1 Centroids and Distance Measures 110\u003c\/p\u003e \u003cp\u003e5.3.1.2 How K-means Works 111\u003c\/p\u003e \u003cp\u003e5.3.2 DBSCAN 115\u003c\/p\u003e \u003cp\u003e5.3.3 Agglomerative Hierarchical Clustering Technique 117\u003c\/p\u003e \u003cp\u003e5.3.3.1 Linkage Criterion 119\u003c\/p\u003e \u003cp\u003e5.3.3.2 Dendrograms 120\u003c\/p\u003e \u003cp\u003e5.4 Final Remarks 122\u003c\/p\u003e \u003cp\u003e5.5 Exercises 123\u003c\/p\u003e \u003cp\u003e\u003cb\u003e6 Frequent Pattern Mining 125\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e6.1 Frequent Itemsets 127\u003c\/p\u003e \u003cp\u003e6.1.1 Setting the \u003ci\u003emin_sup\u003c\/i\u003e Threshold 128\u003c\/p\u003e \u003cp\u003e6.1.2 Apriori – a Join-based Method 131\u003c\/p\u003e \u003cp\u003e6.1.3 Eclat 133\u003c\/p\u003e \u003cp\u003e6.1.4 FP-Growth 134\u003c\/p\u003e \u003cp\u003e6.1.5 Maximal and Closed Frequent Itemsets 138\u003c\/p\u003e \u003cp\u003e6.2 Association Rules 139\u003c\/p\u003e \u003cp\u003e6.3 Behind Support and Conﬁdence 142\u003c\/p\u003e \u003cp\u003e6.3.1 Cross-support Patterns 143\u003c\/p\u003e \u003cp\u003e6.3.2 Lift 144\u003c\/p\u003e \u003cp\u003e6.3.3 Simpson’s Paradox 145\u003c\/p\u003e \u003cp\u003e6.4 Other Types of Pattern 147\u003c\/p\u003e \u003cp\u003e6.4.1 Sequential patterns 147\u003c\/p\u003e \u003cp\u003e6.4.2 Frequent Sequence Mining 148\u003c\/p\u003e \u003cp\u003e6.4.3 Closed and Maximal Sequences 148\u003c\/p\u003e \u003cp\u003e6.5 Final Remarks 149\u003c\/p\u003e \u003cp\u003e6.6 Exercises 149\u003c\/p\u003e \u003cp\u003e\u003cb\u003e7 Cheat Sheet and Project on Descriptive Analytics 151\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e7.1 Cheat Sheet of Descriptive Analytics 151\u003c\/p\u003e \u003cp\u003e7.1.1 On Data Summarization 151\u003c\/p\u003e \u003cp\u003e7.1.2 On Clustering 151\u003c\/p\u003e \u003cp\u003e7.1.3 On Frequent Pattern Mining 153\u003c\/p\u003e \u003cp\u003e7.2 Project on Descriptive Analytics 154\u003c\/p\u003e \u003cp\u003e7.2.1 Business Understanding 154\u003c\/p\u003e \u003cp\u003e7.2.2 Data Understanding 155\u003c\/p\u003e \u003cp\u003e7.2.3 Data Preparation 155\u003c\/p\u003e \u003cp\u003e7.2.4 Modeling 157\u003c\/p\u003e \u003cp\u003e7.2.5 Evaluation 158\u003c\/p\u003e \u003cp\u003e7.2.6 Deployment 158\u003c\/p\u003e \u003cp\u003e\u003cb\u003ePart III Predicting the Unknown 159\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e\u003cb\u003e8 Regression 161\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e8.1 Predictive Performance Estimation 164\u003c\/p\u003e \u003cp\u003e8.1.1 Generalization 164\u003c\/p\u003e \u003cp\u003e8.1.2 Model Validation 165\u003c\/p\u003e \u003cp\u003e8.1.3 Predictive Performance Measures for Regression 169\u003c\/p\u003e \u003cp\u003e8.2 Finding the Parameters of the Model 171\u003c\/p\u003e \u003cp\u003e8.2.1 Linear Regression 171\u003c\/p\u003e \u003cp\u003e8.2.1.1 Empirical Error 173\u003c\/p\u003e \u003cp\u003e8.2.2 The Bias-variance Trade-oﬀ 175\u003c\/p\u003e \u003cp\u003e8.2.3 Shrinkage Methods 177\u003c\/p\u003e \u003cp\u003e8.2.3.1 Ridge Regression 179\u003c\/p\u003e \u003cp\u003e8.2.3.2 Lasso Regression 180\u003c\/p\u003e \u003cp\u003e8.2.4 Methods that use Linear Combinations of Attributes 181\u003c\/p\u003e \u003cp\u003e8.2.4.1 Principal Components Regression 181\u003c\/p\u003e \u003cp\u003e8.2.4.2 Partial Least Squares Regression 182\u003c\/p\u003e \u003cp\u003e8.3 Technique and Model Selection 182\u003c\/p\u003e \u003cp\u003e8.4 Final Remarks 183\u003c\/p\u003e \u003cp\u003e8.5 Exercises 184\u003c\/p\u003e \u003cp\u003e\u003cb\u003e9 Classiﬁcation 187\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e9.1 Binary Classiﬁcation 188\u003c\/p\u003e \u003cp\u003e9.2 Predictive Performance Measures for Classiﬁcation 192\u003c\/p\u003e \u003cp\u003e9.3 Distance-based Learning Algorithms 199\u003c\/p\u003e \u003cp\u003e9.3.1 K-nearest Neighbor Algorithms 199\u003c\/p\u003e \u003cp\u003e9.3.2 Case-based Reasoning 202\u003c\/p\u003e \u003cp\u003e9.4 Probabilistic Classiﬁcation Algorithms 203\u003c\/p\u003e \u003cp\u003e9.4.1 Logistic Regression Algorithm 205\u003c\/p\u003e \u003cp\u003e9.4.2 Naive Bayes Algorithm 207\u003c\/p\u003e \u003cp\u003e9.5 Final Remarks 208\u003c\/p\u003e \u003cp\u003e9.6 Exercises 210\u003c\/p\u003e \u003cp\u003e\u003cb\u003e10 Additional Predictive Methods 211\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e10.1 Search-based Algorithms 211\u003c\/p\u003e \u003cp\u003e10.1.1 Decision Tree Induction Algorithms 212\u003c\/p\u003e \u003cp\u003e10.1.2 Decision Trees for Regression 217\u003c\/p\u003e \u003cp\u003e10.1.2.1 Model Trees 218\u003c\/p\u003e \u003cp\u003e10.1.2.2 Multivariate Adaptive Regression Splines 219\u003c\/p\u003e \u003cp\u003e10.2 Optimization-based Algorithms 221\u003c\/p\u003e \u003cp\u003e10.2.1 Artiﬁcial Neural Networks 222\u003c\/p\u003e \u003cp\u003e10.2.1.1 Backpropagation 224\u003c\/p\u003e \u003cp\u003e10.2.1.2 Deep Networks and Deep Learning Algorithms 230\u003c\/p\u003e \u003cp\u003e10.2.2 Support Vector Machines 233\u003c\/p\u003e \u003cp\u003e10.2.2.1 SVM for Regression 237\u003c\/p\u003e \u003cp\u003e10.3 Final Remarks 238\u003c\/p\u003e \u003cp\u003e10.4 Exercises 239\u003c\/p\u003e \u003cp\u003e\u003cb\u003e11 Advanced Predictive Topics 241\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e11.1 Ensemble Learning 241\u003c\/p\u003e \u003cp\u003e11.1.1 Bagging 243\u003c\/p\u003e \u003cp\u003e11.1.2 Random Forests 244\u003c\/p\u003e \u003cp\u003e11.1.3 AdaBoost 245\u003c\/p\u003e \u003cp\u003e11.2 Algorithm Bias 246\u003c\/p\u003e \u003cp\u003e11.3 Non-binary Classiﬁcation Tasks 248\u003c\/p\u003e \u003cp\u003e11.3.1 One-class Classiﬁcation 248\u003c\/p\u003e \u003cp\u003e11.3.2 Multi-class Classiﬁcation 249\u003c\/p\u003e \u003cp\u003e11.3.3 Ranking Classiﬁcation 250\u003c\/p\u003e \u003cp\u003e11.3.4 Multi-label Classiﬁcation 251\u003c\/p\u003e \u003cp\u003e11.3.5 Hierarchical Classiﬁcation 252\u003c\/p\u003e \u003cp\u003e11.4 Advanced Data Preparation Techniques for Prediction 253\u003c\/p\u003e \u003cp\u003e11.4.1 Imbalanced Data Classiﬁcation 253\u003c\/p\u003e \u003cp\u003e11.4.2 For Incomplete Target Labeling 254\u003c\/p\u003e \u003cp\u003e11.4.2.1 Semi-supervised Learning 254\u003c\/p\u003e \u003cp\u003e11.4.2.2 Active Learning 255\u003c\/p\u003e \u003cp\u003e11.5 Description and Prediction with Supervised Interpretable Techniques 255\u003c\/p\u003e \u003cp\u003e11.6 Exercises 256\u003c\/p\u003e \u003cp\u003e\u003cb\u003e12 Cheat Sheet and Project on Predictive Analytics 259\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e12.1 Cheat Sheet on Predictive Analytics 259\u003c\/p\u003e \u003cp\u003e12.2 Project on Predictive Analytics 259\u003c\/p\u003e \u003cp\u003e12.2.1 Business Understanding 260\u003c\/p\u003e \u003cp\u003e12.2.2 Data Understanding 260\u003c\/p\u003e \u003cp\u003e12.2.3 Data Preparation 265\u003c\/p\u003e \u003cp\u003e12.2.4 Modeling 265\u003c\/p\u003e \u003cp\u003e12.2.5 Evaluation 265\u003c\/p\u003e \u003cp\u003e12.2.6 Deployment 266\u003c\/p\u003e \u003cp\u003e\u003cb\u003ePart IV Popular Data Analytics Applications 267\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e13 Applications for Text, Web and Social Media 269\u003c\/p\u003e \u003cp\u003e13.1 Working with Texts 269\u003c\/p\u003e \u003cp\u003e13.1.1 Data Acquisition 271\u003c\/p\u003e \u003cp\u003e13.1.2 Feature Extraction 271\u003c\/p\u003e \u003cp\u003e13.1.2.1 Tokenization 272\u003c\/p\u003e \u003cp\u003e13.1.2.2 Stemming 272\u003c\/p\u003e \u003cp\u003e13.1.2.3 Conversion to Structured Data 275\u003c\/p\u003e \u003cp\u003e13.1.2.4 Is the Bag of Words Enough? 276\u003c\/p\u003e \u003cp\u003e13.1.3 Remaining Phases 277\u003c\/p\u003e \u003cp\u003e13.1.4 Trends 277\u003c\/p\u003e \u003cp\u003e13.1.4.1 Sentiment Analysis 278\u003c\/p\u003e \u003cp\u003e13.1.4.2 Web Mining 278\u003c\/p\u003e \u003cp\u003e13.2 Recommender Systems 278\u003c\/p\u003e \u003cp\u003e13.2.1 Feedback 279\u003c\/p\u003e \u003cp\u003e13.2.2 Recommendation Tasks 280\u003c\/p\u003e \u003cp\u003e13.2.3 Recommendation Techniques 281\u003c\/p\u003e \u003cp\u003e13.2.3.1 Knowledge-based Techniques 281\u003c\/p\u003e \u003cp\u003e13.2.3.2 Content-based Techniques 282\u003c\/p\u003e \u003cp\u003e13.2.3.3 Collaborative Filtering Techniques 282\u003c\/p\u003e \u003cp\u003e13.2.4 Final Remarks 289\u003c\/p\u003e \u003cp\u003e13.3 Social Network Analysis 291\u003c\/p\u003e \u003cp\u003e13.3.1 Representing Social Networks 291\u003c\/p\u003e \u003cp\u003e13.3.2 Basic Properties of Nodes 294\u003c\/p\u003e \u003cp\u003e13.3.2.1 Degree 294\u003c\/p\u003e \u003cp\u003e13.3.2.2 Distance 294\u003c\/p\u003e \u003cp\u003e13.3.2.3 Closeness 295\u003c\/p\u003e \u003cp\u003e13.3.2.4 Betweenness 296\u003c\/p\u003e \u003cp\u003e13.3.2.5 Clustering Coeﬃcient 297\u003c\/p\u003e \u003cp\u003e13.3.3 Basic and Structural Properties of Networks 297\u003c\/p\u003e \u003cp\u003e13.3.3.1 Diameter 297\u003c\/p\u003e \u003cp\u003e13.3.3.2 Centralization 297\u003c\/p\u003e \u003cp\u003e13.3.3.3 Cliques 299\u003c\/p\u003e \u003cp\u003e13.3.3.4 Clustering Coeﬃcient 299\u003c\/p\u003e \u003cp\u003e13.3.3.5 Modularity 299\u003c\/p\u003e \u003cp\u003e13.3.4 Trends and Final Remarks 299\u003c\/p\u003e \u003cp\u003e13.4 Exercises 300\u003c\/p\u003e \u003cp\u003eApendix A: Comprehensive Description of the CRISP-DM Methodology 303\u003c\/p\u003e \u003cp\u003eReferences 311\u003c\/p\u003e \u003cp\u003eIndex 315\u003c\/p\u003e","brand":"John Wiley \u0026 Sons Inc","offers":[{"title":"Default Title","offer_id":49407030690135,"sku":"9781119296249","price":71.96,"currency_code":"GBP","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0817\/1739\/5799\/files\/9781119296249.jpg?v=1730497934","url":"https:\/\/bookcurl.com\/products\/a-general-introduction-to-data-analytics-9781119296249","provider":"Book Curl","version":"1.0","type":"link"}