課程目錄:Data Science for Big Data Analytics培訓
4401 人關注
(78637/99817)
課程大綱:

         Data Science for Big Data Analytics培訓

 

 

 

Introduction to Data Science for Big Data Analytics
Data Science Overview
Big Data Overview
Data Structures
Drivers and complexities of Big Data
Big Data ecosystem and a new approach to analytics
Key technologies in Big Data
Data Mining process and problems
Association Pattern Mining
Data Clustering
Outlier Detection
Data Classification
Introduction to Data Analytics lifecycle
Discovery
Data preparation
Model planning
Model building
Presentation/Communication of results
Operationalization
Exercise: Case study
From this point most of the training time (80%) will be spent on examples and exercises in R and related big data technology.
Getting started with R
Installing R and Rstudio
Features of R language
Objects in R
Data in R
Data manipulation
Big data issues
Exercises
Getting started with Hadoop
Installing Hadoop
Understanding Hadoop modes
HDFS
MapReduce architecture
Hadoop related projects overview
Writing programs in Hadoop MapReduce
Exercises
Integrating R and Hadoop with RHadoop
Components of RHadoop
Installing RHadoop and connecting with Hadoop
The architecture of RHadoop
Hadoop streaming with R
Data analytics problem solving with RHadoop
Exercises
Pre-processing and preparing data
Data preparation steps
Feature extraction
Data cleaning
Data integration and transformation
Data reduction – sampling, feature subset selection,
Dimensionality reduction
Discretization and binning
Exercises and Case study
Exploratory data analytic methods in R
Descriptive statistics
Exploratory data analysis
Visualization – preliminary steps
Visualizing single variable
Examining multiple variables
Statistical methods for evaluation
Hypothesis testing
Exercises and Case study
Data Visualizations
Basic visualizations in R
Packages for data visualization ggplot2, lattice, plotly, lattice
Formatting plots in R
Advanced graphs
Exercises
Regression (Estimating future values)
Linear regression
Use cases
Model description
Diagnostics
Problems with linear regression
Shrinkage methods, ridge regression, the lasso
Generalizations and nonlinearity
Regression splines
Local polynomial regression
Generalized additive models
Regression with RHadoop
Exercises and Case study
Classification
The classification related problems
Bayesian refresher
Na?ve Bayes
Logistic regression
K-nearest neighbors
Decision trees algorithm
Neural networks
Support vector machines
Diagnostics of classifiers
Comparison of classification methods
Scalable classification algorithms
Exercises and Case study
Assessing model performance and selection
Bias, Variance and model complexity
Accuracy vs Interpretability
Evaluating classifiers
Measures of model/algorithm performance
Hold-out method of validation
Cross-validation
Tuning machine learning algorithms with caret package
Visualizing model performance with Profit ROC and Lift curves
Ensemble Methods
Bagging
Random Forests
Boosting
Gradient boosting
Exercises and Case study
Support vector machines for classification and regression
Maximal Margin classifiers
Support vector classifiers
Support vector machines
SVM’s for classification problems
SVM’s for regression problems
Exercises and Case study
Identifying unknown groupings within a data set
Feature Selection for Clustering
Representative based algorithms: k-means, k-medoids
Hierarchical algorithms: agglomerative and divisive methods
Probabilistic base algorithms: EM
Density based algorithms: DBSCAN, DENCLUE
Cluster validation
Advanced clustering concepts
Clustering with RHadoop
Exercises and Case study
Discovering connections with Link Analysis
Link analysis concepts
Metrics for analyzing networks
The Pagerank algorithm
Hyperlink-Induced Topic Search
Link Prediction
Exercises and Case study
Association Pattern Mining
Frequent Pattern Mining Model
Scalability issues in frequent pattern mining
Brute Force algorithms
Apriori algorithm
The FP growth approach
Evaluation of Candidate Rules
Applications of Association Rules
Validation and Testing
Diagnostics
Association rules with R and Hadoop
Exercises and Case study
Constructing recommendation engines
Understanding recommender systems
Data mining techniques used in recommender systems
Recommender systems with recommenderlab package
Evaluating the recommender systems
Recommendations with RHadoop
Exercise: Building recommendation engine
Text analysis
Text analysis steps
Collecting raw text
Bag of words
Term Frequency –Inverse Document Frequency
Determining Sentiments
Exercises and Case study

主站蜘蛛池模板: 色综合久久久久综合99| 久久久久综合中文字幕| 亚洲另类激情综合偷自拍| 色综合久久综合网观看| 在线综合+亚洲+欧美中文字幕| 国产欧美精品一区二区色综合| 色爱区综合激情五月综合色| 色婷婷六月亚洲综合香蕉 | 久久久久久综合网天天| 亚洲国产综合欧美在线不卡| 色欲综合久久躁天天躁蜜桃| 自拍 偷拍 另类 综合图片| 色综合久久天天综合| 精品综合久久久久久888蜜芽| 综合久久久久久中文字幕亚洲国产国产综合一区首 | 少妇熟女久久综合网色欲| 色欲久久久天天天综合网精品| 五月综合激情婷婷六月色窝| 色综合久久夜色精品国产| 狠狠色丁香久久婷婷综合图片| 亚洲第一区欧美国产不卡综合| 综合久久久久久中文字幕亚洲国产国产综合一区首 | AV色综合久久天堂AV色综合在| 久久婷婷色香五月综合激情| 婷婷久久综合九色综合绿巨人 | 无码专区久久综合久中文字幕| 久久香综合精品久久伊人| 日韩欧美亚洲综合久久影院Ds | 国产成人综合日韩精品无码不卡| 色婷婷综合久久久久中文| 欧美亚洲综合另类成人| 国产综合精品一区二区三区| 人人狠狠综合久久亚洲| 伊人yinren6综合网色狠狠| 色五月丁香六月欧美综合图片 | 久久婷婷五月综合97色一本一本| 国产成人亚洲综合无码| 亚洲狠狠久久综合一区77777| 青青草原综合久久大伊人精品| 久久综合亚洲色HEZYO社区| 丁香五月综合缴情综合|