EC883 Big Data Analytics
Course Name:
EC883 Big Data Analytics
Programme:
Category:
Credits (L-T-P):
Content:
Introduction to Big Data Processing: Introduction to Big Data Analytics. What is Big Data? What are the challenges? Introduction to Apache Hadoop and MapReduce. Apache Spark, Spark programming. (Python and Spark), Spark - Resilient Distributed Dataset (RDDs), Need of big data frameworks, Large-Scale Data Processing with PySpark: Spark - RDDs, DataFrames, Spark SQL, PySpark + NumPy + SciPy, Code Optimization, Cluster Configurations, Linear Algebra Computation in Large Scale, Distributed File Storage Systems, Data Modeling and Optimization Problems: Introduction to modeling: numerical vs. probabilistic vs. Bayesian, Introduction to Optimization Problems, Batch and stochastic Gradient Descent, Newton’s Method, and Expectation-Maximization. Large-Scale Supervised Learning: Introduction to Supervised learning, Generalized Linear Models and Logistic Regression, Regularization, Random Forest, Support Vector Machine (SVM), XGBoost, AdaBoost, Deep Learning Models, Outlier Detection, Spark ML library, Large-Scale Unsupervised Learning: Introduction to Unsupervised learning, K-means / K-medoids, Gaussian Mixture Models, Dimensionality Reduction, Spark MLlib for Unsupervised Learning, Deep Learning Generative models. Applications: Large Scale Text Mining- Latent Semantic Indexing, Topic models, Latent Dirichlet Allocation, Spark ML library for NLP and Time Series Big data analysis