Data Science and Big Data Analytics syllabus
2019 PATTERN. This is the 2019 pattern syllabus, the latest SPPU has published on its site for Third Year Computer Engineering. A 2024 pattern syllabus for this year has not been published there yet, so confirm with your college which pattern applies to you.
310251 · Third Year Computer Engineering, SPPU 2019 pattern. Every unit, the marks scheme, course outcomes and books, copied from the official syllabus PDF.
Unit-wise syllabus
Introduction to Data Science and Big Data
7 hoursBasics and need of Data Science and Big Data, Applications of Data Science, Data explosion, 5 V’s of Big Data, Relationship between Data Science and Information Science, Business intelligence versus Data Science, Data Science Life Cycle, Data: Data Types, Data Collection. Need of Data wrangling, Methods: Data Cleaning, Data Integration, Data Reduction, Data Transformation, Data Discretization.
Statistical Inference
7 hoursNeed of statistics in Data Science and Big Data Analytics, Measures of Central Tendency: Mean, Median, Mode, Mid-range. Measures of Dispersion: Range, Variance, Mean Deviation, Standard Deviation. Bayes theorem, Basics and need of hypothesis and hypothesis testing, Pearson Correlation, Sample Hypothesis testing, Chi-Square Tests, t-test.
Big Data Analytics Life Cycle
7 hoursIntroduction to Big Data, sources of Big Data, Data Analytic Lifecycle: Introduction, Phase 1: Discovery, Phase 2: Data Preparation, Phase 3: Model Planning, Phase 4: Model Building, Phase 5: Communication results, Phase 6: Operation alize.
Predictive Big Data Analytics with Python
7 hoursIntroduction, Essential Python Libraries, Basic examples. Data Preprocessing: Removing Duplicates, Transformation of Data using function or mapping, replacing values, Handling Missing Data. Analytics Types: Predictive, Descriptive and Prescriptive. Association Rules: Apriori Algorithm, FP growth. Regression: Linear Regression, Logistic Regression. Classification: Naïve Bayes, Decision Trees. Introduction to Scikit-learn, Installations, Dataset, mat plotlib, filling missing values, Regression and Classification using Scikit-learn.
Big Data Analytics and Model Evaluation
7 hoursClustering Algorithms: K-Means, Hierarchical Clustering, Time-series analysis. Introduction to Text Analysis: Text-preprocessing, Bag of words, TF-IDF and topics. Need and Introduction to social network analysis, Introduction to business analysis. Model Evaluation and Selection: Metrics for Evaluating Classifier Performance, Holdout Method and Random Sub sampling, Parameter Tuning and Optimization, Result Interpretation, Clustering and Time-series analysis using Scikitlearn, sklearn. metrics, Confusion matrix, AUC-ROC Curves, Elbow plot.
Data Visualization and Hadoop
7 hoursIntroduction to Data Visualization, Challenges to Big data visualization, Types of data visualization, Data Visualization Techniques, Visualizing Big Data, Tools used in Data Visualization, Hadoop ecosystem, Map Reduce, Pig, Hive, Analytical techniques used in Big data visualization. Data Visualization using Python: Line plot, Scatter plot, Histogram, Density plot, Boxplot.
Marks and credits
| Head | Marks | Credit |
|---|---|---|
| Mid-Sem (mid-semester exam) | 30 | 3 |
| End-semester exam | 70 |
Prerequisite: Discrete Mathematics (210241), Database Management Systems (310341).
Course outcomes
- CO1Analyze needs and challenges for Data Science Big Data Analytics
- CO2Apply statistics for Big Data Analytics
- CO3Apply the lifecycle of Big Data analytics to real world problems
- CO4Implement Big Data Analytics using Python programming
- CO5Implement data visualization using visualization tools in Python programming
- CO6Design and implement Big Databases using the Hadoop ecosystem
Books
Text books
- David Dietrich, Barry Hiller, “Data Science and Big Data Analytics”, EMC education services, Wiley publication, 2012, ISBN0-07-120413-X
- Jiawei Han, Micheline Kamber, and Jian Pie, “Data Mining: Concepts and Techniques” Elsevier Publishers Third Edition, ISBN: 9780123814791, 9780123814807
Reference books
- EMC Education Services, “Data Science and Big Data Analytics- Discovering, analyzing Visualizing and Presenting Data”
- DT Editorial Services, “Big Data, Black Book”, DT Editorial Services, ISBN: 9789351197577, 2016 Edition
- Chirag Shah, “A Hands-On Introduction To Data Science”, Cambridge University Press, (2020), ISBN : ISBN 978-1-108-47244-9
- Wes McKinney, “Python for Data Analysis ”, O' Reilly media, ISBN: 978-1-449-31979-3
- Trent Hauk, “Scikit-learn Cookbook”, Packt Publishing, ISBN: 9781787286382
- Jenny Kim, Benjamin Bengfort, “Data Analytics with Hadoop”, OReilly Media, Inc., ISBN: 9781491913703
- Venkat Ankam,“Big Data Analytics”, Packt Publishing, ISBN: 9781785884696 Home
- Seema Acharya, Subhashini Chellappan, “Big Data And Analytics”, Wiley publication, ISBN: 9788126579518
FAQ
How many units are in Data Science and Big Data Analytics?
Data Science and Big Data Analytics (310251) has 6 units: Unit I Introduction to Data Science and Big Data (7 h); Unit II Statistical Inference (7 h); Unit III Big Data Analytics Life Cycle (7 h); Unit IV Predictive Big Data Analytics with Python (7 h); Unit V Big Data Analytics and Model Evaluation (7 h); Unit VI Data Visualization and Hadoop (7 h).
What is the marks scheme for Data Science and Big Data Analytics?
The official Computer Engineering 2019 pattern syllabus lists mid-semester (Mid-Sem) for 30 marks and the end-semester exam for 70 marks, for 3 credits.
What should I know before Data Science and Big Data Analytics?
Prerequisite listed in the syllabus: Discrete Mathematics (210241), Database Management Systems (310341).