Data Visualization and Analytics syllabus
PEC362BCOM · Third Year Computer Engineering, SPPU 2024 pattern. Every unit, the marks scheme, course outcomes and books, copied from the official syllabus PDF.
Unit-wise syllabus
Introduction to Data Analytics and Life Cycle
9 hoursDerived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.
- Introduction to Data Science and Big Data, Applications of Data Science, 5 V’s of Big Data, Relationship between Data Science and Information Science, Business intelligence versus Data Science, Data: Data Types, Data Collection.
- Need of Data wrangling, Methods: Data Cleaning, Data Integration, Data Reduction, Data Transformation, Data Discretization.
- Data Analytic Life cycle: Introduction, Phase 1: Discovery, Phase 2: Data Preparation, Phase 3: Model Planning, Phase 4: Model Building, Phase 5: Communication results, Phase 6: Operationalize.
- Case study: Global Innovation Social Network and Analysis (GINA).
Preserved official unit paragraph
Introduction to Data Science and Big Data, Applications of Data Science, 5 V’s of Big Data, Relationship between Data Science and Information Science, Business intelligence versus Data Science, Data: Data Types, Data Collection. Need of Data wrangling, Methods: Data Cleaning, Data Integration, Data Reduction, Data Transformation, Data Discretization. Data Analytic Life cycle: Introduction, Phase 1: Discovery, Phase 2: Data Preparation, Phase 3: Model Planning, Phase 4: Model Building, Phase 5: Communication results, Phase 6: Operationalize. Case study: Global Innovation Social Network and Analysis (GINA).
Inferential Statistics
9 hoursDerived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.
- Statistical concepts, Measures of Central Tendency: Mean, Median, Mode, Mid-range.
- Measures of Dispersion: Range, Variance, Mean Deviation, Standard Deviation.
- Bayes theorem, Basics and need of hypothesis and hypothesis testing, Pearson Correlation, Sample Hypothesis testing, Chi-Square Tests, t-test.
- Case Study: For an employee dataset, create measure of central tendency and its measure of dispersion for statistical analysis of given data.
Preserved official unit paragraph
Statistical concepts, Measures of Central Tendency: Mean, Median, Mode, Mid-range. Measures of Dispersion: Range, Variance, Mean Deviation, Standard Deviation. Bayes theorem, Basics and need of hypothesis and hypothesis testing, Pearson Correlation, Sample Hypothesis testing, Chi-Square Tests, t-test. Case Study: For an employee dataset, create measure of central tendency and its measure of dispersion for statistical analysis of given data.
Predictive Analytics
9 hoursDerived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.
- Essential Python Libraries, Basic examples.
- Data Preprocessing: Removing Duplicates, Transformation of Data using function or mapping, replacing values, Handling Missing Data.
- Data Analytics Types: Predictive, Descriptive and Prescriptive.
- Association Rules: Apriori Algorithm Regression: Linear Regression, Logistic Regression.
- Classification: Decision Trees.
- Introduction to Scikit-learn: Installations, Dataset, Regression and Classification using Scikit-learn.
- Clustering Algorithms: K-Means.
- Case Study: Use Spam mail dataset and apply data preprocessing methods.
Preserved official unit paragraph
Essential Python Libraries, Basic examples. Data Preprocessing: Removing Duplicates, Transformation of Data using function or mapping, replacing values, Handling Missing Data. Data Analytics Types: Predictive, Descriptive and Prescriptive. Association Rules: Apriori Algorithm Regression: Linear Regression, Logistic Regression. Classification: Decision Trees. Introduction to Scikit-learn: Installations, Dataset, Regression and Classification using Scikit-learn. Clustering Algorithms: K-Means. Case Study: Use Spam mail dataset and apply data preprocessing methods.
Text Analytics and Model Evaluation
9 hoursDerived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.
- History of text mining, Roots of text mining overview of seven practices of text analytic, Application.
- Introduction to Text Analysis: Text-preprocessing, Bag of words, TF-IDF and topics.
- Need and Introduction to social network analysis.
- Naïve Bayes classifier for text classification, Model Evaluation and Selection: Metrics for Evaluating Classifier Performance, Holdout Method and Random Sub sampling, Parameter Tuning and Optimization, Confusion matrix, AUC-ROC Curves.
- Case Study: Use Spam mail dataset and apply Naïve Bayes classifier and evaluate model.
Preserved official unit paragraph
History of text mining, Roots of text mining overview of seven practices of text analytic, Application. Introduction to Text Analysis: Text-preprocessing, Bag of words, TF-IDF and topics. Need and Introduction to social network analysis. Naïve Bayes classifier for text classification, Model Evaluation and Selection: Metrics for Evaluating Classifier Performance, Holdout Method and Random Sub sampling, Parameter Tuning and Optimization, Confusion matrix, AUC-ROC Curves. Case Study: Use Spam mail dataset and apply Naïve Bayes classifier and evaluate model.
Data Visualization and Big Data Processing
9 hoursDerived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.
- Introduction to Data visualization, Challenges to Big data visualization, Data Visualization using Python: Line plot, Scatter plot, Histogram, Density plot, Box-plot, Pie Chart, Area Plot, Heatmap, Funnel Chart.
- Open source data visualization tools, Data Visualization using Tableau.
- Hadoop Ecosystem and Technologies, Hadoop Architecture, Hadoop Storage: HDFS, Common Hadoop Shell commands, Anatomy of File Write and Read, NameNode, Secondary NameNode, and DataNode, Hadoop MapReduce paradigm, Map Reduce tasks, Job, Task trackers, Apache SPARK Framework, Difference between Hadoop and Apache Spark.
- Case Study: Sales Data Visualization using Tableau.
Preserved official unit paragraph
Introduction to Data visualization, Challenges to Big data visualization, Data Visualization using Python: Line plot, Scatter plot, Histogram, Density plot, Box-plot, Pie Chart, Area Plot, Heatmap, Funnel Chart. Open source data visualization tools, Data Visualization using Tableau. Hadoop Ecosystem and Technologies, Hadoop Architecture, Hadoop Storage: HDFS, Common Hadoop Shell commands, Anatomy of File Write and Read, NameNode, Secondary NameNode, and DataNode, Hadoop MapReduce paradigm, Map Reduce tasks, Job, Task trackers, Apache SPARK Framework, Difference between Hadoop and Apache Spark. Case Study: Sales Data Visualization using Tableau.
Marks and credits
| Head | Marks | Credit |
|---|---|---|
| CCE (continuous comprehensive evaluation) | 30 | 3 |
| End-semester exam | 70 |
Course outcomes
- CO1Understand fundamental concepts of Data Science, Big Data, and the Data Analytics Life cycle •
- CO2Apply inferential statistical techniques to analyze datasets •
- CO3Implement predictive analytics models using Python programming •
- CO4Analyze text data using text mining techniques and evaluate machine learning models using standard metrics •
- CO5Implement data visualization using visualization tools in Python programming •
- CO6Design and implement Big Databases using the Hadoop ecosystem
Books
Text books
- David Dietrich, Barry Hiller, “Data Science and Big Data Analytics”, EMC education services, Wiley publication, 2012, ISBN0-07-120413-X
- Jiawei Han, Micheline Kamber, and Jian Pie, “Data Mining: Concepts and Techniques Elsevier Publishers Third Edition, ISBN: 9780123814791, 9780123814807
Reference books
- 1. EMC Education Services, “Data Science and Big Data Analytics- Discovering, analyzing Visualizing and Presenting Data”
- 2. DT Editorial Services, “Big Data, Black Book”, DT Editorial Services, ISBN: 9789351197577, 2016 Edition
- Chirag Shah, “A Hands-On Introduction To Data Science”, Cambridge University Press, (2020), ISBN : ISBN 978-1-108-47244-9
- Wes McKinney, “Python for Data Analysis ”, O’ Reilly media, ISBN: 978-1-449-31979-3
- Trent Hauk, “Scikit-learn Cookbook”, Packt Publishing, ISBN: 9781787286382
- Jenny Kim, Benjamin Bengfort, “Data Analytics with Hadoop”, OReilly Media, Inc., ISBN:9781491913703
- Venkat Ankam,“Big Data Analytics”, Packt Publishing, ISBN: 9781785884696
- Seema Acharya, Subhashini Chellappan, “Big Data And Analytics”, Wiley publication, ISBN: 9788126579518
FAQ
How many units are in Data Visualization and Analytics?
Data Visualization and Analytics (PEC362BCOM) has 5 units and 45 hours of theory: Unit I Introduction to Data Analytics and Life Cycle (9 h); Unit II Inferential Statistics (9 h); Unit III Predictive Analytics (9 h); Unit IV Text Analytics and Model Evaluation (9 h); Unit V Data Visualization and Big Data Processing (9 h).
What is the marks scheme for Data Visualization and Analytics?
The official Computer Engineering 2024 pattern syllabus lists continuous comprehensive evaluation (CCE) for 30 marks and the end-semester exam for 70 marks, for 3 credits.