Data Visualization and Analytics syllabus

PEC362BCOM · Third Year Computer Engineering, SPPU 2024 pattern. Every unit, the marks scheme, course outcomes and books, copied from the official syllabus PDF.

PEC362BCOM3 h/week theoryCCE 30 + End-sem 70
45hours of theory
05.units
03.credits

Unit-wise syllabus

UNIT I

Introduction to Data Analytics and Life Cycle

9 hours

Derived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.

  • Introduction to Data Science and Big Data, Applications of Data Science, 5 V’s of Big Data, Relationship between Data Science and Information Science, Business intelligence versus Data Science, Data: Data Types, Data Collection.
  • Need of Data wrangling, Methods: Data Cleaning, Data Integration, Data Reduction, Data Transformation, Data Discretization.
  • Data Analytic Life cycle: Introduction, Phase 1: Discovery, Phase 2: Data Preparation, Phase 3: Model Planning, Phase 4: Model Building, Phase 5: Communication results, Phase 6: Operationalize.
  • Case study: Global Innovation Social Network and Analysis (GINA).
Preserved official unit paragraph

Introduction to Data Science and Big Data, Applications of Data Science, 5 V’s of Big Data, Relationship between Data Science and Information Science, Business intelligence versus Data Science, Data: Data Types, Data Collection. Need of Data wrangling, Methods: Data Cleaning, Data Integration, Data Reduction, Data Transformation, Data Discretization. Data Analytic Life cycle: Introduction, Phase 1: Discovery, Phase 2: Data Preparation, Phase 3: Model Planning, Phase 4: Model Building, Phase 5: Communication results, Phase 6: Operationalize. Case study: Global Innovation Social Network and Analysis (GINA).

Unit permalink
UNIT II

Inferential Statistics

9 hours

Derived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.

  • Statistical concepts, Measures of Central Tendency: Mean, Median, Mode, Mid-range.
  • Measures of Dispersion: Range, Variance, Mean Deviation, Standard Deviation.
  • Bayes theorem, Basics and need of hypothesis and hypothesis testing, Pearson Correlation, Sample Hypothesis testing, Chi-Square Tests, t-test.
  • Case Study: For an employee dataset, create measure of central tendency and its measure of dispersion for statistical analysis of given data.
Preserved official unit paragraph

Statistical concepts, Measures of Central Tendency: Mean, Median, Mode, Mid-range. Measures of Dispersion: Range, Variance, Mean Deviation, Standard Deviation. Bayes theorem, Basics and need of hypothesis and hypothesis testing, Pearson Correlation, Sample Hypothesis testing, Chi-Square Tests, t-test. Case Study: For an employee dataset, create measure of central tendency and its measure of dispersion for statistical analysis of given data.

Unit permalink
UNIT III

Predictive Analytics

9 hours

Derived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.

  • Essential Python Libraries, Basic examples.
  • Data Preprocessing: Removing Duplicates, Transformation of Data using function or mapping, replacing values, Handling Missing Data.
  • Data Analytics Types: Predictive, Descriptive and Prescriptive.
  • Association Rules: Apriori Algorithm Regression: Linear Regression, Logistic Regression.
  • Classification: Decision Trees.
  • Introduction to Scikit-learn: Installations, Dataset, Regression and Classification using Scikit-learn.
  • Clustering Algorithms: K-Means.
  • Case Study: Use Spam mail dataset and apply data preprocessing methods.
Preserved official unit paragraph

Essential Python Libraries, Basic examples. Data Preprocessing: Removing Duplicates, Transformation of Data using function or mapping, replacing values, Handling Missing Data. Data Analytics Types: Predictive, Descriptive and Prescriptive. Association Rules: Apriori Algorithm Regression: Linear Regression, Logistic Regression. Classification: Decision Trees. Introduction to Scikit-learn: Installations, Dataset, Regression and Classification using Scikit-learn. Clustering Algorithms: K-Means. Case Study: Use Spam mail dataset and apply data preprocessing methods.

Unit permalink
UNIT IV

Text Analytics and Model Evaluation

9 hours

Derived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.

  • History of text mining, Roots of text mining overview of seven practices of text analytic, Application.
  • Introduction to Text Analysis: Text-preprocessing, Bag of words, TF-IDF and topics.
  • Need and Introduction to social network analysis.
  • Naïve Bayes classifier for text classification, Model Evaluation and Selection: Metrics for Evaluating Classifier Performance, Holdout Method and Random Sub sampling, Parameter Tuning and Optimization, Confusion matrix, AUC-ROC Curves.
  • Case Study: Use Spam mail dataset and apply Naïve Bayes classifier and evaluate model.
Preserved official unit paragraph

History of text mining, Roots of text mining overview of seven practices of text analytic, Application. Introduction to Text Analysis: Text-preprocessing, Bag of words, TF-IDF and topics. Need and Introduction to social network analysis. Naïve Bayes classifier for text classification, Model Evaluation and Selection: Metrics for Evaluating Classifier Performance, Holdout Method and Random Sub sampling, Parameter Tuning and Optimization, Confusion matrix, AUC-ROC Curves. Case Study: Use Spam mail dataset and apply Naïve Bayes classifier and evaluate model.

Unit permalink
UNIT V

Data Visualization and Big Data Processing

9 hours

Derived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.

  • Introduction to Data visualization, Challenges to Big data visualization, Data Visualization using Python: Line plot, Scatter plot, Histogram, Density plot, Box-plot, Pie Chart, Area Plot, Heatmap, Funnel Chart.
  • Open source data visualization tools, Data Visualization using Tableau.
  • Hadoop Ecosystem and Technologies, Hadoop Architecture, Hadoop Storage: HDFS, Common Hadoop Shell commands, Anatomy of File Write and Read, NameNode, Secondary NameNode, and DataNode, Hadoop MapReduce paradigm, Map Reduce tasks, Job, Task trackers, Apache SPARK Framework, Difference between Hadoop and Apache Spark.
  • Case Study: Sales Data Visualization using Tableau.
Preserved official unit paragraph

Introduction to Data visualization, Challenges to Big data visualization, Data Visualization using Python: Line plot, Scatter plot, Histogram, Density plot, Box-plot, Pie Chart, Area Plot, Heatmap, Funnel Chart. Open source data visualization tools, Data Visualization using Tableau. Hadoop Ecosystem and Technologies, Hadoop Architecture, Hadoop Storage: HDFS, Common Hadoop Shell commands, Anatomy of File Write and Read, NameNode, Secondary NameNode, and DataNode, Hadoop MapReduce paradigm, Map Reduce tasks, Job, Task trackers, Apache SPARK Framework, Difference between Hadoop and Apache Spark. Case Study: Sales Data Visualization using Tableau.

Unit permalink

Marks and credits

HeadMarksCredit
CCE (continuous comprehensive evaluation)303
End-semester exam70

Course outcomes

  1. CO1Understand fundamental concepts of Data Science, Big Data, and the Data Analytics Life cycle •
  2. CO2Apply inferential statistical techniques to analyze datasets •
  3. CO3Implement predictive analytics models using Python programming •
  4. CO4Analyze text data using text mining techniques and evaluate machine learning models using standard metrics •
  5. CO5Implement data visualization using visualization tools in Python programming •
  6. CO6Design and implement Big Databases using the Hadoop ecosystem

Books

Text books

Reference books

FAQ

How many units are in Data Visualization and Analytics?

Data Visualization and Analytics (PEC362BCOM) has 5 units and 45 hours of theory: Unit I Introduction to Data Analytics and Life Cycle (9 h); Unit II Inferential Statistics (9 h); Unit III Predictive Analytics (9 h); Unit IV Text Analytics and Model Evaluation (9 h); Unit V Data Visualization and Big Data Processing (9 h).

What is the marks scheme for Data Visualization and Analytics?

The official Computer Engineering 2024 pattern syllabus lists continuous comprehensive evaluation (CCE) for 30 marks and the end-semester exam for 70 marks, for 3 credits.

Add to my plan

Browse all syllabus courses