Data Science and Big Data Analytics syllabus

2019 PATTERN. This is the 2019 pattern syllabus, the latest SPPU has published on its site for Third Year Information Technology. A 2024 pattern syllabus for this year has not been published there yet, so confirm with your college which pattern applies to you.

314452 · Third Year Information Technology, SPPU 2019 pattern. Every unit, the marks scheme, course outcomes and books, copied from the official syllabus PDF.

3144523 h/week theoryMid-Sem 30 + End-sem 70
06.units
03.credits

Unit-wise syllabus

UNIT I

Introduction: Data Science and Big Data

6 hours

Introduction to Data science and Big Data, Defining Data science and Big Data, Big Data examples, Data Explosion: Data Volume, Data Variety, Data Velocity and Veracity. Big data infrastructure and challenges Big Data Processing Architectures: Data Warehouse, Re-Engineering the Data Warehouse, shared everything and shared nothing architecture, Big data learning approaches. Data Science – The Big Picture: Relation between AI, Statistical Learning, Machine Learning, Data Mining and Big Data Analytics TE (Information Technology) Syllabus (2019 Course) 67 Curriculum for Third Year of Information Technology (2019 Course), Savitribai Phule Pune University

UNIT II

Mathematical Foundation of Big Data

7 hours

Probability: Random Variables and Joint Probability, Conditional Probability and concept of Markov chains, Tail bounds, Markov chains and random walks, Pair-wise independence and universal hashing Approximate counting, Approximate median. Data Streaming Models and Statistical Methods: Flajole Martin algorithm, Distance Sampling and Random Projections, Bloom filters, Mode, Variance, standard deviation, Correlation analysis and Analysis of Variance.

UNIT III

Big Data Processing

6 hours

Big Data Analytics- Ecosystem and Technologies, Introduction to Google file system, Hadoop Architecture, Hadoop Storage: HDFS, Common Hadoop Shell commands, Anatomy of File Write and Read, NameNode, Secondary NameNode, and DataNode, Hadoop MapReduce paradigm, Map Reduce tasks, Job, Task trackers - Cluster Setup – SSH & Hadoop Configuration, Introduction to NOSQL, Textual ETL processing.

UNIT IV

Big Data Analytics

6 hours

Big Data Analytics- Architecture and Life Cycle, Types of analysis, Analytical approaches, Data Analytics with Mathematical manipulations, Data Ingestion from different sources (CSV, JSON, html, Excel, mongoDB, mysql, sqlite), Data cleaning, Handling missing values, data imputation, Data transformation, Data Standardization, handling categorical data with 2 and more categories, statistical and graphical analysis methods, Hive Data Analytics.

UNIT V

Big Data Visualization

6 hours

Introduction to Data visualization, Challenges to Big data visualization, Conventional data visualization tools, Techniques for visual data representations, Types of data visualization, Visualizing Big Data, Tools used in data visualization, Propriety Data Visualization tools, Open – source data visualization tools,

Case Study: Analysis of a business problem of Zomato using visualization, Analytical techniques used in Big data visualization, Data Visualization using Tableau Introduction to: Candela, D3.js, Google Chart API TE (Information Technology) Syllabus (2019 Course) 68 Curriculum for Third Year of Information Technology (2019 Course), Savitribai Phule Pune University

UNIT VI

Big Data Technologies Application and Impact

5 hours

Social media analytics, Text mining, Mobile analytics, Data analytics life cycle of case studies, Organizational impact, understanding decision theory, creating big data strategy, big data value creation drivers, Michael Porter’s valuation creation models, Big data user experience ramifications, Identifying big data use cases, Big Data Analytics Challenges and Research directions.

Marks and credits

HeadMarksCredit
Mid-Sem (mid-semester exam)303
End-semester exam70

Prerequisite: 1. Engineering and discrete mathematics. 2. Database Management Systems, Data warehousing and Data mining. 3. Programming skill..

Course outcomes

  1. CO1Understand Big Data primitives.
  2. CO2Learn and apply different mathematical models for Big Data.
  3. CO3Demonstrate Big Data learning skills by developing industry or research applications.
  4. CO4Analyze and apply each learning model comes from a different algorithmic approach and it will perform differently under different datasets.
  5. CO5Understand, apply and analyze needs, challenges and techniques for big data visualization.
  6. CO6Learn different programming platforms for big data analytics.

Books

Text books

Reference books

FAQ

How many units are in Data Science and Big Data Analytics?

Data Science and Big Data Analytics (314452) has 6 units: Unit I Introduction: Data Science and Big Data (6 h); Unit II Mathematical Foundation of Big Data (7 h); Unit III Big Data Processing (6 h); Unit IV Big Data Analytics (6 h); Unit V Big Data Visualization (6 h); Unit VI Big Data Technologies Application and Impact (5 h).

What is the marks scheme for Data Science and Big Data Analytics?

The official Information Technology 2019 pattern syllabus lists mid-semester (Mid-Sem) for 30 marks and the end-semester exam for 70 marks, for 3 credits.

What should I know before Data Science and Big Data Analytics?

Prerequisite listed in the syllabus: 1. Engineering and discrete mathematics. 2. Database Management Systems, Data warehousing and Data mining. 3. Programming skill..