Data Science syllabus
PCC-252-AID · Second Year Artificial Intelligence and Data Science, SPPU 2024 pattern. Every unit, the marks scheme, course outcomes and books, copied from the official syllabus PDF.
Unit-wise syllabus
Fundamentals of Data Science
6 hoursBasic Concepts of Data: Types of data: Structured, Semi-structured, Unstructured, Scales of measurement: Nominal, Ordinal, Interval, Ratio, Data formats: CSV, JSON, XML, SQL tables, Data quality dimensions: Accuracy, Completeness, Consistency, Timeliness, Difference Between Data, Information, Knowledge, wisdom. Definition of Data Science, The Data Science Process/Lifecycle (e.g., CRISP-DM, Team Data Science Process (TDSP), SEMMA Methodology), Applications of Data Science in various domains (e.g., Healthcare, Finance, Retail), Distinction between Data Science, Machine Learning, and Artificial Intelligence, Roles in Data Science (e.g., Data Analyst, Data Scientist, Machine Learning Engineer), Ethical and Privacy Issues in Data Science.
Case Study: A telecom business wants to proactively identify customers who are likely to quit because of the high rate of customer turnover. The organization wants to predict loss of customers using data science techniques and take proactive steps to keep at-risk clients. Determine the steps that, as data scientists, we must adhere to in order to resolve this issue.
Applications of Mathematical Statistics
6 hoursImportance of mathematical statistics in Data Science, Role of mathematical statistics in building data models, mathematical statistics foundations needed for machine learning algorithms. Linear Algebra for Data Science: Vectors and Matrices, Vector operations (dot product, cross product) and its application in data science, Matrix operations (addition, multiplication, inverse, determinant), Eigenvalues and Eigenvectors and its uses in dimensionality reduction, Probability Concepts: Random variables and probability distributions, Bayes’ Theorem and conditional probability, Expected value and variance. Statistical Methods: Measures of central tendency (Mean, Median, Mode), Measures of dispersion (Variance, Standard Deviation) and its importance in data preprocessing ,Correlation and covariance and its importance in feature selection in modeling.
Case Study : In a dataset, you have a feature matrix and output vector .Elaborate, How does matrix multiplication help in solving linear regression using the Normal Equation.
Programming for Data Science
6 hoursMachine Learning: Types (Supervised, unsupervised, Reinforcement), Supervised Algorithms (Linear Regression, Logistic Regression, Decision Trees, SVM). Metrics: Confusion Matrix, Accuracy, Precision-Recall, ROC-AUC, F1 Score, Mean Squared Error (MSE),R-squared. NumPy: Arrays, Mathematical Operations, Linear Algebra, Random Number Generation, Broadcasting. Pandas: Data Frames, Cleaning, Transformation, Filtering, Grouping, Import/Export, Missing Data Handling. Matplotlib: Basic & Advanced Plots, Customization, Subplots, Interactive Features. Seaborn: Distribution, Categorical, Relationship, Regression Plots, Multi-Plot Grids, Themes.
Case study: Write a case study on E-commerce Data. Key Components of the Analytical Plan are: Problem statement and objectives, Data Collection and Preparation, Data Analysis and Model Development, Analysis Outcomes and Takeaways
Data Preprocessing and Visualization
6 hoursData Preprocessing, Data Cleaning: Handling missing data, outliers. - Data Transformation Techniques: Feature Extraction and Selection (PCA), encoding. - Exploratory Data Analysis (EDA): Univariate, bivariate, multivariate analysis. Outliers: (Z-Score Method, Interquartile Range) Data Visualization: Overview of Data Visualization, Need of Data Visualization, Shapes of data, input for data visualization, Types of Data Visualization: Cognitive and perceptual, Practicing good ethics in Data Visualization, Principles of visual perception, Data Visualization Tool (Tableau, Power BI)
Case study : 1.
Case Study on Exploratory Data Analysis (EDA) and Visualizations 2. Create simple plot to visualize a distribution of variables using python
Automating AI Workflows with Pandas
6 hoursETL, ETL Challenges, ETL in Data Preprocessing for AutoML, AutoML: AutoML Workflow, AutoML Libraries: Auto-sklearn, H2O.ai, and TPOT, Benefits and Limitations of AutoML, AutoML vs Traditional Machine Learning, Pandas for data preprocessing for AutoML, data normalization and standardization for AutoML, Open-source Tools and Environments: Python, R, Jupyter, Git, VS Code and its roles in data science. Introduction to Big Data: Characteristics, Architecture and Ecosystem, Tools and Technologies, Applications and Challenges of Big Data in Data Science. AI-powered Data Cleaning: Using AI models to clean and structure raw data before analysis.
Case study: A retail company wants to predict its future sales revenue based on various factors such as product pricing, marketing spends, seasonal trends, and customer demographics. The company aims to use regression and classification models to analyze historical sales data, identify trends, and make data-driven decisions to enhance profitability and optimize resource allocation. Justify how does the choice between regression and classification models affect sales predictions?
Marks and credits
| Head | Marks | Credit |
|---|---|---|
| CCE (continuous comprehensive evaluation) | 30 | 2 |
| End-semester exam | 70 |
Prerequisite: Programming and Problem Solving, Artificial Intelligence.
Course outcomes
- CO1Discuss core concept of data science and its practical applications.
- CO2Apply mathematical tools like linear algebra, probability, and statistics to model datadriven problem solutions.
- CO3Analyze core machine learning algorithms and methodologies to address diverse problem sets.
- CO4Recommend effective data cleaning, transformation, and visualization techniques to extract meaningful insights from data.
- CO5Use automation tools for AI workflows to enhance the scalability and efficiency of AIdriven solutions.
Books
Text books
- Vijay Kotu, Bala Deshpande, “Data Science Concepts and Practice”, 2nd Edition, Morgan Kaufmann, ISBN 978-0-12-814761-0.
- Suresh Kumar Mukhiya, Usman Ahmed, “Hands-On Exploratory Data Analysis with Python” 1st Edition, 2020, Packt Publish, ISBN 978-1-80323-110-5.
- Dirk P. Kroese et.al.,“Data Science and Machine Learning: Mathematical and Statistical Methods”,1st Edition, CRC Press, ISBN 978-1-138-49253-0.
Reference books
- Davy Cielen, Arno D.B. Meysman, Mohamed Ali, ”Introducing Data Science: Big Data, Machine Learning, and More, Using Python Tools”, 1st Edition, Dreamtech Press, ISBN 978-1-63343-003-7.
- Arockia Liborious, Rik Das, “Fun with Machine Learning”, 1st Edition, BPB Publications, ISBN 978-93-555-1785-2
NPTEL and SWAYAM links
Listed in the official syllabus:
FAQ
How many units are in Data Science?
Data Science (PCC-252-AID) has 5 units and 30 hours of theory: Unit I Fundamentals of Data Science (6 h); Unit II Applications of Mathematical Statistics (6 h); Unit III Programming for Data Science (6 h); Unit IV Data Preprocessing and Visualization (6 h); Unit V Automating AI Workflows with Pandas (6 h).
What is the marks scheme for Data Science?
The official Artificial Intelligence and Data Science 2024 pattern syllabus lists continuous comprehensive evaluation (CCE) for 30 marks and the end-semester exam for 70 marks, for 2 credits.
What should I know before Data Science?
Prerequisite listed in the syllabus: Programming and Problem Solving, Artificial Intelligence.