Information Retrieval syllabus
PEC362ACOM · Third Year Computer Engineering, SPPU 2024 pattern. Every unit, the marks scheme, course outcomes and books, copied from the official syllabus PDF.
Unit-wise syllabus
Introduction to IR & Text Processing
9 hoursDerived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.
- Introduction to Information Retrieval – Definition, Need, IR vs Database Systems, Applications of IR, Components of an IR System, Architecture of a Search Engine, Document Representation, Types of Documents (Structured, Semi-structured, Unstructured), Text Preprocessing: Tokenization, Stop Word Removal, Stemming (Porter Algorithm), Lemmatization, Normalization Techniques, Term Weighting: Bag of Words, Term Frequency (TF), Inverse Document Frequency (IDF), TF-IDF, Vector Space Model (Basics), Cosine Similarity, Introduction to Inverted Index (Structure & Need) Case Study: Spam Email Filtering using Text Preprocessing
Preserved official unit paragraph
Introduction to Information Retrieval – Definition, Need, IR vs Database Systems, Applications of IR, Components of an IR System, Architecture of a Search Engine, Document Representation, Types of Documents (Structured, Semi-structured, Unstructured), Text Preprocessing: Tokenization, Stop Word Removal, Stemming (Porter Algorithm), Lemmatization, Normalization Techniques, Term Weighting: Bag of Words, Term Frequency (TF), Inverse Document Frequency (IDF), TF-IDF, Vector Space Model (Basics), Cosine Similarity, Introduction to Inverted Index (Structure & Need) Case Study: Spam Email Filtering using Text Preprocessing
Information Retrieval Models
9 hoursDerived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.
- Boolean Retrieval Model – Boolean Queries (AND, OR, NOT), Query Processing, Limitations of Boolean Model, Introduction to Ranked Retrieval, Vector Space Model (Detailed) – Document Vectors, Query Vectors, Similarity Measures, Document Ranking using Cosine Similarity, Probabilistic Retrieval Model – Basic Probability, Binary Independence Model, Relevance Feedback, BM25 (Basic Concept), Language Models for IR (Query Likelihood, Smoothing Basics) + Evaluation Metrics (Precision, Recall, F1, MAP Overview) Case Study: Vector Space Model in Academic Digital Libraries
Preserved official unit paragraph
Boolean Retrieval Model – Boolean Queries (AND, OR, NOT), Query Processing, Limitations of Boolean Model, Introduction to Ranked Retrieval, Vector Space Model (Detailed) – Document Vectors, Query Vectors, Similarity Measures, Document Ranking using Cosine Similarity, Probabilistic Retrieval Model – Basic Probability, Binary Independence Model, Relevance Feedback, BM25 (Basic Concept), Language Models for IR (Query Likelihood, Smoothing Basics) + Evaluation Metrics (Precision, Recall, F1, MAP Overview) Case Study: Vector Space Model in Academic Digital Libraries
Index Compression and Dynamic Inverted Indices
9 hoursDerived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.
- Performance evaluation: Precision and recall, alternative measures Ontology: Ontology based information sharing, Ontology languages for semantic web, Ontology creation.
- General-Purpose Data Compression, Modeling and Coding, Huffman Coding, Arithmetic Coding, Symbol wise ,Text Compression Compressing Postings Lists: Nonparametric Gap Compression, Parametric Gap Compression, Context- Aware Compression Methods, Index Compression for High Query Performance, Compression Effectiveness, Decoding Performance, Document Reordering.
- Dynamic Inverted Indices: Incremental Index Updates, Contiguous Inverted Lists, Noncontiguous Inverted, Document Deletions: Invalidation List, Garbage Collection, Document Modifications, Case Study: Ontology-Based Medical Information Retrieval (e.g., PubMed + MeSH Ontology)
Preserved official unit paragraph
Performance evaluation: Precision and recall, alternative measures Ontology: Ontology based information sharing, Ontology languages for semantic web, Ontology creation. General-Purpose Data Compression, Modeling and Coding, Huffman Coding, Arithmetic Coding, Symbol wise ,Text Compression Compressing Postings Lists: Nonparametric Gap Compression, Parametric Gap Compression, Context- Aware Compression Methods, Index Compression for High Query Performance, Compression Effectiveness, Decoding Performance, Document Reordering. Dynamic Inverted Indices: Incremental Index Updates, Contiguous Inverted Lists, Noncontiguous Inverted, Document Deletions: Invalidation List, Garbage Collection, Document Modifications, Case Study: Ontology-Based Medical Information Retrieval (e.g., PubMed + MeSH Ontology)
Web Searching
9 hoursDerived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.
- Introduction, Challenges, Web Characteristics, Search Engines: Centralized Architecture, Distributed Architecture, User Interfaces, Ranking, Crawling the web, Indices, Browsing, Meta-searchers, Searching using Hyperlinks, Trends and Research Issues, Introduction to Web Scraping: Python for web scraping, Request, HTML parsing, Beautiful Soup.
- Case Study: Web Crawling Architecture (Googlebot)
Preserved official unit paragraph
Introduction, Challenges, Web Characteristics, Search Engines: Centralized Architecture, Distributed Architecture, User Interfaces, Ranking, Crawling the web, Indices, Browsing, Meta-searchers, Searching using Hyperlinks, Trends and Research Issues, Introduction to Web Scraping: Python for web scraping, Request, HTML parsing, Beautiful Soup. Case Study: Web Crawling Architecture (Googlebot)
Advanced Information Retrieval
9 hoursDerived reading outline. Source text split at semicolons, line breaks and sentence boundaries, not an official topic hierarchy.
- XML Retrieval: Basic XML concepts, Challenges in XML retrieval, Vector space model for XML retrieval, Evaluation of XML retrieval, Text-Centric vs. Data-Centric XML retrieval.
- Recommendation System: Collaborative Filtering and Content Based Recommendation of Documents and Products.
- Introduction to Semantic Web.
- Case Study: Personalized Search using User Profiling Focus Areas
Preserved official unit paragraph
XML Retrieval: Basic XML concepts, Challenges in XML retrieval, Vector space model for XML retrieval, Evaluation of XML retrieval, Text-Centric vs. Data-Centric XML retrieval. Recommendation System: Collaborative Filtering and Content Based Recommendation of Documents and Products. Introduction to Semantic Web. Case Study: Personalized Search using User Profiling Focus Areas
Marks and credits
| Head | Marks | Credit |
|---|---|---|
| CCE (continuous comprehensive evaluation) | 30 | 3 |
| End-semester exam | 70 |
Course outcomes
- CO1To Understand IR concepts and apply clustering techniques in IR •
- CO2To Study indexing structures for Information Retrieval. •
- CO3To evaluate the performance of IR systems and understand user interfaces for searching •
- CO4Map IR concepts to recent developments in the field of Information Retrieval. •
- CO5Interpret recent trends and applications in Information Retrieval
Books
Text books
- Yates & Neto, "Modern Information Retrieval", Pearson Education, ISBN 81-297-0274-6.
- C.J. Rijsbergen, "Information Retrieval", (www.dcs.gla.ac.uk).
- Heiner Stuckenschmidt, Frank van Harmelen, “Information Sharing on th Semantic Web”, Springer International Edition, ISBN 3-540-20594-2.
Reference books
- 1. Data Mining Techniques, Arun K Pujari, 3rd Edition, UniversitiesPress
- Data Warehousing Fundament’s, Pualraj Ponnaiah, Wiley Student Edition.
- The Data Warehouse Life Cycle Toolkit — Ralph Kimball, Wiley StudentEdition.
- Data Mining, Vikaram Pudi, P Radha Krishna, Oxford University Press
FAQ
How many units are in Information Retrieval?
Information Retrieval (PEC362ACOM) has 5 units and 45 hours of theory: Unit I Introduction to IR & Text Processing (9 h); Unit II Information Retrieval Models (9 h); Unit III Index Compression and Dynamic Inverted Indices (9 h); Unit IV Web Searching (9 h); Unit V Advanced Information Retrieval (9 h).
What is the marks scheme for Information Retrieval?
The official Computer Engineering 2024 pattern syllabus lists continuous comprehensive evaluation (CCE) for 30 marks and the end-semester exam for 70 marks, for 3 credits.