Researcher at Inria
The objective of this course is to present the principles of the knowledge discovery and data mining pipeline, covering acquisition, storage, and extraction. From storing data to searching for similarities, the course will include a mix of lectures and practical sessions.
Each lecture has two 1.5 h parts (with a 15 min break in between), starts at 9 am and ends at 12.15 pm.
This exam lasts 2.5 h. The only documents allowed are 6 handwritten A4 sheets (both sides). Communicating devices are strictly forbidden. When writing code, imprecision in language syntax will be tolerated. The exam is graded out of 20 points. The questions will cover Lectures 1 (Sep. 18) to 7 (Oct. 30). There is only one bonus question on Lecture 8 (Nov. 06).
| Session | Type | |
|---|---|---|
| Dec. 01 | Written exam of 2.5h |
This course has two practice sessions covering both database systems (Lab 1) and Similarity-based algorithms (Lab 2). Each practice session is divided into Part 1, tackled in Class, and Part 2, to do as homework. The notebook, with code, results, and explanation, should be filled out (for both Part 1 and Part 2) and sent to me by email in both PDF and IPYNB formats (at paul [dot] boniol [at] inria [dot] fr) before the corresponding deadline (at 5 pm CEST (Central European Summer Time)). Each Practice session is graded out of 20 points.
| Session | Title | topic | Notebook | Data | Deadline |
|---|---|---|---|---|---|
| Oct. 09 | Database of documents | Oct. 22 | |||
| Oct. 23 | Similarity Search | Nov. 05 |
This course is highly inspired by previous courses of other professors and researchers listed below: