Data Science Toolkit and Applications
Last updatedAugust 21, 2026Presentation & objectives
This course covers the entire data science pipeline, from initial data exploration to the implementation and evaluation of machine learning models. It emphasizes hands-on experience with essential tools and libraries, preparing students to tackle real-world data science challenges. The course also incorporates best practices in software development and version control.
Key concepts covered:
- Python programming for data science (including key libraries like Pandas, NumPy, Scikit-learn)
- Data manipulation and storage (structured data formats, databases)
- Data exploration and preprocessing techniques
- Data visualization principles and tools
- Model training and evaluation
- Software development best practices (version control, testing, documentation)
Prerequisites :
- Being familiar and efficient with Python programming
- Being familiar with basic Linux commands
You can refer to the course presentation slides if needed: Course presentation.
Assessment
By the end of this course, you will be assessed on the core competencies defined by the program. The evaluation is based on a machine learning case study, whose deliverable is a public project deposited on GitLab that also serves as the foundation of your portfolio. See the Assessment Guidelines for how your final grade is determined, and Academic Integrity for the rules you must follow.
License
This content is made available under a Creative Commons license. See License & Citation for the full terms and how to cite this material.