Data Science 2 (DS2) is an elective in the computer science bachelor's programme at HHU Düsseldorf: algorithms for data that does not fit on one machine, and for data that arrives as a stream. It follows the Data Science module and is built on two books — Leskovec, Rajaraman and Ullman's Mining of Massive Datasets for the algorithms, and Martin Kleppmann's Designing Data-Intensive Applications for the systems they run on.
The course is taught entirely in German: lectures, slides, exercise sheets, sample solutions and recordings. Only the two textbooks are in English. This page and the module description below are translations; nothing else has an English version.
Konrad Völkel teaches DS2 in winter 2026/27. Students can see rooms and dates for lectures, exercises and exams in LSF. Once registered in LSF, students get access to the ILIAS page with current course materials.
What the course covers
- Big data, tail latency, MapReduce
- Communication cost, partitioning, column stores
- Similarity: shingling, MinHash, locality sensitive hashing
- Streams: sampling, sliding windows, stream systems, online algorithms
- Bloom filters, Flajolet-Martin, HyperLogLog, Count-Min
- PageRank in graphs, and PageRank in practice
- Clustering, and BFR for data that does not fit in memory
- Recommender systems and latent factors
- SVD for large matrices

Recorded lectures, freely accessible
Twelve lectures from the winter 2024/25 run are public on HHU's media server — in German, no login needed. They are the only part of the course that is openly accessible: slides, exercise sheets, sample solutions and the mock exam live in the ILIAS course behind an HHU login. New recordings for the run of winter 2026/27 are planned but not promised.
- VL 1: Map-Reduce
- VL 2: Jaccard-Abstand
- VL 3: MinHash — recording lost
- VL 4: LSH
- VL 5: Streams
- VL 6: Bloom Filter und Flajolet-Martin
- VL 7: PageRank part 1
- VL 8: untitled — the session on advertising on the web
- VL 9: Clustering
- VL 10: BFR-Clustering, Motivation Empfehlungssysteme
- VL 11: Empfehlungssysteme (Inhaltsbasiert, kollaboratives Filtern)
- VL 12: Empfehlungssysteme: latente-Faktoren-Modelle
- VL 13: Wiederholung & Vertiefung probabilistische Datenstrukturen
Material
The two books the course is taught from, both freely readable:
- Mining of Massive Datasets — Leskovec, Rajaraman and Ullman, 3rd edition 2020. The algorithmic spine of the course; the authors publish it for free. Chapter 11 carries the SVD session.
- Martin Kleppmann, Designing Data-Intensive Applications (O'Reilly, 2017). Not background reading: chapters 1, 3, 6 and 11 are taught — tail latency, column stores, partitioning, and stream systems.
My own two scripts, both in German, both with sources:
- Data Science script (PDF) — the preceding module. This is where the probability theory, the estimators and the Numpy come from.
- Einführung in Python script (PDF) — three chapters of it are used directly here: Big Data & High-Performance Pipelines, Von Generatoren zu Async/Await and Profiling und Performance-Analyse.
There is no script for Data Science 2.
Using this material
If you teach something in this area and want the slides, the exercise sheets or the sample solutions, write to me and I will send them. The upcoming slides for this term are made with LaTeX Beamer, and I am happy to hand over the sources rather than the PDFs. They are in German.
The official module description
This is what the module is formally examined against, and what an examination office needs if the course is to count for something elsewhere. It is the entry in the Modulhandbuch B.Sc. Informatik PO 2021 (version of 22 July 2026, page 54), freely translated from the German original:
Contents. Data science is the application of statistical and machine-learning methods to data of any kind, by computer, in order to model systems and predict behaviour. Data Science 2 takes the foundations laid in Data Science and deepens them towards application, above all in two areas: large volumes of data (big data), and data-parallel processing — the algorithms and the software packages.
Learning outcomes. Having taken part in this module successfully, students can analyse algorithms for large volumes of data and for data streams with respect to running time and communication complexity — in particular for nearest-neighbour search in high-dimensional data, locality sensitive hashing (LSH), dimensionality reduction, recommender systems, clustering (variants of k-means), link analysis (PageRank) and advertising on the web. They can judge for which kind and volume of data algorithms without parallelisation become unsuitable, and which techniques of parallel processing suit a given problem — map-reduce above all —, and they can either implement those or put existing software packages to appropriate use.
The module is worth 5 ECTS — lecture 2 SWS, exercise class 2 SWS — and is assessed by a written exam, with admission earned through the exercise sheets. Formally there are no prerequisites; in substance the handbook expects the contents of Data Science or of Machine Learning — rudimentary Python, the basics of machine learning, and vectorised programming with Numpy.
When it runs again
The handbook schedules the module for every second winter term. The last run is winter 2026/27; after it, the next one is expected in winter 2028/29, and nothing later is fixed. Earlier runs, in the university's course registry (in German): winter 2024/25, winter 2023/24, winter 2022/23.
Data Science 2 is one of the courses I teach at HHU Düsseldorf; the others are on the teaching page.
