Project Machine Learning Systems WiSe 2026/27
(PR, 41341 Project Machine Learning Systems)

Project Machine Learning Systems (PMLS) is a 12 ECTS module offered to master's students in Computer Science, Computer Engineering, Information Systems Management, and Electrical Engineering. Large Language Models (LLMs) are increasingly used in scientific research and discovery, coding agents, robotics, conversational assistants, information retrieval, and many more areas. At the core of these applications are inference engines. These systems load trained model(s), process user input, efficiently execute the model and generate output.

PMLS provides a hands-on introduction to the architecture and implementation of transformer-based LLM inference systems. Students will learn how trained models are stored and loaded, how input text is tokenized and processed, how attention layers are executed, and how output tokens are generated by an autoregressive model.

Content

The practical has a total capacity of 20 students. PMLS is ungraded, however, the following will be used for evaluation:

  • Project implementation (source code) [45%]
  • Tests, benchmarks, experiments [30%]
  • Design document [10%]
  • Presentation of the final prototype (30min interview) [15%]


Practical: LLM Inference Engine

Project Code: (link coming soon)

Task Description: The goal of the PMLS project is to implement an LLM inference engine in C++. Starting from a small reference code base, students will progressively build an initial prototype containing the core components required to run pretrained transformer-based language models locally. The engine will load model weights and configuration files from existing Hugging Face models, process tokenized user input, execute the transformer architecture, and perform autoregressive token generation. The implementation will cover the essential building blocks of modern LLM inference, including positional encodings, tensor operations, model loading, multi-head attention, feed-forward blocks, layer normalization, a KV cache, and sampling strategies. After the prototype implementation, students will extend their inference engine with one or more advanced features, such as support for newer architectures, quantization, hand-optimized kernels, continuous batching, paged attention, prefix caching, speculative decoding, constrained decoding, sparse attention, support for mixture-of-experts models, watermarking, or other techniques used in modern LLM inference systems.

Lab Sessions:

The first two lab sessions will take place in FR 706, afterwards lab sessions will be in FR 713.
  • 14.10.2026, FR 706 at 16:00 s.t.: Kickoff and Introduction [pdf coming soon]


Organization

  • Lecturer: Univ.-Prof. Dr.-Ing. Matthias Boehm, DAMS
  • Teaching Assistant: Philipp Ortner, DAMS
  • Project submission 1: TBD
  • Project submission 2: TBD
  • Project interviews: TBD
  • Grading: pass with ≥ 50% of the points