Fault Tolerance in Distributed Stream Processing Engines

IndoSys 2026 Tutorial

Speakers

  • Abhilash Jindal, IIT Delhi
  • Satyam Jay, IIT Delhi

List of Topics

  • Basics of stream processing
  • The problem of faults
  • Distributed checkpointing algorithm used in Apache Flink
    • Stop-the-world vs asynchronous checkpointing
    • Consistent vs inconsistent checkpoints
  • Hands-on: Implement asynchronous, consistent, distributed checkpointing in a simple Python-based standalone system

Pre-requisites

  • Have done basic courses in OS, computer networks, and parallel programming
  • Comfortable with Python

Expected Outcomes

  • Understand how distributed stream processing engines work
  • Learn about fault tolerance algorithms
  • Implement an asynchronous consistent checkpointing and a recovery procedure

Reference Material

Additional Setup / Equipment Requirements

A laptop with functional Python shall suffice.

← Back to Call for Tutorials