Simulating AI Infrastructure at Scale
Simulation-Driven Research on AI Systems, Networks, Memory, and Fabrics — IndoSys 2026 Tutorial
Speakers
- Abed Mohammad Kamaluddin, Marvell Technology
- Hemant Singh, Marvell Technology
- Sriram Vatala, Marvell Technology
- Ramanjeet Singh, Marvell Technology
Tutorial Overview
Modern AI systems are increasingly limited by the performance of the underlying infrastructure rather than compute alone. Communication fabrics, memory systems, network topology, and system architecture play a critical role in scaling large language models and foundation models.
This tutorial introduces researchers to simulation-driven methodologies for studying large-scale AI infrastructure. Participants will learn the fundamentals of distributed AI systems, gain hands-on experience with Astra-Sim for system-level simulation, and explore emerging research directions in AI networks, fabrics, memory systems, and LLM infrastructure.
The tutorial combines conceptual foundations with guided hands-on exercises and demonstrations of advanced AI infrastructure simulation techniques.
Day 1 – Foundations and AI Systems Simulation (1.5 Hours)
1. AI Infrastructure Fundamentals (Conceptual)
- Evolution of AI infrastructure beyond the GPU
- Distributed AI training and inference
- Parallelism strategies:
- Data Parallelism
- Tensor Parallelism
- Pipeline Parallelism
- Expert (MoE) Parallelism
- Collective communication primitives: AllReduce, AllGather, ReduceScatter
- AI communication fabrics: Ultra Ethernet, InfiniBand, NVLink, CXL
- Why simulation is essential for AI infrastructure research
2. Introduction to Astra-Sim (Conceptual + Hands-on)
- Astra-Sim architecture and design philosophy
- Workload modeling
- System topology modeling
- Running baseline simulations
- Understanding simulation outputs
3. Hands-on Lab (Hands-on)
Participants will:
- Configure AI workloads
- Modify network topologies
- Compare communication configurations
- Evaluate different collective communication strategies
- Measure: training iteration time, communication overhead, collective completion time, scalability
- Analyze system bottlenecks using simulation results
Day 2 – Research Applications and Emerging Directions (1.5 Hours)
4. Simulation-Driven AI Infrastructure Research (Conceptual)
- Design-space exploration
- Performance modeling and scalability analysis
- Bottleneck identification
- Using simulation to formulate and validate AI systems research
5. Advanced AI Infrastructure Simulation (Conceptual + Demonstration)
Network and Fabric Simulation:
- Fine-grained network simulation using HTSim
- AI transport protocols and congestion control
- AI fabrics and collective communication
- Network topology exploration and performance analysis
Memory Systems Simulation:
- Memory hierarchy in AI systems
- Modeling HBM, DDR, and CXL memory
- Memory capacity versus bandwidth trade-offs
- Memory disaggregation and near-memory computing
LLM Systems Simulation:
- Modeling LLM training and inference
- Communication and memory behavior of foundation models
- KV-cache modeling and optimization
- Mixture-of-Experts (MoE) communication patterns
6. Research Frontiers in AI Infrastructure (Conceptual)
- Co-design of compute, memory, storage, and networks
- AI-native fabrics and next-generation communication protocols
- Memory-centric AI architectures
- Topology-aware scheduling and collective optimization
- Emerging directions in AI infrastructure simulation
Pre-requisites
- Basic knowledge of Computer Architecture
- Basic knowledge of Computer Networks
- Familiarity with Linux command-line usage is desirable
No prior experience with AI infrastructure simulators is required.
Expected Outcomes
By the end of the tutorial, participants will be able to:
- Understand the architecture of modern AI infrastructure beyond GPUs
- Explain the impact of communication, networks, fabrics, and memory on AI system performance
- Use Astra-Sim to model distributed AI workloads
- Evaluate AI infrastructure design choices through simulation
- Analyze simulation outputs to identify system bottlenecks
- Apply simulation methodologies to AI systems and infrastructure research
- Identify emerging research opportunities in AI networks, fabrics, memory systems, and large-scale AI infrastructure
Reference Material
Participants will be provided with:
- Astra-Sim documentation and tutorials
- HTSim documentation
- Tutorial slides and hands-on lab guide
- Sample workloads and configuration files
- Curated reading list covering: distributed AI systems, AI communication fabrics, collective communication, memory systems for AI, large-scale AI infrastructure simulation
Additional Setup / Equipment Requirements
Venue Requirements:
- Projector and presentation system
- Reliable Wi-Fi connectivity
- Power outlets for participant laptops
Participant Requirements:
- Personal laptop
- Docker installed (preferred) or Python 3.10+ with required dependencies
- Approximately 10–15 GB free disk space
Software: A pre-configured tutorial package (or Docker image) will be provided containing Astra-Sim, required software dependencies, example workloads, configuration templates, and hands-on lab scripts.
