Modern processors, memory systems, and I/O devices employ increasingly complex hardware mechanisms, making application performance difficult to understand and predict. We develop performance models that capture how applications interact with these mechanisms, explain observed behavior, and predict the impact of system changes. We use these models to identify bottlenecks, evaluate design trade-offs, and guide hardware and software optimizations.
AI services combine compute-intensive model execution with memory- and data-intensive processing. As models and datasets grow, accelerator performance alone is insufficient: memory capacity, data movement, and communication can limit system performance. We develop hardware and software techniques that address these bottlenecks in AI training and inference, as well as supporting infrastructure such as vector databases.
Datacenters spend substantial computing resources on infrastructure tasks in addition to application execution. As systems scale, the overhead of memory management, I/O processing, and data management can limit overall efficiency. We design CPU architectures and system software that reduce these overheads, and explore hardware offloading with SmartNICs, DPUs, and FPGAs to make more resources available to applications.
Emerging high-speed network technologies, such as Ultra Ethernet, are changing how large-scale AI and distributed systems communicate. However, higher link bandwidth alone does not eliminate bottlenecks in host processing, data movement, and network congestion. We optimize communication across hosts, NICs, and switches through efficient transport mechanisms and in-network acceleration to improve end-to-end performance and scalability.