CloudX 2026: How Do You Forecast Hardware at Billion-User Scale?
Invited speaker at CloudX 2026 (Santa Clara, Sept 1-3), co-located with API World + AI TechWorld. Speaking on how to forecast hardware capacity for AI infrastructure at billion-user scale.
I'll be speaking at CloudX 2026 (co-located with API World + AI TechWorld) in Santa Clara, CA (September 1-3) on how to forecast hardware capacity for AI infrastructure at billion-user scale.
Capacity planning breaks down fast once GPU demand, model iteration speed, and traffic growth stop moving in sync. This session shares a practical framework for forecasting hardware and compute needs before they become a bottleneck, drawn from building and operating capacity-planning systems for production AI inference at scale.
Key takeaways
- Why traditional capacity planning models fail once GPU/MIG-based inference enters the picture
- A replica-centric approach to forecasting hardware demand
- Balancing headroom against cost at billion-user scale
- Signals that predict a capacity crunch before it hits production