Kong API + AI Summit 2026: Operational Lessons from Running Large-Scale AI Inference
Invited speaker at the Kong API + AI Summit 2026 in Los Angeles (Sep 30 – Oct 1). A practitioner session on the operational realities of running large-scale AI inference: deployment coordination, GPU and infrastructure constraints, performance tuning, and reliability under rapid model iteration.
I'll be speaking at the Kong API + AI Summit 2026 in Los Angeles (September 30 – October 1) on the operational realities of running large-scale AI inference in production.
Deploying a model is only the beginning. The real work starts when those models must serve millions of requests under strict latency, availability, and cost constraints, while infrastructure demand fluctuates and models iterate rapidly. This session shares practical lessons on deployment coordination, managing GPU and infrastructure constraints, performance tuning, and maintaining reliability under continuous change, drawn from operating inference platforms at scale.
Key takeaways
- Why running ML models in production introduces operational challenges that model-architecture work never surfaces
- Common reliability and scaling issues that show up in large AI inference systems
- How to balance rapid model iteration with production stability
- Operational strategies for infrastructure constraints and performance tuning
- Lessons from coordinating large engineering efforts around AI infrastructure