About Me

Ankur Gupta
Senior Staff Technical Program Manager
AI Infrastructure, LinkedIn · San Jose, CA
I turn ambiguous, high-stakes technical problems into systems that actually ship, and stay reliable once they do.
I write about AI infrastructure, reliability, systems, and platform engineering, and the technical program management that ships it.
Hi, I'm Ankur. I'm a product-minded technology leader with over fifteen years in AI infrastructure, cloud platforms, and large-scale distributed systems. As a Senior Staff Technical Program Manager at LinkedIn, I lead high-impact AI infrastructure and platform programs serving over a billion members, work that spans GPU-accelerated inference, AI/ML deployment pipelines, inference optimization, reliability engineering, capacity planning, data center buildouts, and a first-of-its-kind hardware forecasting control plane. Across roles at LinkedIn and Walmart, my focus has stayed the same: turn complex technical challenges into scalable platforms that deliver measurable business outcomes. My work sits at the intersection of program management, platform engineering, and AI infrastructure. I'm an IEEE Senior Member, and I write here about the craft behind it all.
Talks, publications & appearances
A running record of talks, publications, podcasts, and presentations. This updates automatically as new ones go live.
Budapest Data + AI Forum 2026
Speaker: why ML deployments fail in production, and how to engineer reliability at scale.
How automation is changing managerial roles
BSmart, Manager's Mantra: on automation, cloud strategy, and the evolving manager.
PM, TPM & EM: A Practical Framework
Published on Stackademic / Medium: a clear way to think about three leadership roles.
Also publishing with Stackademic / Medium and the IEEE Computer Society (ongoing), plus conference speaking on AI infrastructure and reliability. See all highlights →
Work I'm proud of
GPU Efficiency & AI Inference
Partnered with principal engineers on a disaggregated serving architecture for ML inference at scale.
Model Deployment & Velocity
Scaled the ML deployment and experimentation platform powering ranking and retrieval across LinkedIn's core products.
Infrastructure Capacity Forecasting
Defined the architecture, data models, and product design for a centralized forecasting platform; delivered with 22 teams in 10 sprints.
SRE3 Reliability Transformation
Shifted reliability ownership to product teams across 225 teams and 1,400+ services via platform automation and self-service tooling.
Next-Gen Data Center Bring-up
Led the software platform bring-up for LinkedIn's next-generation data center on Kubernetes-based deployment frameworks.
SAP Hybrid Cloud Migration
Designed a highly available SAP hybrid-cloud platform and migrated $30M of on-prem infrastructure to Azure for 40 teams across 11 countries.
The path here
Senior Staff Technical Program Manager
Staff Technical Program Manager
M.S., Information Management (Data Science)
Program Management Intern
Systems Engineer (Java Developer, SDET)
Credentials & awards
Awards & recognition
- Mountain Mover Award — LinkedIn (2024)
- SRE TPM Excellence Award — LinkedIn (2023)
- Associate of the Quarter — Walmart Global Tech (2020)
- Global Recognition Award — Bank of America (2015)
- INSTA Award — Infosys (2014)
- Most Valuable Player (MVP) — Infosys (2012)
Education
- M.S., Information Management
Syracuse University, iSchool (2014–2016) · GPA 3.94/4.0 - Certificate of Advanced Studies, Data Science
Syracuse University (2016) · 4.0/4.0 - B.E., Computer Science
Maharshi Dayanand University, India (2006–2010)
What I write about
Why I write, who it's for, and the questions I keep coming back to.
I'm a technical program manager who has spent more than fifteen years in the engine room of large engineering organizations, the place where ambitious ideas run into messy reality. I've led reliability transformations across hundreds of teams, built capacity and forecasting platforms, helped bring up new data centers, and today I focus on the AI infrastructure that deploys and serves machine learning models at scale. Most of that work lives in the space between teams, where the real problem is rarely a single piece of code.
I write because the hardest parts of building large systems are usually the least visible: ambiguity, coordination, reliability, and the long gap between something that works in a demo and something that survives production. Early in my career I would have given a lot for clear, honest explanations of how these systems actually fit together, and why programs really succeed or fail. So I write the pieces I wish I'd had, plain-language mental models for engineers and for the people who lead them.
Most large programs don't fail because nobody saw the risks coming. They fail because the warning signs were visible for months, and no one forced a decision. My job is to surface what's being avoided, and make sure it gets acted on.— On the craft of program management
My writing falls into three threads. First, the craft of technical program management: leading without authority, turning ambiguity into roadmaps, managing dependencies, and escalating without burning bridges. Second, the engineering fundamentals every TPM and product leader should understand, from databases and networking to caching, observability, and system design. Third, AI infrastructure and production reliability: why machine learning deployments fail, how to serve models efficiently, and what it takes to make modern AI systems dependable.
Published research
IEEE papers (2026)
- Characterizing GPU Capacity and Computational Cost in Production AI Inference Using a Replica-Centric Measurement Framework
IEEE International Workshop on Metrology for Industry 4.0 & IoT (MetroInd4.0 & IoT), 2026 - Operationalizing Site Reliability in Large-Scale Distributed Systems: Shifting Ownership Left
IEEE 15th International Conference on Communication Systems and Network Technologies (CSNT), 2026 · pp. 1438–43
Focus
- Peer-reviewed research on production AI inference efficiency, GPU capacity measurement, and operating reliability across large-scale distributed systems.
Reviewing for the field
126 verified peer reviews recorded on Web of Science, across 34 international conferences and journals spanning AI, security, data mining, materials, and engineering. I also serve as a reviewer for Elsevier's Engineering Applications of Artificial Intelligence.
For editors and press
I write deeply-researched, practitioner-grounded pieces on production AI infrastructure, GPU efficiency, and reliability at scale. I am available for bylined articles, op-eds, technical deep-dives, and expert commentary, and I am happy to write to a brief.
- From 60% to 99%: engineering reliable model deployments at scale
- Getting 4.4× more from the same GPUs without changing the model
- Everyone wants to scale AI. Nobody has figured out how to afford it.
- Stop forecasting hardware, start forecasting intent
- Replacing a thousand spreadsheets with one forecasting platform
- Shifting reliability left: giving ownership back to product teams
- Why centralized SRE teams don't scale, and what to do instead
- Chaos engineering isn't about breaking things, it's about proving them
- The hardest part of building a data center has nothing to do with hardware
- Moving a billion users to a data center that didn't exist six months ago
- Why enterprise cloud migration is not lift-and-shift, and never should be
- Lessons from a multi-phase SAP cloud transformation
- Delivering billion-dollar technical programs without writing a line of code
- Leading large engineering programs without authority