About Me

Ankur Gupta
About

Ankur Gupta

Senior Staff Technical Program Manager
AI Infrastructure, LinkedIn · San Jose, CA

I turn ambiguous, high-stakes technical problems into systems that actually ship, and stay reliable once they do.

I write about AI infrastructure, reliability, systems, and platform engineering, and the technical program management that ships it.

Hi, I'm Ankur. I'm a product-minded technology leader with over fifteen years in AI infrastructure, cloud platforms, and large-scale distributed systems. As a Senior Staff Technical Program Manager at LinkedIn, I lead high-impact AI infrastructure and platform programs serving over a billion members, work that spans GPU-accelerated inference, AI/ML deployment pipelines, inference optimization, reliability engineering, capacity planning, data center buildouts, and a first-of-its-kind hardware forecasting control plane. Across roles at LinkedIn and Walmart, my focus has stayed the same: turn complex technical challenges into scalable platforms that deliver measurable business outcomes. My work sits at the intersection of program management, platform engineering, and AI infrastructure. I'm an IEEE Senior Member, and I write here about the craft behind it all.

15+
Years in infrastructure & AI platforms
4.4×
GPU inference efficiency gain
99.98%
Service availability delivered
225
Teams in reliability transformation
1B+
Members served by systems I build
Selected work

Talks, publications & appearances

A running record of talks, publications, podcasts, and presentations. This updates automatically as new ones go live.

Also publishing with Stackademic / Medium and the IEEE Computer Society (ongoing), plus conference speaking on AI infrastructure and reliability. See all highlights →

Flagship programs

Work I'm proud of

LinkedIn · AI Serving

GPU Efficiency & AI Inference

Partnered with principal engineers on a disaggregated serving architecture for ML inference at scale.

90 → 400 QPS/GPU at 50ms p99 — 4.4× efficiency, ~$20M saved
LinkedIn · ML Platform

Model Deployment & Velocity

Scaled the ML deployment and experimentation platform powering ranking and retrieval across LinkedIn's core products.

Deploy success ~60% → toward 99%; rollouts days → <30 min
LinkedIn · Capacity

Infrastructure Capacity Forecasting

Defined the architecture, data models, and product design for a centralized forecasting platform; delivered with 22 teams in 10 sprints.

10,000+ manual hours removed — ~$350M (FY24), ~$480M (FY25) reduced
LinkedIn · Reliability

SRE3 Reliability Transformation

Shifted reliability ownership to product teams across 225 teams and 1,400+ services via platform automation and self-service tooling.

99.98% availability, 61% faster mitigation, 33% faster detection
LinkedIn · Data Center

Next-Gen Data Center Bring-up

Led the software platform bring-up for LinkedIn's next-generation data center on Kubernetes-based deployment frameworks.

200+ teams · automated turn-up · repeatable DC launches
Walmart Global Tech · SAP

SAP Hybrid Cloud Migration

Designed a highly available SAP hybrid-cloud platform and migrated $30M of on-prem infrastructure to Azure for 40 teams across 11 countries.

350+ apps · 84% faster deploys, $2M vendor savings
Career

The path here

Dec 2021 — Present

Senior Staff Technical Program Manager

LinkedIn · Mountain View, CA
Jun 2016 — Dec 2021

Staff Technical Program Manager

Walmart Global Tech · Bentonville, AR
Jul 2014 — May 2016

M.S., Information Management (Data Science)

Syracuse University · New York
Jun 2015 — Aug 2015

Program Management Intern

Bank of America
Nov 2010 — Jun 2014

Systems Engineer (Java Developer, SDET)

Infosys · Chandigarh, India
Recognition

Credentials & awards

Awards & recognition

  • Mountain Mover Award — LinkedIn (2024)
  • SRE TPM Excellence Award — LinkedIn (2023)
  • Associate of the Quarter — Walmart Global Tech (2020)
  • Global Recognition Award — Bank of America (2015)
  • INSTA Award — Infosys (2014)
  • Most Valuable Player (MVP) — Infosys (2012)

Education

  • M.S., Information Management
    Syracuse University, iSchool (2014–2016) · GPA 3.94/4.0
  • Certificate of Advanced Studies, Data Science
    Syracuse University (2016) · 4.0/4.0
  • B.E., Computer Science
    Maharshi Dayanand University, India (2006–2010)
Writing & ideas

What I write about

Why I write, who it's for, and the questions I keep coming back to.

I'm a technical program manager who has spent more than fifteen years in the engine room of large engineering organizations, the place where ambitious ideas run into messy reality. I've led reliability transformations across hundreds of teams, built capacity and forecasting platforms, helped bring up new data centers, and today I focus on the AI infrastructure that deploys and serves machine learning models at scale. Most of that work lives in the space between teams, where the real problem is rarely a single piece of code.

I write because the hardest parts of building large systems are usually the least visible: ambiguity, coordination, reliability, and the long gap between something that works in a demo and something that survives production. Early in my career I would have given a lot for clear, honest explanations of how these systems actually fit together, and why programs really succeed or fail. So I write the pieces I wish I'd had, plain-language mental models for engineers and for the people who lead them.

Most large programs don't fail because nobody saw the risks coming. They fail because the warning signs were visible for months, and no one forced a decision. My job is to surface what's being avoided, and make sure it gets acted on.— On the craft of program management

My writing falls into three threads. First, the craft of technical program management: leading without authority, turning ambiguity into roadmaps, managing dependencies, and escalating without burning bridges. Second, the engineering fundamentals every TPM and product leader should understand, from databases and networking to caching, observability, and system design. Third, AI infrastructure and production reliability: why machine learning deployments fail, how to serve models efficiently, and what it takes to make modern AI systems dependable.

TPM Craft System Design AI Infrastructure Reliability Engineering Fundamentals
Publications

Published research

IEEE papers (2026)

  • Characterizing GPU Capacity and Computational Cost in Production AI Inference Using a Replica-Centric Measurement Framework
    IEEE International Workshop on Metrology for Industry 4.0 & IoT (MetroInd4.0 & IoT), 2026
  • Operationalizing Site Reliability in Large-Scale Distributed Systems: Shifting Ownership Left
    IEEE 15th International Conference on Communication Systems and Network Technologies (CSNT), 2026 · pp. 1438–43

Focus

  • Peer-reviewed research on production AI inference efficiency, GPU capacity measurement, and operating reliability across large-scale distributed systems.
Peer review

Reviewing for the field

126 verified peer reviews recorded on Web of Science, across 34 international conferences and journals spanning AI, security, data mining, materials, and engineering. I also serve as a reviewer for Elsevier's Engineering Applications of Artificial Intelligence.

126
Verified peer reviews
34
Conferences & journals
13IEEE International Conference on AI and Security for Industrial IoT Systems
12International Conference on Network, Multimedia and Information Technology (NMITCON)
10International Conference on Advanced Materials for Sustainable Clean Energy and Healthcare Technologies
8IEEE International Conference on Electro-Information Technology (EIT)
6IEEE International Conference on Distributed Computing, VLSI, Electrical Circuits and Robotics
6Research Software Asia Australia Conference
5International Conference on Intelligent Computing and Knowledge Extraction (ICICKE)
5Intelligent Data Engineering & Artificial Intelligence
5International Conference on Future Technologies
4International Conference on Circuits, Control, Communication and Computing (I4C)
4IEEE Global Emerging Technologies Conference
4International Conference on Ambient Intelligence, Knowledge Informatics and Industrial Electronics
4ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
3International Conference on Communication Systems and Network Technologies
3International Conference on Emerging Trends in Networks and Computer Communications (ETNCC)
3International Conference on Emerging Technologies in Computing
3International Conference of Research Software in Africa
3BuildSim Nordic Conference
3International Conference on Synergies in Next-Generation Cyber-Physical Systems
2Programming and Computer Software
2UMaT Biennial International Mining and Mineral Conference
2International Conference on Innovative Trends in Electrical, Electronics and Bio-Technology Engineering
2IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB)
2International Conference on Industrial Cyber-Physical Systems
2International Conference on Future Technologies (ICFT)
2International Conference on Artificial Intelligence
1International Conference on Computational Intelligence and Communication Networks (CICN)
1Discover Computing
1Engineering Applications of Artificial Intelligence
1International Conference on Computer Vision & Image Processing
1UPADI International Conference on Artificial Intelligence
1International Conference on Sustainable Energy and Thermal Systems
1Conference on Uncertainty in Artificial Intelligence
1International Conference on Advanced Materials Science and Civil Engineering (AMSCE)
Write with me

For editors and press

I write deeply-researched, practitioner-grounded pieces on production AI infrastructure, GPU efficiency, and reliability at scale. I am available for bylined articles, op-eds, technical deep-dives, and expert commentary, and I am happy to write to a brief.

Topics I can write about
AI Infrastructure & ML Systems
  • From 60% to 99%: engineering reliable model deployments at scale
  • Getting 4.4× more from the same GPUs without changing the model
  • Everyone wants to scale AI. Nobody has figured out how to afford it.
Capacity & Hardware Forecasting
  • Stop forecasting hardware, start forecasting intent
  • Replacing a thousand spreadsheets with one forecasting platform
Reliability at Scale
  • Shifting reliability left: giving ownership back to product teams
  • Why centralized SRE teams don't scale, and what to do instead
  • Chaos engineering isn't about breaking things, it's about proving them
Data Centers & Large-Scale Migration
  • The hardest part of building a data center has nothing to do with hardware
  • Moving a billion users to a data center that didn't exist six months ago
Enterprise Cloud & SAP
  • Why enterprise cloud migration is not lift-and-shift, and never should be
  • Lessons from a multi-phase SAP cloud transformation
Technical Program Leadership
  • Delivering billion-dollar technical programs without writing a line of code
  • Leading large engineering programs without authority

Formats: bylined articles, op-eds, explainers, and expert quotes. I also share new work with a professional network on LinkedIn. Recent bylines, papers, and talks are in Highlights.

© Ankur Gupta