Tanmay Sahay · तन्मय
Reliability for AI Systems
ymcy@tnl.iaa9m4aosamhag 66-177-50551-0+
Abstract. Site reliability engineer with 8+ years in production, at Google since 2019 across three platforms: LLM inference serving for Gemini and Vertex AI (Google DeepMind), Google's network infrastructure, and Google Cloud's serverless platform (Cloud Run, Cloud Functions, App Engine). I automate complex infrastructure, cut my team's oncall load by 50%, and build AI-assisted incident tooling that runs against live production.
1Positions
Software Engineer, SRE (Gemini & Vertex AI LLM serving · Google DeepMind)
Google · Apr '25 - Present · US-MTV
tl;dr — Making Google Cloud's large language model offerings (Gemini, Veo, Imagen) more reliable. Check out cloud.google.com/vertex-ai
— Reliability owner for accelerator-backed inference serving at frontier-model scale — GPU/TPU fleet capacity, serving-stack rollouts, and the incident path for Gemini, Veo and Imagen traffic.
— Led the migration of Vertex AI's ML model-serving configurations to dedicated hermetic ML-Ops, removing a class of release-time serving outages and materially improving rollout safety.
— Automated the months-long turnup of Vertex AI in new regions, including accelerator capacity validation before traffic admission.
— Built a Gemini-powered incident framework that cuts MTTx through LLM-assisted triage and gated actuation — an agentic system running against live production.
Software Engineer, SRE (Network Infrastructure)
Google · Feb '24 - Apr '25 · US-MTV
tl;dr — Ensuring reliability for Google's global backbone network telemetry and monitoring systems.
— Built a system that enables Netflix-style chaos-monkey testing for Google's network infrastructure.
— Played a pivotal role in migrating networking from the legacy monitoring system to the next-generation monitoring system.
— Built tooling to enhance network monitoring and alerting capabilities.
Software Engineer, SRE (Cloud Infrastructure)
Google · Feb '23 - Feb '24 · CH-ZRH
tl;dr — Continued driving reliability improvements and tooling development from Zurich.
— Seamlessly migrated fragmented internal users from 15-year-old + 4-year-old alert visualization tools to the next-gen tool by adopting the Google-wide experiment framework.
— Shipped the Alert2Bug migration in 1 week against a multi-month estimate by leaning on existing primitives instead of a bespoke rebuild.
— Continued development and expansion of Khoj (InvDash), scaling adoption across Google SRE teams.
— Mentored junior engineers and drove knowledge transfer across regions.
— Contributed to cross-functional infrastructure reliability initiatives.
— Prepared for and executed smooth transition to US-based Network Infrastructure team.
Software Engineer, SRE (Serverless Platform)
Google · Mar '19 - Feb '23 · UK-LON / CH-ZRH
tl;dr — Making Google Cloud's Serverless compute offerings (App Engine, Cloud Functions, Cloud Run) more reliable.
— Automated the months-long process of turning up Cloud Run in new regions.
— Enabled safe, progressive rollouts of schema and config changes across 560+ Spanner databases, eliminating the prior failure mode where bad changes affected customers globally for multiple hours.
— Built an actionable-metrics framework that democratized signal across SRE and partner teams, driving a 50% reduction in oncall load while the team onboarded 40% more partner services.
— Drove down resource ceilings across all Serverless Bigtables by adopting Autocap, materially improving fleet-wide compute efficiency.
— Led the Log4J code-red response across four Serverless products under incident pressure.
— Led 4 interns to build Khoj, a Google-wide automated incident root-causing system that reduced mean time to response and mitigation from multiple hours to just minutes.
Creator & Lead Developer — Khoj (InvDash)
Google (Internal Project) · 2021 - Present · Global
tl;dr — Built and continuously evolved an automated incident investigation and root-causing system.
— Conceived and built Khoj in 2021 to automate tedious incident root-cause analysis.
— System correlates logs, metrics, and change events to surface probable causes during outages.
— Adopted Google-wide across multiple SRE teams, significantly reducing Mean Time To Diagnose (MTTD).
— Continuously maintained and enhanced over 4+ years, adapting to new infrastructure patterns.
Software Developer
Booking.com · Jun '17 - Feb '19 · NL-AMS
tl;dr — Machine Learning Services & Image Infrastructure.
— Reduced image serving latency by 50% and storage costs by 80% via on-the-fly resizing service.
— Built ML platform features used by 200+ Data Scientists.
— Migrated image building pipelines to Google Cloud Dataproc.
2Methods & instruments
programming — Python, Go, Java, C++, SQL, Shell.
sre & cloud — Kubernetes, Terraform, Bazel, GCP, Incident Response, Observability.
ai & ml — Vertex AI, LLM Ops, Gemini CLI, AI Agents, Model Serving, TPU Fleet Mgmt.
tools — Prodspec, Spanner, Bigtable, Prometheus/Monarch.
3Selected peer review
43 peer bonuses, spot bonuses and kudos from 37 colleagues, 2019–2025 — for collaboration (12), technical excellence (10), incident response (8), mentoring (6), leadership (4), innovation (2), automation (1). In their words:
"Thanks Tanmay for introducing me and keeping me up to date with all the innovative things happening in the world of AI. Your presentation on gemini-cli and how to prompt was awesome. Your push towards using AI to…"— colleague, Vertex AI, 2025 · Innovation
"Thank you for going above and beyond the call of duty in your response to the log4j security vulnerabilities in December 2021. Your commitment to securing Google and our customers is truly appreciated!"— colleague, Serverless, 2022 · Incident Response
"Being oncall in an understaffed rotation takes time away from your project work and personal life. Thank you Tanmay for enabling our team to persevere through this challenging time!"— colleague, Serverless, 2022 · Collaboration
"For collaborating with Cloud Functions SREs to ensure alignment and shared understanding of observability requirements during the GCF v2 launch preparation."— colleague, Serverless, 2019 · Collaboration
4Education & honors
B.Tech in Computer Science
IIIT Hyderabad · 2013 - 2017 · Hyderabad, India
— ACM ICPC Regional Finalist (2014, 2015) — among top competitive programmers in Asia.
— JEE Mains Rank 719 (Top 0.05%) out of 1.4M candidates nationwide.
— National Talent Scholar — recognized for academic excellence.
— CodeChef Campus Chapter Ambassador — organized coding competitions.
— Teaching Assistant for Algorithms & Data Structures courses.
5Languages
germanic — English (native), Dutch (conversational), German (basic).
romance — French (conversational), Spanish (basic).
indo-aryan — Hindi (native), Urdu (conversational), Kannada (fluent), Sanskrit (exposure).
scripts — Cyrillic.
6Citation
@misc{sahay2026,
author = {Sahay, Tanmay},
title = {Reliability for AI Systems},
year = {2026},
note = {ICPC regionalist; polyglot (9 languages);
reachable at the address above}
}