Hi, I’m Santosh Koti, an IEEE Senior Member and Senior Staff Software Engineer specializing in distributed systems, cloud-native infrastructure, Kubernetes control planes, and AI infrastructure.

My work focuses on designing reliable, scalable, and low-latency platforms for production environments. I am particularly interested in applying distributed-systems and control-plane techniques to improve the reliability, efficiency, observability, and predictability of production LLM inference systems.

I build high-throughput backend, data-plane, and control-plane systems primarily using Go, along with Rust, Java, Kubernetes, AWS, and Terraform. My experience includes distributed state management, replication, caching, low-latency request paths, resource orchestration, observability, fault tolerance, and production-scale platform engineering.

I have designed and operated distributed platforms spanning more than 1,000 sites, supporting over 100 million requests per day, with strong consistency and single-digit-millisecond request-path latency requirements.

LinkedIn · GitHub · ORCID · Dev.to · Résumé

Research Focus

My current research is centered on the following question:

How can distributed-systems and control-plane techniques make production LLM inference more reliable, efficient, and predictable?

I am especially interested in problems that connect systems research with measurable production outcomes, including latency, throughput, availability, resource utilization, and operational safety.

Research Interests

Areas I am currently studying and exploring include:

Peer Review Interests

I am available to review scholarly and technical work in areas aligned with my professional and research expertise, including:

I am particularly interested in reviewing work that combines rigorous systems design with realistic evaluation under production-like workloads.

Selected Technical Work

Distributed Cache Operator

An open-source Kubernetes operator built using controller-runtime and consistent hashing to manage the lifecycle and topology of distributed-cache clusters.

The project explores Kubernetes reconciliation, declarative APIs, failure recovery, membership changes, state management, and safe distributed-system automation.

View the Distributed Cache Operator on GitHub

Production Distributed Data Plane

Designed a strongly consistent, low-latency distributed data plane supporting runtime configuration, experimentation, and feature delivery across more than 1,000 sites and over 100 million requests per day.

The platform included:

Technical Writing

I write about:

My articles are published on this site and on Dev.to.

Current Research Directions

My current work explores practical research questions such as:

Professional Background

Current position: Senior Staff Software Engineer Organization: Fanatics, Inc. Location: San Mateo, California, United States

My broader professional experience includes:

Professional Membership and Recognition

Technologies

Languages: Go, Rust, Java, Python Platforms: Kubernetes, AWS, Terraform Distributed systems: Replication, caching, leader election, consistency, reconciliation, rate limiting, circuit breaking Observability: Prometheus, Grafana, OpenTelemetry, Loki, Tempo AI infrastructure: LLM inference, model serving, scheduling, routing, performance analysis

Research and Professional Profiles

Collaboration and Reviewing

I welcome opportunities involving:

Contact

For research, peer-review, speaking, and professional correspondence:

[email protected]


The articles and opinions published here are my own and do not necessarily represent the views of my employer.