Hi, I’m Santosh Koti, an IEEE Senior Member and Senior Staff Software Engineer specializing in distributed systems, cloud-native infrastructure, Kubernetes control planes, and AI infrastructure.
My work focuses on designing reliable, scalable, and low-latency platforms for production environments. I am particularly interested in applying distributed-systems and control-plane techniques to improve the reliability, efficiency, observability, and predictability of production LLM inference systems.
I build high-throughput backend, data-plane, and control-plane systems primarily using Go, along with Rust, Java, Kubernetes, AWS, and Terraform. My experience includes distributed state management, replication, caching, low-latency request paths, resource orchestration, observability, fault tolerance, and production-scale platform engineering.
I have designed and operated distributed platforms spanning more than 1,000 sites, supporting over 100 million requests per day, with strong consistency and single-digit-millisecond request-path latency requirements.
LinkedIn · GitHub · ORCID · Dev.to · Résumé
Research Focus
My current research is centered on the following question:
How can distributed-systems and control-plane techniques make production LLM inference more reliable, efficient, and predictable?
I am especially interested in problems that connect systems research with measurable production outcomes, including latency, throughput, availability, resource utilization, and operational safety.
Research Interests
Areas I am currently studying and exploring include:
- Reliable and efficient LLM inference systems
- Distributed and disaggregated model serving
- Inference-aware, locality-aware, and KV-cache-aware request routing
- Prefill and decode scheduling
- Resource scheduling, admission control, and overload protection
- Kubernetes-based AI platforms and inference control planes
- Autoscaling and fleet lifecycle automation
- Stale telemetry and delayed control-plane state
- Metric fidelity and misleading infrastructure signals
- Distributed control planes and reconciliation systems
- Fault tolerance and graceful degradation
- Observability and OpenTelemetry
- Performance engineering for cloud-native systems
- Edge and hybrid-cloud AI infrastructure
Peer Review Interests
I am available to review scholarly and technical work in areas aligned with my professional and research expertise, including:
- Distributed systems
- Cloud computing
- Kubernetes and container orchestration
- Cloud-native systems
- Control planes and reconciliation
- Resource management and scheduling
- Autoscaling and admission control
- Fault tolerance and service reliability
- Distributed caching and state management
- Observability and telemetry systems
- Performance evaluation and benchmarking
- AI infrastructure and distributed machine learning
- LLM inference and model serving
- Edge computing and hybrid-cloud platforms
I am particularly interested in reviewing work that combines rigorous systems design with realistic evaluation under production-like workloads.
Selected Technical Work
Distributed Cache Operator
An open-source Kubernetes operator built using controller-runtime and consistent hashing to manage the lifecycle and topology of distributed-cache clusters.
The project explores Kubernetes reconciliation, declarative APIs, failure recovery, membership changes, state management, and safe distributed-system automation.
View the Distributed Cache Operator on GitHub
Production Distributed Data Plane
Designed a strongly consistent, low-latency distributed data plane supporting runtime configuration, experimentation, and feature delivery across more than 1,000 sites and over 100 million requests per day.
The platform included:
- Full and incremental state propagation
- Replication and consistency mechanisms
- Low-latency in-memory request paths
- Distributed caching
- Failure recovery
- Observability and operational safeguards
- Support for concurrent experiments and configuration changes
Technical Writing
I write about:
- Distributed systems
- Kubernetes controllers and control planes
- Cloud-native reliability
- Production observability
- AI infrastructure
- LLM-serving systems
- Performance and failure analysis
My articles are published on this site and on Dev.to.
Current Research Directions
My current work explores practical research questions such as:
- How stale telemetry affects autoscaling and routing decisions
- How delayed control-plane state can trigger unsafe infrastructure actions
- How metric aggregation can hide inference overload
- How request routing affects KV-cache reuse and time to first token
- How inference systems should degrade under resource pressure
- How Kubernetes controllers can safely reconcile rapidly changing AI workloads
- How to distinguish queueing, compute, memory-bandwidth, and cache-capacity bottlenecks
Professional Background
Current position: Senior Staff Software Engineer Organization: Fanatics, Inc. Location: San Mateo, California, United States
My broader professional experience includes:
- Distributed backend and platform systems
- Kubernetes operators and controllers
- Cloud infrastructure and automation
- High-throughput, low-latency services
- Runtime configuration and experimentation platforms
- Distributed caching and replication
- Reliability engineering and fault tolerance
- Metrics, logging, tracing, and production observability
Professional Membership and Recognition
- IEEE Senior Member
- Member, IEEE Computer Society
- ORCID researcher profile
Technologies
Languages: Go, Rust, Java, Python Platforms: Kubernetes, AWS, Terraform Distributed systems: Replication, caching, leader election, consistency, reconciliation, rate limiting, circuit breaking Observability: Prometheus, Grafana, OpenTelemetry, Loki, Tempo AI infrastructure: LLM inference, model serving, scheduling, routing, performance analysis
Research and Professional Profiles
Collaboration and Reviewing
I welcome opportunities involving:
- IEEE and ACM peer reviewing
- Research collaboration
- Applied systems research
- Conference and workshop reviewing
- Technical program committees
- Open-source systems projects
- Industry–academic collaboration
- Technical speaking and panel participation
Contact
For research, peer-review, speaking, and professional correspondence:
The articles and opinions published here are my own and do not necessarily represent the views of my employer.