Microsoft
Jun 2026 – Sep 2026
Applied Scientist InternAzure community infra
- Built a dynamic-batching request scheduler in Go for Foundation Model serving (routing, admission control, deadlines), sustaining 2K+ concurrent streams; cut p99 TTFT by 28% under bursty traffic.
- Integrated EAGLE-3 speculative decoding into the serving runtime, with per-request KV-cache rollback so batches with unequal acceptance lengths verify in one pass; 1.8× tokens/s at mean acceptance length 3.6.
- Profiled the scheduler-to-GPU path with pprof and Nsight Systems; traced a TPOT regression to per-token scheduler–runtime RPC overhead and removed it with batched streaming, lowering TPOT by 17%.
- Built benchmarking and observability for TTFT, TPOT, tokens/s, and KV-cache usage across batch sizes and speculation configs; its load sweeps set the concurrency threshold above which speculation is disabled.