LLM Systems · Efficient Inference · Compute Allocation

Jingbo Wen (Tony)

I build and study infrastructure for large language models: inference serving, speculative decoding, and distributed training.

I am completing a B.Eng. (Hons) in Software Engineering at the University of Sydney. Most recently I was an Applied Scientist Intern at Microsoft, working on foundation-model serving; before that, an Agent Engineering Intern at Xiaohongshu (RedNote).

Open to LLM Engineer / LLM Infrastructure roles.

Places

Drag to turn the globe · pick a place to fly there

My research is on allocating inference compute: consequence-aware reasoning budgets and visual token compression, budget-robust speculative decoding, reward-aware execution gating for agents, and market-aware routing across inference providers.

  1. Not All Errors Are Equal: Consequence-Aware Reasoning Compute Allocation

    Liang He, Jingbo Wen, Haoyu Wang, Ziqi He, Yixiong Chen, Kangning Cui, Xilu Wang

    Preprint · arXiv:2606.04402 [cs.AI] · Jun 2026

    22–33% lower cost-weighted loss than difficulty-aware compute routing on SWE-bench Lite.

  2. BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding

    Liang He, Jingbo Wen, Qishi Zhan, Yixiong Chen, Kangning Cui, Qizhen Lan, Xilu Wang

    Preprint · arXiv:2606.00144 [cs.LG] · May 2026

    One budget-robust drafter for sparse KV Cache; 6.55× end-to-end speedup at 4K context.

  3. Not All Visual Tokens Are Equally Safe to Remove: Consequence-Sensitive Visual Token Compression

    Jingbo Wen, Liang He, Mingyu Cao, Haoyu Wang, Minxuan Hu, Kangning Cui, Xilu Wang

    Preprint · arXiv:2608.09176 [cs.CV] · Aug 2026

    High-stakes VLM errors 0.300 → 0.133 at fixed token budget; 38% lower cost-weighted error.

  4. From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents

    Liang He, Jingbo Wen, Hongyu Gu, Hao Li, Haoyu Wang, Yixiong Chen, Kangning Cui, Xilu Wang

    Preprint · arXiv:2608.09168 [cs.AI] · Aug 2026

    RADEG skips 68% of agent calls while retaining 61% of total reward.

  5. You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference

    Liang He, Jingbo Wen, Yixiong Chen, Yue Yang, Qizhen Lan, Kangning Cui, Xilu Wang

    Preprint · arXiv:2609.37902 [cs.AI] · Sep 2026

    Price does not predict quality; measured routing saves ~50% at matched quality.

Microsoft

Applied Scientist InternAzure community infra

  • Built a dynamic-batching request scheduler in Go for Foundation Model serving (routing, admission control, deadlines), sustaining 2K+ concurrent streams; cut p99 TTFT by 28% under bursty traffic.
  • Integrated EAGLE-3 speculative decoding into the serving runtime, with per-request KV-cache rollback so batches with unequal acceptance lengths verify in one pass; 1.8× tokens/s at mean acceptance length 3.6.
  • Profiled the scheduler-to-GPU path with pprof and Nsight Systems; traced a TPOT regression to per-token scheduler–runtime RPC overhead and removed it with batched streaming, lowering TPOT by 17%.
  • Built benchmarking and observability for TTFT, TPOT, tokens/s, and KV-cache usage across batch sizes and speculation configs; its load sweeps set the concurrency threshold above which speculation is disabled.

Xiaohongshu (RedNote)

Agent Engineering InternSocial Networking Engineering

  • Designed and developed an internal Coding Agent from 0→1, enabling autonomous codebase understanding, multi-file editing, tool use, code execution, and iterative debugging from natural-language instructions.
  • Built a 0→1 OCR verification pipeline (extraction, validation, exception handling) for a production mobile app and took it from prototype to launch: 98% verification accuracy, 21K+ users verified on day one.
  • Optimized LLM serving (vLLM, INT4 quantization, continuous batching): +45% QPS, −40% cost.
  • Built and maintained internal LLM training and evaluation pipelines supporting SFT / DPO / RLHF workflows.

Distributed LLM Training & Inference System

Personal projectLLM Systems / Distributed Computing

  • Implemented a 350M–1.3B GPT-style Transformer from scratch (GQA, RoPE, SwiGLU, RMSNorm).
  • Built Megatron-style Tensor Parallelism (Column/RowParallelLinear, VocabParallelEmbedding) and Pipeline Parallelism (GPipe / 1F1B scheduling, P2P activation transfer, pipeline bubble analysis).
  • Integrated DeepSpeed ZeRO-1/2/3 with CPU Offload into a 3D parallel training stack (TP × PP × DP), pre-training a 1.3B model on 4×A100 (TP=2, PP=2) over 1.5B tokens.
  • Built a KV Cache inference engine (Prefill/Decode separation, dynamic batching, O(n) per-step attention).
  • Wrote custom CUDA C++ and Triton kernels (elementwise fusion, reduction, softmax/LayerNorm, tiled matmul) and profiled them against PyTorch native ops with Nsight Compute.
LLM Training
LoRA · SFT / DPO / RLHF · DDP · Tensor / Pipeline Parallelism · DeepSpeed ZeRO-1/2/3 · NCCL
LLM Inference
KV Cache · Paged Attention · Speculative Decoding · Quantization · vLLM · TensorRT-LLM
GPU & Kernel
CUDA C++ · Triton · Nsight Compute · CUDA Streams / Events
ML Engineering
PyTorch · Go · FAISS · ChromaDB · RAG · FastAPI · Docker · Linux · Git

University of Sydney

B.Eng. (Hons), Software Engineering (ECE)

  • GPA 3.8 / 4.0
  • 2025 Dean’s List
  • TOEFL iBT 110 / 120 (5.5 / 6)

I write 拆解大模型 (“Taking LLMs Apart”), a series in Chinese on Juejin about how large language models work: language modeling, Transformer internals, attention, and training.

Read on Juejin