Papers
Longer, structured write-ups: reproducible benchmarks, methodology notes, and deep dives on LLM configuration. Each one ships with its code and raw data, and preprints will be cross-posted when they're ready.
A practical evaluation framework for LLM inference configurations
A structured way to compare serving configurations, including quantization, batching, and context length, across throughput, latency, and cost, with an open-source harness.
Benchmarking vLLM serving parameters on single-GPU deployments
An empirical study of vLLM scheduler and memory settings on commodity hardware, mapping concurrency against tail latency and memory headroom.
Notes on reproducibility in LLM benchmarking
What an inference benchmark has to report, from hardware and warmup to workload shape and sampling settings, for someone else to reproduce the number.
How I run an experiment
The same six steps, every time.
Ask one narrow question
Something a number can answer, not an opinion.
Fix the setup
Same hardware, model, and workload. Warm up properly and record every parameter that could move the result.
Measure the spread
Multiple runs, reported as a range. Latency gets percentiles, never an average alone.
Publish the raw data
Configs, scripts, and unedited results go up next to the write-up.
Report what failed
Runs that contradict the hypothesis get written up too. They're usually the useful ones.
Find the next question
Every result raises a sharper question, and that becomes the next experiment.
Want to collaborate on a write-up?
If you're working on LLM evaluation or inference and want a second pair of hands, I'd like to hear about it.
Send a message