Ankit Dongare, AI engineer
I measure what language models cost to run.
I benchmark LLM inference: GPUs, serving configs, throughput, and price. When I run a benchmark, I publish the configs with it so you can rerun it and check my numbers.
Will your model fit?
Pick a model size and precision to see how many GPUs it takes to hold, and what that costs per hour on the cheapest published neocloud rates.
Rough estimate. Rates are the low end of published prices, Aug 2026: H200 $2.60/hr, MI300X $1.71/hr. How I got these numbers
Writing
GPU comparisons, inference benchmarks, and notes on the AI market. Posts go up when the work is done.
Latest post
H200 vs MI300X: what you're actually paying for
AMD sells memory, NVIDIA sells interconnect and a software stack that already works. Which is the better buy depends on whether your problem is fitting the model or making many GPUs act like one.
Memory per GPU
Quantization in practice: quality vs. cost on consumer GPUs
What happens to output quality and tokens per second going from FP16 to INT8 and INT4, measured on real prompts.
How batch size changes vLLM throughput
A single-GPU sweep of max_num_seqs, measuring tokens per second against p50 and p95 latency.
KV cache is the real memory budget
Why context length, not parameter count, usually decides how many users one GPU can hold.
Projects
Things I've built and run. More on GitHub.
My product
Aurapal
A career platform I design, build, and run end to end. It's where my engineering, automation, and product work come together.
Visit aurapal.org- Built soloFrom design and code to deployment and support.
- Automation-firstRepetitive work is scripted, so the time goes into the product.
- LiveRunning in production at aurapal.org.
Automated YouTube pipeline
Generates, renders, captions, and publishes videos with almost no manual steps.
Python, FFmpeg, Whisper, YouTube Data APILLM inference benchmarks
An ongoing series on how batch size, quantization, and context length change throughput, latency, and cost.
vLLM, PyTorch, PythonEWG, Elite White Glove
A conversion-focused website for a white-glove service business, from design to deployment.
HTML, CSS, VercelHNH Handyman
A booking-focused web app for a handyman service, with a typed, maintainable codebase.
TypeScriptAbout
I started as a data analyst turning messy tables into answers, and kept following the harder questions until they led to building AI systems full time.
Today I work on the practical side of large language models: how they're served, how fast they run, what they cost, and how configuration choices change all three. Most of my week goes into benchmarking models and tuning serving stacks like vLLM.
That data background still shapes how I work. I trust measurements over opinions, and if I do something twice, I write a script for it.
Day to day: Python, PyTorch, vLLM, Hugging Face, Docker, and AWS. The full list is on the uses page.
Stay in touch
Get new posts by email, or send me a note about a benchmark, a collaboration, or Aurapal.
The newsletter
New benchmarks and write-ups, sent when they're published. No schedule and no filler. Unsubscribe any time.
Signups come straight to me while the list is small.