DatatalkwithAnkit

Writing

GPU comparisons, LLM inference benchmarks, and notes on the AI market. Benchmarks ship with the configs that produced them.

Aug 21, 2026

H200 vs MI300X: what you're actually paying for

Memory, interconnect, and price compared. What the extra 51 GB buys, why PCIe vs NVLink decides multi-GPU scaling, and what a GPU-hour really costs.

9 min read
Coming soon

Quantization in practice: quality vs. cost on consumer GPUs

What happens to output quality and tokens per second going from FP16 to INT8 and INT4, measured on real prompts.

In progress
Coming soon

How batch size changes vLLM throughput

A single-GPU sweep of max_num_seqs, measuring tokens per second against p50 and p95 latency.

In progress
Coming soon

KV cache is the real memory budget

Why context length, not parameter count, usually decides how many users one GPU can hold.

In progress
Coming soon

Cost per million tokens: self-hosted vs. API

A break-even analysis including the costs people forget: idle GPU time, operations, and engineering hours.

In progress
Coming soon

Building a fully automated YouTube pipeline in Python

Agents orchestrating FFmpeg, Whisper captions, loudness normalization, and scheduled uploads.

In progress

Get new posts by email

New GPU and AI infrastructure posts, sent when they're published.

Subscribe