Writing
GPU comparisons, LLM inference benchmarks, and notes on the AI market. Benchmarks ship with the configs that produced them.
H200 vs MI300X: what you're actually paying for
Memory, interconnect, and price compared. What the extra 51 GB buys, why PCIe vs NVLink decides multi-GPU scaling, and what a GPU-hour really costs.
Quantization in practice: quality vs. cost on consumer GPUs
What happens to output quality and tokens per second going from FP16 to INT8 and INT4, measured on real prompts.
How batch size changes vLLM throughput
A single-GPU sweep of max_num_seqs, measuring tokens per second against p50 and p95 latency.
KV cache is the real memory budget
Why context length, not parameter count, usually decides how many users one GPU can hold.
Cost per million tokens: self-hosted vs. API
A break-even analysis including the costs people forget: idle GPU time, operations, and engineering hours.
Building a fully automated YouTube pipeline in Python
Agents orchestrating FFmpeg, Whisper captions, loudness normalization, and scheduled uploads.
No posts on that topic yet. Pick another, or subscribe to hear when one lands.