What our V100 + vLLM stack actually feels like (Qwen2.5-7B)

Four Tesla V100s, one vLLM engine and 360 real ShareGPT conversations. Every metric defined before it is reported, every serving flag explained, and an honest read on which workloads this hardware fits.

How to read LLM inference numbers (before you trust a tok/s chart)

A big tok/s chart is not enough. Learn TTFT, ITL, goodput, and why ISL, OSL, and concurrency belong on every inference label before you trust the number.