
We Need a Better Benchmark Than Tokens Per Second
When the benchmark ends, the graph is clean, and the card is pushing out tokens at a rate that makes the old machine look absurd. Someone has dutifully recorded the VRAM, the wattage, the price, the context length, the throughput. On paper, everything that counts has been counted. Then the person who bought the hardware goes back to the terminal. The model doesn’t load. Maybe it loads once, then crashes the second time. Yesterday’s driver fix has broken today’s container. A GitHub issue from ...

We Need a Better Benchmark Than Tokens Per Second
When the benchmark ends, the graph is clean, and the card is pushing out tokens at a rate that makes the old machine look absurd. Someone has dutifully recorded the VRAM, the wattage, the price, the context length, the throughput. On paper, everything that counts has been counted. Then the person who bought the hardware goes back to the terminal. The model doesn’t load. Maybe it loads once, then crashes the second time. Yesterday’s driver fix has broken today’s container. A GitHub issue from ...