#llm

4 posts tagged llm.

Jul 16, 2026 · 07:0013 min
I Was Measuring the Wrong Weights
Two weeks ago I said the DGX Spark's bandwidth wall was physics, not a driver update — fast or big, pick one. Then I ran a 122B model on the same box at six times the speed of a 70B. The law was right. The number I kept feeding it wasn't.
Jul 4, 2026 · 06:009 min
I Bought the GPU Anyway
I bought Andromeda, a DGX Spark, to fine-tune models — but first I owed my last post a rematch. I re-ran the same local/free/paid LLM inference benchmark on 120 GB of Grace-Blackwell instead of a CPU NUC. It fixes the one thing the NUC got fatally wrong, hits a bandwidth ceiling I didn't see coming, and makes clear why the real reason I bought it isn't inference at all.
Jun 27, 2026 · 17:007 min
Reading Android Bench: scores, cost, and the efficiency trap
Google's Android Bench scores LLMs on real Android engineering tasks. A look at what the June 2026 leaderboard actually says — why the top is a statistical tie, where the real cost-efficiency winners are, and the trap hiding in the cost and latency columns.
May 13, 2026 · 22:006 min
What Free LLMs Actually Cost
I bought a NUC to run local LLMs and escape API costs. Then I tested local, free cloud, and paid cloud against the same prompts. The result wasn't 'pay for inference' or 'run it yourself' — it was that 'fast and free' isn't a quadrant that exists right now, and that has implications for anyone planning a personal AI setup.