
Blog
Articles tagged self-hosted-llm from Whatever Works.

AMD's Threadripper Halo pairs a 96-core CPU with up to 576GB of HBM3e — the first deskside box that can hold a trillion-parameter model without a datacenter. Here's what that memory actually buys, what it costs, and who should seriously consider it.
Read article →
GPT-6 Astra is being reported at roughly $10 in / $50 out per million tokens with a ~1M-token context window — flagship-grade money. Here is the honest math on why that bill is set by your architecture, not the model, and how to keep it affordable.
Read article →