Insights
Thinking on AI Infrastructure
and Inference Economics.
Perspectives on where and how to run production AI more economically. No provider names, no invented numbers—frameworks you can apply to your own workload.
- AI Inference Economics
- GPU Infrastructure
- Neocloud Market Trends
- Infrastructure Optimization
- Model Serving
- Capacity Planning
Capacity Planning for Production Inference
How to size reserved capacity to the predictable base of a workload, handle variability, and avoid paying for idle GPUs.
Model Serving Levers That Change Your Unit Economics
Batching, quantization, memory utilization, and accelerator selection—the serving decisions that determine how much output a GPU actually produces.
Why $/GPU-Hour Is the Wrong Number to Optimize
GPU hourly pricing is easy to compare and easy to get wrong. A framework for evaluating AI inference infrastructure on cost per unit of output instead.
Reading the Specialized AI Infrastructure Market
A framework for understanding how Neocloud and specialized GPU infrastructure providers differ from hyperscale clouds—and when each fits.
Stay informed
Infrastructure notes, when there’s something worth saying.
Occasional analysis on AI infrastructure economics and market structure. No marketing cadence.
Next step
Prefer to Talk About Your Workload Instead?
Reading about unit economics is useful. Measuring yours is better.