Our inference bill was on track to pass our payroll. This talk is the story of cutting tokens-per-dollar by 7x without a visible quality drop, in the order the savings actually arrived: response caching, prompt-prefix reuse, routing easy qu…
Amina Okafor
Director of ML Efficiency · Common Thread