Define the production problem.
We begin with a bottleneck observed in a real inference workload and turn it into a measurable validation target.
- Capture the model, hardware, workload, and traffic profile.
- Choose the primary metric: memory, latency, or throughput.
- Set quality and reliability limits before testing.