How many conversations fit on one GPU
MHA 32L×32h
GQA 32L×8h
GQA 80L×8h
context length
8K
HBM left for cache
60 GB