TensorRT-LLM in Numbers: FP16 vs. FP8 on the RTX A6000 Ada

In the first three posts of this little series I explained why I’m tackling TensorRT-LLM on...

Read More