Glossary · Infrastructure
TensorRT-LLM
TensorRT-LLM is an open-source NVIDIA library that optimizes and runs large language model inference on NVIDIA GPUs, using techniques such as kernel fusion, quantization and in-flight batching.
TensorRT-LLM sits in the Infrastructure part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.