NVIDIA has introduced a practical pathway for running private, high-performance AI inference on trusted hardware through its Confidential Computing (CC) platform. As large language model (LLM) inference increasingly processes sensitive data—such as proprietary business models, personal user inputs, or regulated health information—there is a growing need for secure execution environments that preserve both performance and confidentiality.
What Happened: A Controlled Performance Evaluation
In a controlled study, the NVIDIA performance engineering team evaluated the impact of enabling Confidential Computing on production-grade AI inference using the DeepSeek-R1 model on eight NVIDIA B200 GPUs. The test compared two identical configurations: one with confidential computing disabled (CC off) and one enabled (CC on), while keeping model parameters, sequence lengths, and concurrency levels constant.
The results show that confidential computing retained 96.1% to 98.2% of baseline output-token throughput across concurrency levels from 1 to 16. Meanwhile, the mean time per output token (TPOT) increased by only 1.2% to 4.3%—a small overhead that remains within acceptable bounds for most real-world applications.
Key Facts from the Study
- Workload configuration: 32K input, 1K output tokens, low concurrency (1–16 requests).
- Hardware: 8 NVIDIA B200 GPUs in a DGX B200 system.
- Framework: NVIDIA TensorRT LLM with PyTorch backend.
- Performance retention: 96.1%–98.2% of baseline throughput under CC.
- Latency overhead: 1.2%–4.3% increase in per-token latency.
- Key adaptation: TensorRT LLM switched from CUDA events to GPU %globaltimer for kernel autotuning under CC to avoid unstable timing signals.
- Communication: NVLink multicast is unavailable in B200 CC configurations, so frameworks must detect availability and select low-latency communication algorithms.
Background: How Confidential Computing Works in AI Inference
NVIDIA Confidential Computing leverages hardware-enforced encryption to protect data in use. In a confidential virtual machine (CVM), memory is encrypted at the hardware level, and data movement between host and GPU passes through a software-encrypted bounce buffer. This architecture prevents untrusted software or system-level attacks from accessing sensitive data.
Traditional AI inference frameworks like TensorRT LLM assume direct, unencrypted access to GPU memory. In a confidential environment, these assumptions break down. For example:
- Host-to-device transfers: Cannot use pinned memory directly due to encryption. TensorRT LLM adapts by selecting pageable memory and offloading repeated token readbacks to asynchronous workers.
- Timing measurements: CUDA events are unreliable under CC due to encryption-induced timing instability. TensorRT LLM now uses the GPU’s %globaltimer for kernel autotuning, ensuring accurate performance measurements.
- Multi-GPU communication: NVLink multicast is not available in B200 confidential configurations. Frameworks must detect this limitation and choose communication strategies that minimize latency based on message size and topology.
These adaptations are not optional—they are critical for maintaining performance while preserving security. Without them, the overhead from secure execution could degrade inference throughput significantly.
Why It Matters: Security Without Sacrificing Performance
For enterprises, financial institutions, healthcare providers, and government agencies, AI inference often involves processing sensitive or regulated data. Without a trusted execution environment, models may inadvertently expose private inputs or internal logic.
By enabling secure inference at the hardware level, NVIDIA Confidential Computing allows organizations to deploy AI models in private, isolated environments—without compromising performance. This is especially valuable in regulated sectors where data privacy laws (like GDPR or HIPAA) require strict control over data access and processing.

The study demonstrates that performance degradation is minimal. This makes confidential AI inference viable for production workloads, not just experimental or test environments. It supports a shift toward privacy-preserving AI at scale.
Limitations and Open Questions
While the results are promising, several limitations remain:
- Workload dependency: The performance overhead is most visible in long-context, low-concurrency workloads. In high-concurrency or short-context scenarios, the overhead may be less noticeable—but the study does not validate this across diverse use cases.
- Framework-level adaptability: The performance benefits depend on framework-level optimizations like TensorRT LLM’s CC-aware tuning. Not all inference frameworks have such adaptations, limiting broader adoption.
- Hardware constraints: NVLink multicast is unavailable in B200 CC configurations, which may limit performance in multi-GPU setups. Future iterations may need to address this gap.
- Scalability: The evaluation used a single model (DeepSeek-R1). Performance under different architectures, model sizes, or data types remains unverified.
Additionally, while the study isolates CC overhead, real-world deployments may face additional challenges such as integration with existing orchestration tools, compliance workflows, and monitoring systems.
What to Watch Next
As AI models grow in size and complexity, the demand for secure, private inference will only increase. Key developments to watch include:
- Wider framework support: Expansion of CC-aware optimizations beyond TensorRT LLM to other AI inference frameworks like llama.cpp or vLLM.
- Integration with Kubernetes and cloud platforms: How container orchestration tools will support confidential AI workloads in production environments.
- Real-world deployment cases: Case studies from industries like finance, healthcare, or manufacturing that adopt confidential AI inference for compliance and operational security.
- Performance benchmarks across models: How different LLMs (e.g., Llama 3, Mistral, DeepSeek) perform under CC, especially with varying context lengths and parallelism.
For engineers evaluating secure AI deployment, the NVIDIA Confidential Computing documentation and TensorRT LLM release notes provide essential guidance. Read the original source for technical details and implementation steps.
For those interested in the broader challenge of deploying AI at scale with privacy in mind, explore SWE-Serve, which highlights the gap between local AI testing and real-world serving. Also consider Token Forecaster for insights into managing AI response timing in production.
Ultimately, this work represents a foundational step toward trustworthy AI—where performance and privacy are not mutually exclusive, but rather, can be achieved together.
Sources & further reading
Featured image: The original finding aid described this as: Capture Date: 11/16/1979 Photographer: DONALD HUEBLER Keywords: Larsen Scan Location Building No: 106 by Donald Huebler, Public domain, via Wikimedia Commons. Image source
