PNY NVIDIA L40S 48GB - versatile Ada Lovelace data center accelerator
The L40S covers two workloads at once in a single rack: LLM and generative AI inference plus rendering and graphics visualization. The Ada Lovelace architecture provides 18,176 CUDA cores, 568 4th generation Tensor cores with Transformer Engine and 142 3rd generation RT cores. In FP8 mode the card delivers up to 733 TFLOPS (1,466 with sparsity), and on classic compute - 91.6 TFLOPS FP32.
48 GB of GDDR6 with ECC error correction and a 384-bit bus provide 864 GB/s of bandwidth - enough to keep large models and Omniverse scenes in memory without offloading to disk. Three NVENC units and three NVDEC units with hardware AV1 encoding offload the CPU during parallel video transcoding and 4K/8K streaming.
- NVIDIA Ada Lovelace architecture, 18,176 CUDA cores
- 48 GB GDDR6 ECC, 384-bit bus, 864 GB/s
- 568 4th gen Tensor cores (FP8) + 142 3rd gen RT cores
- PCIe Gen4 x16 interface, TGP 350 W, 16-pin power
- Dual-slot FHFL, passive cooling, 4x DisplayPort 1.4a
- Support for NVIDIA vGPU, Secure Boot, NEBS Level 3
The TCSL40SPCIE-PB version ships in branded PNY packaging with a full manufacturer warranty. This model does not support NVLink - for scaling use multi-GPU over PCIe. An ETE.UA engineer will calculate the server configuration for your stack and the number of cards. Tel. +380 93,594 00 77.
Frequently Asked Questions
What power supply does this card need?
Auxiliary power connector: 1x16-pin. With two of these cards in one system, size the supply with headroom and check the circuit rating.
Will this card fit my case?
The card takes 2 slots. Measure the clearance from the rear panel to the drive cage before ordering: that is where a few millimetres usually go missing.
What model size fits into 48 GB of memory?
Roughly a 30B model at 4-bit or a 13B model at 8-bit. The exact figure depends on context length: longer conversations consume more cache. Memory: 48 GB GDDR6, 384 bit bus.
Which GPU does this card use?
NVIDIA L40S, 18176 compute cores. The GPU generation decides which compute formats are accelerated in hardware, and that affects inference speed more than raw memory capacity does.
What warranty is provided and what is in the box?
Warranty: 36 months. Ships in the manufacturer's original packaging. Ask us to confirm the bundle contents before ordering.
Not sure this card fits your system?
We will check compatibility with your platform, calculate the power draw for your configuration and tell you whether the memory is enough for your models.
The model you need is out of stock? ETE is listed among official PNY partners, so we bring professional NVIDIA cards to order with the manufacturer warranty.
Message us on Telegram or leave a request.
Read the review and measurements: NVIDIA L40S: one card for inference, rendering and VDI — review →