UT Southwestern Events Calendar
View map

BioHPC at UTSW leverages Ollama and OnDemand ChatUI to provide a localized, high-performance environment for large language model (LLM) workloads. While Ollama serves as the backend engine for running models on compute-node resources, the OnDemand ChatUI, built on Open WebUI, offers a streamlined, self-hosted graphical interface tailored for researchers requiring rapid, disposable access to large-scale models.

In this training, we will go over challenges related to the availability of powerful GPU nodes and resource-constrained environments. As model and context sizes continue to grow, limited GPU VRAM creates a bottleneck often referred to as the “memory wall,” leading to deployment and serving challenges. At BioHPC, we address this through optimization strategies at both the hardware and algorithmic levels to enable efficient serving of local LLMs:

(i) NVIDIA Unified Memory (UM) architecture facilitates a hardware-abstracted safety net for memory oversubscription by offloading weight management to the CUDA driver’s demand-paging logic, resulting in faster (up to 1.8× with GPU A100) LLM workloads.

(ii) Google’s TurboQuantization, a novel algorithmic approach for high-ratio, low-bitwidth KV-cache compression, enables up to 6x reduction in active memory footprint (for Llama 3: 70B).

Integrating these approaches allows high-parameter model serving where hardware manages static weight spillover while the algorithm minimizes dynamic activation volume. Overall, these optimizations enable users to deploy, interact with, and serve massive LLMs with negligible loss in precision.

Instructor: Dr. Ramcharan Chandrashekar

Time and Location: Wednesday, April 22, 10:30 AM, G9.102A or hybrid

Registration: not necessary, feel free to walk in

Microsoft Teams: Meeting link

Meeting ID: 223 293 098 775 17

Passcode: oT9XV35K

Event Details

See Who Is Interested

0 people are interested in this event


 

Microsoft Teams: Meeting link

Meeting ID: 223 293 098 775 17

Passcode: oT9XV35K