1 <!--
  2 Copyright (c) 2026, Oracle and/or its affiliates. All rights reserved.
  3 DO NOT ALTER OR REMOVE COPYRIGHT NOTICES OR THIS FILE HEADER.
  4 
  5 This code is free software; you can redistribute it and/or modify it
  6 under the terms of the GNU General Public License version 2 only, as
  7 published by the Free Software Foundation.  Oracle designates this
  8 particular file as subject to the "Classpath" exception as provided
  9 by Oracle in the LICENSE file that accompanied this code.
 10 
 11 This code is distributed in the hope that it will be useful, but WITHOUT
 12 ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or
 13 FITNESS FOR A PARTICULAR PURPOSE.  See the GNU General Public License
 14 version 2 for more details (a copy is included in the LICENSE file that
 15 accompanied this code).
 16 
 17 You should have received a copy of the GNU General Public License version
 18 2 along with this work; if not, write to the Free Software Foundation,
 19 Inc., 51 Franklin St, Fifth Floor, Boston, MA 02110-1301 USA.
 20 
 21 Please contact Oracle, 500 Oracle Parkway, Redwood Shores, CA 94065 USA
 22 or visit www.oracle.com if you need additional information or have any
 23 questions.
 24 -->
 25 
 26 # Building HAT (NVidia Notes)
 27 [Back to Index ../](../index.md)
 28 
 29 This guide provides step-by-step instructions to configure Babylon/HAT for CUDA-based GPU acceleration.
 30 
 31 ### Prerequisites
 32 
 33 To run Babylon/HAT with the CUDA backend, ensure the following components are properly installed:
 34 
 35 1. **The NVIDIA GPU Driver:** Download and install the latest appropriate driver for your operating system from the [NVIDIA Drivers page](https://www.nvidia.com/en-us/drivers/)
 36 2. **The CUDA Toolkit (SDK):** Download the matching CUDA Toolkit from the [NVIDIA CUDA Downloads page](https://developer.nvidia.com/cuda-downloads).
 37 
 38 The CUDA Toolkit version must be compatible with your installed NVIDIA driver.
 39 For example, as of June 2025, the stable NVIDIA driver for Linux is `570.169`, which supports CUDA Toolkit version `12.8`.
 40 
 41 Always verify compatibility before installation to prevent runtime errors:
 42 - You can find previous CUDA Toolkit versions on the [NVIDIA CUDA Toolkit Archive](https://developer.nvidia.com/cuda-toolkit-archive)
 43 - Review supported CUDA versions and PTX ISA implementations in the [NVIDIA Parallel Thread Execution documentation](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#release-notes).
 44 
 45 
 46 ### Troubleshooting: Unsupported CUDA/PTX Versions
 47 
 48 If Babylon/HAT runs with an incompatible CUDA/PTX version, you may encounter an error similar to:
 49 
 50 ```bash
 51 cuModuleLoadDataEx CUDA error = 222 CUDA_ERROR_UNSUPPORTED_PTX_VERSION
 52       /<path/to/babylon>/hat/backends/ffi/cuda/src/main/native/cpp/cuda_backend.cpp line 220
 53 ```
 54 
 55 If this is your case, you can:
 56   - **update your GPU driver**.
 57   - **or downgrade your CUDA version.**
 58 
 59 ### Building HAT for the CUDA backend
 60 
 61 If the NVIDIA driver and the CUDA Toolkit SDK are installed, the HAT build process will automatically compile
 62 all sources to dispatch with the CUDA backend.
 63 
 64 ```bash
 65 mvn clean package
 66 ```
 67 
 68 ### Run HAT examples with the CUDA Backend
 69 You can enable the CUDA backend by using the `ffi-cuda` option from HAT.
 70 For example, to run the Matrix Multiplication example:
 71 
 72 ```bash
 73 java @.ffi-cuda-example matmul.Main --kernel=1D
 74 ```
 75 
 76 ### CUDA source compiler selection
 77 
 78 The CUDA backend compiles generated CUDA source with `nvcc` by default. To
 79 compile with NVRTC instead, set `HAT_CUDA_COMPILER=nvrtc`:
 80 
 81 ```bash
 82 HAT_CUDA_COMPILER=nvrtc java @.ffi-cuda-example matmul.Main --kernel=1D
 83 ```
 84 
 85 `HAT_CUDA_COMPILER` accepts `nvcc` or `nvrtc`. Any other value is a
 86 configuration error.
 87 
 88 - **`nvcc`:** compile with the `nvcc` executable. SIMT kernels are emitted as
 89   PTX, and Tile kernels as cubin.
 90 - **`nvrtc`:** compile in-process with NVRTC. SIMT kernels are emitted as PTX,
 91   and Tile kernels as cuda_tile IR.
 92 
 93 SIMT kernels require no extra setup. Tile kernels compiled with NVRTC do. The
 94 NVRTC library registers signal handlers while compiling and executing Tile
 95 kernels. Those handlers receive the signals first and forward them with a
 96 polluted signal state, which breaks JVM signal handling. Preload JDK `libjsig`
 97 so the JVM signal handlers process each signal first and then dispatch to the
 98 NVRTC handlers via chained handlers:
 99 
100 ```bash
101 LD_PRELOAD=$JAVA_HOME/lib/libjsig.so HAT_CUDA_COMPILER=nvrtc \
102   java @.ffi-cuda-test hat.test.TestTileAPI
103 ```
104 
105 If `libjsig` is not preloaded, HAT warns once and compiles Tile kernels with
106 `nvcc`.
107 
108 When NVRTC is selected, HAT loads `libnvrtc.so` from the CUDA Toolkit library
109 directory recorded at build time, then from the default library search path. To
110 use a specific library, set `HAT_CUDA_NVRTC_LIBRARY` to its path or
111 loader-visible name. If that library cannot be loaded, HAT exits:
112 
113 ```bash
114 HAT_CUDA_COMPILER=nvrtc HAT_CUDA_NVRTC_LIBRARY=/path/to/libnvrtc.so java \
115   @.ffi-cuda-example matmul.Main --kernel=1D
116 ```