1 <!--
2 Copyright (c) 2026, Oracle and/or its affiliates. All rights reserved.
3 DO NOT ALTER OR REMOVE COPYRIGHT NOTICES OR THIS FILE HEADER.
4
5 This code is free software; you can redistribute it and/or modify it
6 under the terms of the GNU General Public License version 2 only, as
7 published by the Free Software Foundation. Oracle designates this
8 particular file as subject to the "Classpath" exception as provided
9 by Oracle in the LICENSE file that accompanied this code.
10
11 This code is distributed in the hope that it will be useful, but WITHOUT
12 ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or
13 FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License
14 version 2 for more details (a copy is included in the LICENSE file that
15 accompanied this code).
16
17 You should have received a copy of the GNU General Public License version
18 2 along with this work; if not, write to the Free Software Foundation,
19 Inc., 51 Franklin St, Fifth Floor, Boston, MA 02110-1301 USA.
20
21 Please contact Oracle, 500 Oracle Parkway, Redwood Shores, CA 94065 USA
22 or visit www.oracle.com if you need additional information or have any
23 questions.
24 -->
25
26 # Building HAT (NVidia Notes)
27 [Back to Index ../](../index.md)
28
29 This guide provides step-by-step instructions to configure Babylon/HAT for CUDA-based GPU acceleration.
30
31 ### Prerequisites
32
33 To run Babylon/HAT with the CUDA backend, ensure the following components are properly installed:
34
35 1. **The NVIDIA GPU Driver:** Download and install the latest appropriate driver for your operating system from the [NVIDIA Drivers page](https://www.nvidia.com/en-us/drivers/)
36 2. **The CUDA Toolkit (SDK):** Download the matching CUDA Toolkit from the [NVIDIA CUDA Downloads page](https://developer.nvidia.com/cuda-downloads).
37
38 The CUDA Toolkit version must be compatible with your installed NVIDIA driver.
39 For example, as of June 2025, the stable NVIDIA driver for Linux is `570.169`, which supports CUDA Toolkit version `12.8`.
40
41 Always verify compatibility before installation to prevent runtime errors:
42 - You can find previous CUDA Toolkit versions on the [NVIDIA CUDA Toolkit Archive](https://developer.nvidia.com/cuda-toolkit-archive)
43 - Review supported CUDA versions and PTX ISA implementations in the [NVIDIA Parallel Thread Execution documentation](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#release-notes).
44
45
46 ### Troubleshooting: Unsupported CUDA/PTX Versions
47
48 If Babylon/HAT runs with an incompatible CUDA/PTX version, you may encounter an error similar to:
49
50 ```bash
51 cuModuleLoadDataEx CUDA error = 222 CUDA_ERROR_UNSUPPORTED_PTX_VERSION
52 /<path/to/babylon>/hat/backends/ffi/cuda/src/main/native/cpp/cuda_backend.cpp line 220
53 ```
54
55 If this is your case, you can:
56 - **update your GPU driver**.
57 - **or downgrade your CUDA version.**
58
59 ### Building HAT for the CUDA backend
60
61 If the NVIDIA driver and the CUDA Toolkit SDK are installed, the HAT build process will automatically compile
62 all sources to dispatch with the CUDA backend.
63
64 ```bash
65 mvn clean package
66 ```
67
68 ### Run HAT examples with the CUDA Backend
69 You can enable the CUDA backend by using the `ffi-cuda` option from HAT.
70 For example, to run the Matrix Multiplication example:
71
72 ```bash
73 java @.ffi-cuda-example matmul.Main --kernel=1D
74 ```
75
76 ### CUDA source compiler selection
77
78 The CUDA backend compiles generated CUDA source with `nvcc` by default. To
79 compile with NVRTC instead, set `HAT_CUDA_COMPILER=nvrtc`:
80
81 ```bash
82 HAT_CUDA_COMPILER=nvrtc java @.ffi-cuda-example matmul.Main --kernel=1D
83 ```
84
85 `HAT_CUDA_COMPILER` accepts `nvcc` or `nvrtc`. Any other value is a
86 configuration error.
87
88 - **`nvcc`:** compile with the `nvcc` executable. SIMT kernels are emitted as
89 PTX, and Tile kernels as cubin.
90 - **`nvrtc`:** compile in-process with NVRTC. SIMT kernels are emitted as PTX,
91 and Tile kernels as cuda_tile IR.
92
93 SIMT kernels require no extra setup. Tile kernels compiled with NVRTC do. The
94 NVRTC library registers signal handlers while compiling and executing Tile
95 kernels. Those handlers receive the signals first and forward them with a
96 polluted signal state, which breaks JVM signal handling. Preload JDK `libjsig`
97 so the JVM signal handlers process each signal first and then dispatch to the
98 NVRTC handlers via chained handlers:
99
100 ```bash
101 LD_PRELOAD=$JAVA_HOME/lib/libjsig.so HAT_CUDA_COMPILER=nvrtc \
102 java @.ffi-cuda-test hat.test.TestTileAPI
103 ```
104
105 If `libjsig` is not preloaded, HAT warns once and compiles Tile kernels with
106 `nvcc`.
107
108 When NVRTC is selected, HAT loads `libnvrtc.so` from the CUDA Toolkit library
109 directory recorded at build time, then from the default library search path. To
110 use a specific library, set `HAT_CUDA_NVRTC_LIBRARY` to its path or
111 loader-visible name. If that library cannot be loaded, HAT exits:
112
113 ```bash
114 HAT_CUDA_COMPILER=nvrtc HAT_CUDA_NVRTC_LIBRARY=/path/to/libnvrtc.so java \
115 @.ffi-cuda-example matmul.Main --kernel=1D
116 ```