FZGPUModules 2.0
GPU-accelerated modular compression pipelines
Loading...
Searching...
No Matches
nvrtc_jit.h File Reference

Shared NVRTC compile/cache used by every runtime-generated fusion kernel. More...

#include <string>

Go to the source code of this file.

Namespaces

namespace  fz
 

Functions

bool fz::fused::nvrtcAvailable ()
 True if NVRTC + the CUDA driver are usable in this process.
 
void * fz::fused::nvrtcGetKernel (const std::string &src, const char *entry)
 

Detailed Description

Shared NVRTC compile/cache used by every runtime-generated fusion kernel.

Both fusion strategies generate CUDA source at runtime and JIT it: the chunk- cooperative path (one entry point) and the warp-register path (two entries in one module). This is the common machinery — compile a source to a CUBIN for the device's real SM (no driver-side PTX JIT), load it, and hand back a named device function, caching the compiled module by (arch, source) so only the first compile of a given kernel pays the cost. See nvrtc_chunk_fusion.cpp / nvrtc_warp_fusion.cpp.

Function Documentation

◆ nvrtcGetKernel()

void * fz::fused::nvrtcGetKernel ( const std::string &  src,
const char *  entry 
)

Compile src (cached by device arch + source text) and return the device function named entry from the resulting module. Several entries in the same source share one compiled module (compiled once, looked up per entry). Returned as void* so the header stays free of the CUDA driver headers; callers that launch cast it back to CUfunction. Throws std::runtime_error on failure.