FZGPUModules 2.0
GPU-accelerated modular compression pipelines
Loading...
Searching...
No Matches
atomics.h File Reference

Backend-neutral block-scoped atomics. More...

#include "backend/api.h"

Go to the source code of this file.

Namespaces

namespace  fz
 

Functions

template<typename T >
fz::backend::atomicAddBlock (T *addr, T val)
 
template<typename T >
fz::backend::atomicOrBlock (T *addr, T val)
 
template<typename T >
fz::backend::atomicMaxBlock (T *addr, T val)
 

Detailed Description

Backend-neutral block-scoped atomics.

atomicAdd_block/atomicOr_block (CUDA's block-memory-scope atomic family, as opposed to the default device-scope atomicAdd/atomicOr) are used by the RARE/RAZE device code (modules/coders/lc_common/lc_chunk_components.cuh) for a per-block histogram accumulation and for OR-ing a straddling partial word into shared memory — both cases where every participating thread is known to be in the same block, so scoping the atomic down from "device" saves the extra memory-fence cost of a full device-wide atomic.

The _block suffix family is CUDA-only spelling. Contrary to this header's original expectation ("ROCm declares them with identical names"), ROCm 6.4.1 declares neither atomicAdd_block nor atomicOr_block anywhere under hip/ — verified by grep against the installed toolchain, and by the ‘use of undeclared identifier 'atomicAdd_block’` error the first HIP build of the RARE/RAZE stages produced. This is exactly the divergence the header was created to absorb, so the patch lands here and the call sites are untouched.

HIP's equivalent is the scoped-atomic builtin family (__hip_atomic_fetch_add/_or + __HIP_MEMORY_SCOPE_WORKGROUP), where "workgroup" is AMD's name for a CUDA block. Memory ordering is __ATOMIC_RELAXED to match CUDA's atomic*_block, which carry no ordering guarantees beyond the atomicity of the read-modify-write itself — the RARE/ RAZE call sites (per-block histogram accumulation; OR-ing a straddling partial word into shared memory) separately __syncthreads() before reading the accumulated result, so they rely on that barrier for ordering, not on the atomic.

Unlike the warp shuffle/ballot family (see warp.h), a block-scoped atomic has no wavefront-width dependency, so there is no 32-vs-64-lane subtlety here.

Function Documentation

◆ atomicAddBlock()

template<typename T >
T fz::backend::atomicAddBlock ( T *  addr,
val 
)
inline

Block-scoped atomic add. T must be one of the types CUDA's atomicAdd_block overloads on (int, unsigned, unsigned long long, float, double).

◆ atomicOrBlock()

template<typename T >
T fz::backend::atomicOrBlock ( T *  addr,
val 
)
inline

Block-scoped atomic bitwise-OR. T must be one of the types CUDA's atomicOr_block overloads on (int, unsigned, unsigned long long).

◆ atomicMaxBlock()

template<typename T >
T fz::backend::atomicMaxBlock ( T *  addr,
val 
)
inline

Block-scoped atomic max. T must be one of the types CUDA's atomicMax_block overloads on (int, unsigned, unsigned long long).