ast
無料Manual AST construction with FunctionBuilder for kernels and callables without DSL sugar.
日本語の概要は準備中です。原文の説明を表示しています。
Optimize LuisaCompute DSL kernels using warp/wave primitives, shared-memory aggregation, and block-level collectives. Use when kernels bottleneck on atomics, reductions, or inter-thread communication.
インストールする前に、エージェントに与えられる指示の中身を確認できます。
Zero-initialization note: All temporary local variables in DSL kernels — scalars, vectors, matrices, structs, and arrays (excluding shared arrays
Shared<T>) — are created with a zero value automatically. There is no need to manually set them to zero before use. This applies to variables declared withVar<T>,auto, or type-inferred syntax inside a kernel or callable body.
LuisaCompute exposes the following warp-level (subgroup) intrinsics via luisa/dsl/builtin.h. All operate on active lanes within the current warp.
| DSL Call | Returns | Description |
|---|---|---|
warp_lane_count() | UInt | Total lanes in the warp (e.g. 32 or 64). |
warp_lane_id() | UInt | Current lane index [0, warp_lane_count()). |
warp_is_first_active_lane() | Bool | True if this lane is the first active lane in the warp. |
warp_first_active_lane() | UInt | Lane index of the first active lane. |
device.compute_warp_size() | uint | Host-side query for backend's native warp size. |
Each lane receives the same reduced value.
| DSL Call | Signature | Semantics |
|---|---|---|
warp_active_sum(v) | T -> T | Sum across active lanes. |
warp_active_product(v) | T -> T | Product across active lanes. |
warp_active_min(v) | T -> T | Minimum across active lanes. |
warp_active_max(v) | T -> T | Maximum across active lanes. |
warp_active_all(v) | Bool -> Bool | Logical AND across active lanes. |
warp_active_any(v) | Bool -> Bool | Logical OR across active lanes. |
warp_active_all_equal(v) | T -> Bool (or Vec<Bool,N>) | True if all active lanes have the same value. |
warp_active_bit_and(v) | Int -> Int | Bitwise AND across active lanes. |
warp_active_bit_or(v) | Int -> Int | Bitwise OR across active lanes. |
warp_active_bit_xor(v) | Int -> Int | Bitwise XOR across active lanes. |
warp_active_count_bits(v) | Bool -> UInt | Population count of true predicates. |
warp_active_bit_mask(v) | Bool -> UInt4 | 128-bit ballot mask of true predicates. |
All accept scalar or vector types (except bit ops which require integral types).
Each lane receives the exclusive prefix of all preceding active lanes. Lane 0 receives identity (0 for sum, 1 for product, 0u for count_bits).
| DSL Call | Signature | Description |
|---|---|---|
warp_prefix_sum(v) | T -> T | Exclusive prefix sum (T: arithmetic). |
warp_prefix_product(v) | T -> T | Exclusive prefix product (T: arithmetic). |
warp_prefix_count_bits(v) | Bool -> UInt | Exclusive prefix popcount of true predicates. |
| DSL Call | Signature | Description |
|---|---|---|
warp_read_lane(v, lane_idx) | (T, UInt) -> T | Read v from lane lane_idx. T can be scalar, vector, or matrix. |
warp_read_first_active_lane(v) | T -> T | Read v from the first active lane. |
| DSL Call | Description |
|---|---|
set_warp_size(uint8_t) | Must be a power-of-two in [8,128]. Call inside kernel lambda before compilation. |
sync_block() | Full block barrier. All threads in the block must reach it. |
Critical rule: Warp operations only communicate within the same warp. No sync_block() is needed for warp collectives — they are guaranteed to complete within the warp without barriers.
Pattern: each warp computes one output tile via warp_active_sum reduction over the K dimension.
auto warp_size = device.compute_warp_size();
Kernel2D mat_mul = [&](BufferFloat lhs, BufferFloat rhs, BufferFloat result, UInt lhs_row_size) {
set_block_size(128, 1, 1);
set_warp_size(warp_size);
UInt lhs_y = dispatch_id().x / warp_size;
UInt rhs_x = dispatch_id().y;
UInt warp_lane = warp_lane_id();
UInt tile_count = (lhs_row_size + warp_size - 1) / warp_size;
Float acc = 0.f;
for (auto t : dynamic_range(tile_count)) {
UInt lhs_x = t * warp_size + warp_lane;
Float local_v;
$if (lhs_x < lhs_row_size) {
local_v = lhs.read(lhs_y * lhs_row_size + lhs_x)
* rhs.read(rhs_x * lhs_row_size + lhs_x);
} $else {
local_v = 0.f;
};
acc += warp_active_sum(local_v); // all-reduce sum within warp
}
$if (warp_lane == 0) {
result.write(rhs_x * /*M*/ ... + lhs_y, acc);
};
};
Key insight: Warp-active reductions eliminate the need for shared memory entirely. All lanes in a warp already execute in lockstep, so warp_active_sum is a single hardware instruction on most GPUs.
warp_read_laneFor finding the maximum across a logical group smaller than the warp, use pairwise warp_read_lane with XOR lane masks (butterfly / tree-reduction pattern):
// 8-lane max reduction within a group of 8
Float m = input;
m = max(m, warp_read_lane(m, lane ^ 4u)); // distance 4
m = max(m, warp_read_lane(m, lane ^ 2u)); // distance 2
m = max(m, warp_read_lane(m, lane ^ 1u)); // distance 1
// Now lane 0..7 all have the max of lanes 0..7 (modulo diverged lanes)
This is used for the softmax normalization constant. Each logical group of 8 lanes computes its own max independently, without a barrier.
When a warp contains multiple independent logical groups, compute the inclusive prefix sum per group, then use warp_read_lane to fetch the last element of the previous group:
constexpr uint kWarpSize = 32u;
constexpr uint kGroupLanes = 8u;
constexpr uint kGroupsPerWarp = kWarpSize / kGroupLanes; // 4
auto lane = warp_lane_id();
auto group_id = lane / kGroupLanes; // which group (0..3)
auto group_lane = lane % kGroupLanes; // position within group (0..7)
auto prefix = warp_prefix_sum(value); // exclusive prefix across whole warp
auto inclusive = prefix + value;
// Last lane of this group
auto last_lane = group_id * kGroupLanes + (kGroupLanes - 1u);
auto incl_last = warp_read_lane(inclusive, last_lane);
// Last lane of previous group (or 0 for group 0)
auto prev_last = ite(group_id == 0u, 0u, last_lane - kGroupLanes);
auto prev_incl = warp_read_lane(inclusive, prev_last);
auto group_sum = incl_last - ite(group_id == 0u, make_float2(0.f), prev_incl);
This avoids separate warp_prefix_sum calls per group and instead uses a single warp-wide prefix plus lane reads to extract group boundaries.
For inter-block scan, tiles publish their status and other tiles poll via warp collectives:
// Poll across warp: any lane sees INVALID?
$while (warp_active_any(status == SCAN_TILE_INVALID)) {
delay();
status = tile_status.volatile_read(predecessor_idx);
};
// All lanes agree predecessor is inclusive?
$while (warp_active_all(predecessor_status != SCAN_TILE_INCLUSIVE)) {
predecessor_idx -= 32;
// poll next window...
exclusive = scan_op(window_aggregate, exclusive);
};
Also uses warp_active_bit_mask for segmented reductions within warps:
UInt warp_flags = warp_active_bit_mask(flag == 1u).x;
warp_flags >>= 1; // for HEAD_SEGMENT mode
warp_flags &= get_lane_mask_ge(); // mask of lanes with id >= mine
warp_flags |= 1u << (LOGIC_WARP_SIZE - 1u); // sentinel
UInt last_lane = ctz(warp_flags); // first set bit = end of my segment
A software implementation of warp reduce using warp_read_lane with increasing offsets:
Var<T> result = input;
UInt offset = 1u;
$while (offset < warp_lane_count()) {
Var<T> temp = warp_read_lane(result, lane_id + offset, valid_item);
$if (lane_id + offset <= valid_item) {
result = reduce_op(result, temp);
};
offset <<= 1;
};
This is a fallback pattern; prefer warp_active_sum / warp_active_min / warp_active_max when the operation matches the built-in.
Warp-level GEMM where each warp computes one output tile. Threads cooperatively load quantized weights via warp_read_lane to assemble dequantized values, then accumulate with warp_active_sum:
auto warp_lane = warp_lane_id();
UInt tile_count = (K + warp_size - 1) / warp_size;
for (auto t : dynamic_range(tile_count)) {
UInt tile_begin = t * warp_size;
UInt tile_size = min(warp_size, K - tile_begin);
// Each lane loads one quantized word, then shares via warp_read_lane
UInt rel_byte = warp_lane * kElementByteSize;
UInt word = warp_read_lane(warp_word, rel_byte / 4u);
// ... dequantize and multiply ...
acc += warp_active_sum(local_v);
}
Before: Block-level atomic on shared memory.
Shared<int> shared{1u};
shared[0u] = 0;
sync_block();
shared.atomic(0u).fetch_add(1); // contended within block
sync_block();
$if (thread_x() == 0u) {
global_counter.atomic(0u).fetch_add(shared.read(0u));
};
After: Use warp_active_sum per warp, then one lane writes.
// Each warp computes its own partial sum
Int warp_partial = warp_active_sum(1);
// First lane of each warp writes to shared
$if (warp_is_first_active_lane()) {
shared.atomic(warp_lane_id() / warp_size).fetch_add(warp_partial);
};
sync_block();
// One thread (e.g. thread 0) combines across warps
$if (thread_x() == 0u) {
Int block_total = 0;
for (auto w : range(num_warps_per_block)) {
block_total += shared.read(w);
};
global_counter.atomic(0u).fetch_add(block_total);
};
Before: Entire block reduces into shared memory.
Shared<float> smem{block_size};
smem[tid] = value;
sync_block();
for (uint stride = block_size / 2; stride > 0; stride >>= 1) {
$if (tid < stride) {
smem[tid] += smem[tid + stride];
};
sync_block();
};
After: Warp-level reduction + cross-warp shared reduction.
// Phase 1: warp-level reduction (no shared memory, no barrier)
Float warp_sum = warp_active_sum(value);
// Phase 2: cross-warp reduction in shared memory (much smaller)
$if (warp_is_first_active_lane()) {
Shared<float> warp_results{num_warps_per_block};
warp_results[warp_id()] = warp_sum;
};
sync_block();
// Only num_warps_per_block threads participate in phase 2
$if (thread_x() < num_warps_per_block) {
Float v = warp_results[thread_x()];
Float warp_partial = warp_active_sum(v); // second warp reduce
$if (warp_is_first_active_lane()) {
result = warp_partial;
};
};
Before: Butterfly pattern with warp_read_lane.
Float m = value;
m = max(m, warp_read_lane(m, lane ^ 4u));
m = max(m, warp_read_lane(m, lane ^ 2u));
m = max(m, warp_read_lane(m, lane ^ 1u));
After: Single warp_active_max when reduction spans the whole warp.
Float m = warp_active_max(value);
When to keep butterfly: Only when reducing over a logical group smaller than the warp (e.g. 8-lane groups inside a 32-lane warp). In that case warp_active_max would reduce over all 32 lanes, which is incorrect.
warp_prefix_sumBefore: Manually accumulating values from lower lanes via a loop of warp_read_lane.
Float prefix = 0.f;
for (uint i = 0; i < warp_lane_id(); i++) {
prefix += warp_read_lane(value, i);
};
After: Single warp_prefix_sum call.
Float prefix = warp_prefix_sum(value);
When only a subset of lanes should participate in a warp collective, wrap the call in a conditional. Lanes that don't execute the call are excluded from the reduction/scan:
$if (thread_x() % 2u == 0u) {
// Only even-index threads contribute
auto result = warp_prefix_sum(make_half4(.5_h));
device_log("{} -> {}", dispatch_x(), result);
};
// Odd threads don't participate; they receive no result.
This pattern is commonly used for partial-warp scans where only a subset of lanes needs results.
Use warp_active_bit_mask to build a lane mask, then ctz / popcount to locate lanes or count participants:
UInt4 mask = warp_active_bit_mask(condition);
UInt flag_mask = mask.x; // first 32 lanes
UInt first_true = ctz(flag_mask);
UInt num_true = popcount(flag_mask);
Warp collectives (section 3) only communicate within one warp. When cooperation must span the whole thread block (multiple warps), or you need persistent per-block scratch, arbitrary cross-thread indexing, or block-local privatization of a global atomic, use a shared array (Shared<T> / $shared<T>). Shared memory is on-chip and orders of magnitude faster than global memory, so staging data there once and reusing it, or aggregating locally before touching global memory, is a core optimization.
include/luisa/dsl/shared.h, include/luisa/dsl/sugar.h:105)| DSL | Description |
|---|---|
Shared<T> s{n} / $shared<T> s{n} | Allocate n elements of T in workgroup memory. Must be constructed inside the kernel/callable body (uses FunctionBuilder::current()). |
s[i] | Reference access (read or write); i must be an integral expr. |
s.read(i) / s.write(i, v) | Explicit read / write helpers (alias for s[i]). |
s.atomic(i).fetch_add(v) / .compare_exchange(e, v) / ... | Atomic ops on a shared slot. Available for scalar/vector element types (disabled for custom structs). |
s.size() | Element count. |
new Shared<T>{n} | Heap-allocate so helper classes can own shared scratch (Shared<T> is move-only, non-copyable). Lifetime is tied to the enclosing kernel's function builder. See WarpReduce in test_decoupled_look_back.cpp. |
Always set_block_size(...) and size the array to the block (Shared<T> s{block_size}). Use sync_block() to make writes visible across warps.
| Situation | Use |
|---|---|
| Reduction/scan fits in a single warp | Warp collective (section 3) — no barrier, single instruction. |
| Reduction spans a whole block (block_size > warp_size) | Two-level: warp collective → shared → block (section 4.5), or full shared-memory tree reduction (4.4). |
| Many threads append to one global counter/queue | Block-local privatization in shared, then one global atomic per block (4.3). |
| Global data reused by many threads in a block | Stage global → shared once, sync_block(), then reuse (4.6). |
| Arbitrary cross-thread indexing (not just lane shuffles) | Shared array indexed by thread_id(). |
The biggest shared-memory win: replace up to block_size contended global atomics with per-thread shared atomics plus a single global atomic per block. Pattern from test_atomic_queue.cpp (push_if) and test_shared_memory.cpp (AtomicQueue::push):
// Append `value` to a global queue when `pred` holds, minimizing global contention.
Shared<uint> index{1};
$if (thread_x() == 0u) { index.write(0u, 0u); }; // init block counter
sync_block();
auto local_index = def(0u);
$if (pred) { local_index = index.atomic(0).fetch_add(1u); }; // cheap SHARED atomic
sync_block();
$if (thread_x() == 0u) { // ONE global atomic for the whole block
auto local_count = index.read(0u);
auto global_offset = _counter->atomic(0u).fetch_add(local_count);
index.write(0u, global_offset); // reuse slot to broadcast the base
};
sync_block();
$if (pred) { // scatter to reserved, contiguous range
auto global_index = index.read(0u) + local_index;
_buffer->write(global_index, value);
};
Insight: Global-atomic traffic drops from O(active threads) to O(1) per block. Contention moves from device-wide global memory to fast on-chip shared memory. This is the standard stream-compaction / queue-append optimization.
When the reduction spans the whole block, stage each thread's value in shared memory and reduce pairwise with a halving loop. Pattern from test_softmax.cpp (block sum for softmax) and test_complex_kernel.cpp:
set_block_size(block_size, 1, 1); // power of two
Shared<float> shared_arr(block_size); // one slot per thread
auto tid = thread_id().x;
shared_arr[tid] = value; // stage per-thread value
UInt half = block_size / 2u;
sync_block();
$while (half > 0u) {
$if (tid < half) { // compute into a register FIRST
value = shared_arr[tid * 2] + shared_arr[tid * 2 + 1];
};
sync_block(); // barrier between read and write-back
$if (tid < half) {
shared_arr[tid] = value; // write reduced value back
};
half /= 2u;
sync_block(); // barrier before next iteration's reads
};
$if (tid == 0u) { output.write(block_id().x, shared_arr[0]); }; // thread 0 emits block result
Why two sync_block() per step: reducing into a local value register and only writing back after a barrier avoids the read-after-write / write-after-read hazard where one thread overwrites a slot another thread is still reading. Prefer this whole-block form only when block_size > warp_size; inside a single warp, warp_active_sum (section 3.2) is faster and barrier-free.
Combine both tools: reduce within each warp with a warp collective (no barrier), write one partial per warp to a small shared array, then reduce those partials. This minimizes both shared traffic and barriers vs. a full block tree reduction (see also sections 3.1/3.2):
UInt warp_id = thread_x() / warp_size; // which warp in the block (warp_size = device.compute_warp_size())
Float warp_sum = warp_active_sum(value); // phase 1: intra-warp, no barrier
Shared<float> warp_results{num_warps_per_block};
$if (warp_is_first_active_lane()) {
warp_results[warp_id] = warp_sum; // one write per warp
};
sync_block();
$if (thread_x() < num_warps_per_block) { // phase 2: reduce the few partials
Float block_sum = warp_active_sum(warp_results[thread_x()]);
$if (warp_is_first_active_lane()) { /* thread 0 has the block total */ };
};
Load global data into shared once, then reuse it many times or exchange it between threads, avoiding repeated global reads. Patterns from test_shared_mem.cpp, test_async_copy.cpp, and hierarchical mip reduction in test_mipmap.cpp:
set_block_size(N, 1u, 1u);
Shared<uint> s_src{N};
auto tid = thread_x();
s_src[tid] = src_buf.read(dispatch_x()); // global -> shared, once
sync_block(); // publish to the whole block
// ... now reuse s_src[...] / read neighbors' values without touching global memory ...
test_async_copy.cpp fills a shared staging buffer with async_copy(...) (thread 0 issues the copy, then sync_block() before consumers read). test_mipmap.cpp writes 2×2 block averages into Shared<float3> and reduces level-by-level with a sync_block() between levels.
Shared<T> needs FunctionBuilder::current(); declaring it outside a kernel/callable is invalid.sync_block() is required (a) after initializing/filling shared before other threads read, and (b) between the read and write-back phases of each reduction step. Unlike warp collectives, shared memory is NOT self-synchronizing across warps.set_block_size(...); use a power-of-two block for the halving tree reduction, and pad out-of-range lanes with the reduction identity (e.g. 0.f for sum) — see the $if (id < size) {...} $else { value = 0.f; } guards in test_softmax.cpp.Var/register and write back only after a barrier to avoid RAW/WAR hazards.Shared<T> cannot be copied; store Shared<T> * (via new) when a helper class must hold shared scratch.| GPU Backend | Warp/Lane Terminology | Native Width |
|---|---|---|
| CUDA | Warp (32 lanes) | 32 |
| HIP | Wavefront (32/64 lanes) | 32 or 64 |
| Vulkan | Subgroup | Varies (usually 32 or 64) |
| DirectX | Wave | 32 or 64 |
| Metal | SIMD group | 32 |
Always query device.compute_warp_size() on the host and warp_lane_count() on the device rather than hardcoding 32.
Prefer warp collectives over shared memory. warp_active_sum, warp_active_max, warp_prefix_sum compile to single hardware instructions (e.g. __shfl_xor_sync on CUDA, OpGroupNonUniformFAdd on SPIR-V). No barrier needed.
Set warp size explicitly when using warp collectives: set_warp_size(device.compute_warp_size()) inside the kernel lambda.
Don't mix warp and block assumptions. warp_active_sum only reduces within the current warp. If you have multiple warps per block, use a two-level reduction (warp → shared → block).
sync_block() is NOT needed between warp collectives within the same warp. Warp ops are guaranteed complete for the calling lane immediately.
Divergence matters. Lanes that don't execute the warp collective call are excluded. Use this for conditional participation (section 3.5).
warp_prefix_sum is exclusive (not inclusive). Lane 0 always gets 0 (for sum) or 1 (for product).
Vector types work. All warp collectives accept float2, float3, float4, int2, etc. The operation applies component-wise.
Logic warp size. You can logically group lanes (e.g. 4 groups of 8 within a 32-lane warp) using lane % kGroupLanes and lane / kGroupLanes arithmetic. Use warp_read_lane to communicate across groups.
Use shared memory for block-wide cooperation (section 4). When cooperation exceeds one warp, or you need privatization/staging/arbitrary cross-thread indexing, Shared<T> beats warp collectives. Always set_block_size and size the array to the block.
Privatize global atomics into shared memory (section 4.3). Aggregate per-thread contributions with cheap shared atomics, then issue one global atomic per block. This is the key stream-compaction / queue-append win.
Shared memory needs sync_block(); warp collectives do not. Barrier after filling shared and between the read/write-back phases of a tree reduction (section 4.4). Reduce into a register first, then write back after the barrier to avoid RAW/WAR hazards.
GPU kernel optimization (sections 1–6) focuses on what happens inside a single kernel dispatch. Equally important is how you submit work from the host: every stream << command call adds driver overhead. When a per-frame or per-iteration hot loop issues many small stream submissions (upload, dispatch A, dispatch B, download, ...), the accumulated driver cost can become a bottleneck, especially on D3D12 and Vulkan where command submission is not free.
CommandList lets you batch multiple commands into a single submission. Commands are recorded into a CommandList object, then committed to the stream in one shot:
CommandList cmdlist = CommandList::create();
cmdlist << upload_command
<< dispatch_a
<< dispatch_b
<< download_command;
stream << cmdlist.commit() << synchronize();
All commands execute in FIFO order on the GPU, exactly as if they were submitted individually — but with a single driver round-trip instead of many.
CommandList cannot span multiple streams.stream << statements inside a loop or per-frame function.CommandList.CommandList::create() once at the start of the group, then append commands with <<.stream << cmdlist.commit() submits the batch; follow with a single synchronize() if host-readback is needed.Batching N separate stream << cmd submissions into one CommandList reduces:
commit() + synchronize() replaces N pairs of stream << ... << synchronize().In practice, replacing 5+ stream submissions per iteration with a single CommandList::create() → commit() can yield measurable wall-clock speedups in offline rendering or training-data export loops, where the CPU-side submission overhead is a meaningful fraction of the iteration time.
| Optimization | Scope | Impact |
|---|---|---|
| Warp collectives (section 3) | GPU kernel — replaces shared memory and atomics | Reduces latency/contention within a warp |
| Shared memory privatization (section 4.3) | GPU kernel — aggregates block atomics | Reduces global atomic contention from O(block_size) to O(1) per block |
| CommandList batching (this section) | Host submission — batches stream commands | Reduces driver overhead from O(N) to O(1) per iteration |
CommandList batching is orthogonal to kernel-level optimizations. Apply both: optimize the kernel with warp/shared-memory techniques, then batch the host submissions for maximum throughput.
CommandList for a group of dependent commands, commit it once, and synchronize once rather than submitting each command separately.commit() the list is consumed; create a fresh one for the next batch.stream <<. Batched submission is more efficient than long chains of stream << a << b << c << synchronize() because it reduces internal queue flushes.synchronize() with Non-Blocking CompletionCommandList provides two callback hooks that decouple host work from GPU execution:
// Runs AFTER all GPU commands in this CommandList finish.
cmdlist.add_callback([](auto &&... captured) noexcept {
// GPU work is done — safe to read back buffers, write files, etc.
});
// Runs when the CommandList is committed/destructed, BEFORE GPU starts.
cmdlist.add_dtor_callback([](auto &&... captured) noexcept {
// Host-side cleanup of resources no longer needed.
});
| Callback | When it fires | GPU status | Typical use |
|---|---|---|---|
add_dtor_callback | At commit() (or when Commit object is destroyed) | Not started yet | Release temporary host buffers, close files, or free staging memory that was only needed to construct the commands. |
add_callback | After all GPU commands in the list have completed | Done | Read back downloaded buffers, write output files, signal host work queues, or launch dependent host tasks. |
Critical difference: add_dtor_callback runs before GPU execution begins — it is not a completion callback. Only add_callback guarantees GPU work is finished.
synchronize() StallsWithout callbacks, host→GPU data exchange typically looks like:
// Blocking pattern: host stalls until GPU finishes.
stream << cmdlist.commit() << synchronize(); // host waits here
process_results(host_buffer); // then processes
add_callback lets you flip this into a non-blocking, continuation-passing style:
// Non-blocking pattern: callback processes results when GPU is done.
cmdlist.add_callback([host_buffer = std::move(host_buf)]() noexcept {
process_results(host_buffer); // runs after GPU finishes
});
stream << cmdlist.commit(); // host returns immediately, no stall
// Host can begin preparing the next frame/iteration NOW...
This is most impactful when:
The classic pattern for hiding latency: overlap GPU execution of iteration N+1 with host processing of iteration N's results.
// Host-side buffers must outlive the GPU work.
// Use double-buffering or shared ownership (e.g., shared_ptr).
for (int i = 0; i < num_iterations; i++) {
auto host_buf = std::make_shared<std::vector<float>>(size);
CommandList cmdlist = CommandList::create();
// Record GPU commands (upload, dispatch, download into host_buf)
cmdlist << upload << kernel.dispatch(...) << download(host_buf->data());
// Install the completion callback — captures host_buf by shared_ptr
cmdlist.add_callback([host_buf, i]() noexcept {
// GPU done: safely read host_buf, write to disk, etc.
save_result(*host_buf, i);
});
// Submit without blocking — host returns immediately
stream << cmdlist.commit();
// Host can prepare iteration i+1's work right away...
}
// Final sync: wait for the very last iteration to finish.
stream << synchronize();
With this pattern:
synchronize() per iteration — only one final sync at the very end.synchronize() stall × (N−1) iterations.Lambdas passed to add_callback / add_dtor_callback must own their captured resources because the callback outlives the CommandList object. Use:
// ✅ Shared ownership (recommended for buffers)
auto data = std::make_shared<std::vector<float>>(size);
cmdlist.add_callback([data]() noexcept { /* safe */ });
// ✅ Move semantics for unique resources
auto data = std::make_unique<std::vector<float>>(size);
cmdlist.add_callback([data = std::move(data)]() noexcept { /* safe */ });
// ❌ Capturing raw pointers or references to stack/local variables is UAF.
float *raw = ...;
cmdlist.add_callback([raw]() noexcept { /* DANGER: raw may be dangling */ });
add_dtor_callback has the same ownership rules, despite running earlier — it still fires after commit() returns, so stack variables captured by reference would be invalid.
| Situation | Use |
|---|---|
| Read back GPU results and save/process them | add_callback |
| Free host staging buffers after upload | add_dtor_callback |
| Close files or decrement refcounts after submission | add_dtor_callback |
| Signal a host work queue that GPU output is ready | add_callback |
| Launch the next iteration's host prep work | Just place after commit() on the host (no callback needed) |
add_callback fires after GPU completion — it is the non-blocking replacement for synchronize().add_dtor_callback fires before GPU starts — use only for host-side cleanup, never for reading back GPU results.synchronize() is still needed at the end of a pipeline to ensure the last iteration's callbacks have fired before the program exits.まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Manual AST construction with FunctionBuilder for kernels and callables without DSL sugar.
日本語の概要は準備中です。原文の説明を表示しています。
Backend plugin architecture, DeviceInterface, dynamic loading, and command encoding.
日本語の概要は準備中です。原文の説明を表示しています。
CMake build options, custom functions, and backend patterns for LuisaCompute.
日本語の概要は準備中です。原文の説明を表示しています。
C++ naming, formatting, static analysis, and RTTI rules for LuisaCompute.
日本語の概要は準備中です。原文の説明を表示しています。
Debug crashes and test failures via stack-traces, host/device logging, and DSL buffer inspection.
日本語の概要は準備中です。原文の説明を表示しています。
Show uncommitted changes and commit history via git. Use when: (1) inspecting what changed in a file before committing, (2) checking diff of tracked/new/deleted files, (3) reviewing a specific commit's changes, (4) listing files changed in a commit.
日本語の概要は準備中です。原文の説明を表示しています。