A dispatch is a grid of workgroups, each a block of @workgroup_size invocations. local_invocation_id locates an invocation in its workgroup, workgroup_id locates the workgroup, global_invocation_id combines them (workgroup_id * workgroup_size + local_invocation_id), and num_workgroups is the dispatch size. Chrome 154 1 adds the flat global_invocation_index and workgroup_index (the linear_indexing language extension), and the subgroups feature adds subgroup_size and subgroup_invocation_id:
<script type="module">
const adapter = await navigator.gpu.requestAdapter();
const device = await adapter.requestDevice({ requiredFeatures: ['subgroups'] });
const module = device.createShaderModule({ code: /* wgsl */ `
enable subgroups;
@group(0) @binding(0) var<storage, read_write> out: array<vec4u>;
@compute @workgroup_size(4, 2) fn main(@builtin(global_invocation_id) g: vec3u,
@builtin(local_invocation_id) l: vec3u, @builtin(workgroup_id) w: vec3u,
@builtin(local_invocation_index) li: u32, @builtin(global_invocation_index) gi: u32,
@builtin(num_workgroups) n: vec3u, @builtin(subgroup_size) sg: u32) {
out[gi] = vec4u(g.x * 10 + g.y, l.x * 10 + l.y, w.x * 10 + li, n.x * 100 + sg);
}` });
const { STORAGE, COPY_SRC, COPY_DST, MAP_READ } = GPUBufferUsage;
const out = device.createBuffer({ size: 256, usage: STORAGE | COPY_SRC });
const read = device.createBuffer({ size: 256, usage: COPY_DST | MAP_READ });
const pipeline = device.createComputePipeline({ layout: 'auto', compute: { module } });
const encoder = device.createCommandEncoder(), pass = encoder.beginComputePass();
pass.setPipeline(pipeline), pass.setBindGroup(0, device.createBindGroup({
layout: pipeline.getBindGroupLayout(0), entries: [{ binding: 0, resource: out }] }));
pass.dispatchWorkgroups(2), pass.end(); // 2 workgroups of 4 x 2 = 16 invocations
encoder.copyBufferToBuffer(out, 0, read, 0, 256);
device.queue.submit([encoder.finish()]), await read.mapAsync(GPUMapMode.READ);
const v = new Uint32Array(read.getMappedRange());
for (const [k, name] of ['global xy', 'local xy', 'workgroup,local'].entries()) {
console.log(name.padEnd(16) + [...Array(16)].map((_, i) =>
String(v[i * 4 + k]).padStart(2, '0')).join(' '));
}
console.log(`num_workgroups.x ${Math.floor(v[3] / 100)}, subgroup_size ${v[3] % 100}`);
</script>global xy 00 10 20 30 40 50 60 70 01 11 21 31 41 51 61 71 local xy 00 10 20 30 00 10 20 30 01 11 21 31 01 11 21 31 workgroup,local 00 01 02 03 10 11 12 13 04 05 06 07 14 15 16 17 num_workgroups.x 2, subgroup_size 32
Columns follow global_invocation_index across the 8 x 2 grid; local_invocation_index restarts per workgroup, and the GTX 1650 has 32-wide subgroups. Dispatches round up to whole workgroups, so guard the global ID against the data size.