A shader's resource interface is normally written down in three places and held together by hand: once in the shader, once in C++, and once in a reflection table that tries to reconcile the two after the shader has compiled. In SETech, our in-house engine, it is written once. One header is compiled twice, as shader code and as C++, and the engine binds resources on D3D11, D3D12, Vulkan, Metal and WebGPU without ever looking a name up.

This article walks through how that works, why the grouping underneath it is the part that matters, and what it bought us on D3D12. Everything here is from the engine itself. You can see it running in the World Alone demo, which is this same code compiled to WebAssembly.

The usual way: reconcile the two sides at runtime

The traditional flow has three steps. Declare the resource in the shader. Declare it a second time on the C++ side. Then match the two, usually by name, through a reflection table built when the shader compiles.

This is a reasonable trade, and the cost is not performance: the lookup resolves once at load, and the runtime carries integers from then on. The cost is where the check can happen. The table cannot exist until the shader has compiled, so nothing checks the two declarations against each other at the point where you write them. Rename a texture on one side and the C++ compiles, the shader compiles, and the texture is simply never bound. You get an unbound parameter, not a build error.

There is a real reason that cannot simply be turned into an error. An unused shader parameter is legitimately optimized out during shader compilation, so "this name is not in the parameter map" cannot be told apart from "the shader stopped using it". Once binding is by name against a compiled artifact, that ambiguity is unavoidable. Removing the name is what removes the ambiguity.

One pass, three files, one declaration

Screen-space ambient occlusion is a good small example. The pass needs a noise texture, a point sampler and some constants. In SETech the declaration lives in a shared group header, the shader just uses it, and the C++ points at the declaration instead of repeating it.

SEFixedGroup.srg.h, the only declaration
SE_DECL_STATIC_SAMPLER( gSamplerPointWrap )
SE_DECL_TEXTURE_2D(     NoiseTexture )
SSAO.fx, which declares nothing
#include "SEBaseEngine.srl.h"

float3 rand = SE_SAMPLE_2D( NoiseTexture, gSamplerPointWrap, noiseUV ).xyz;
SECommonResources.cpp, which names nothing
aResources[ 5 ].m_pDesc      = SE_RESOURCE_DESC( BaseEngine, Fixed, NoiseTexture );
aResources[ 5 ].m_ppTextures = &m_pNoiseTexture;

m_pFixedResourceGroup->Update( m_uiFixedRingSlot, aResources, 12 );

SE_RESOURCE_DESC is not a string lookup. It expands to a struct member access into the layout that the header itself declared. Rename NoiseTexture in the header and this line stops compiling. The mistake that would silently unbind a texture under name matching does not build here.

One declaration, two compilers

A layout is declared in a .srl.h file (shader resource layout), built out of .srg.h fragments (shader resource groups). The same file is consumed by both toolchains, and the fork between them is one #ifdef __cplusplus in SEShaderMacros.h.

SINGLE SOURCE SESlugText.srl.h SE_DECL_TEXTURE_2D( … ) #ifdef __cplusplus SRL PREPROCESSOR picks slots for the target API Texture2D SlugCurveTexture : register( t1, space2 ); → fxc / dxc C++ COMPILER same tokens, struct members const SEResourceDesc SlugCurveTexture = { …, m_uiOffset++, m_uiSRVCount++ }; both branches walk the declarations in the same order, so both land on the same slot. by construction, not by agreement.
One source, two consumers. Neither branch can drift from the other because there is only one list of declarations to walk.

On the shader side, a small preprocessor walks the declaration list before the HLSL compiler sees it and rewrites each two-argument declaration into one that carries its register and space. The macros then expand to ordinary HLSL: a cbuffer, a Texture2D, a SamplerState. The target backend is a parameter to that preprocessor, which matters later.

On the C++ side, the same macros expand to a struct. Each group becomes a nested struct, and each resource becomes a member whose default initializer bumps a per-kind counter:

What the C++ compiler sees, roughly
struct SELayout_BaseEngine {
    struct _SEGroup_Fixed {
        uint32 m_uiOffset = 0, m_uiCBVCount = 0, m_uiSRVCount = 0, m_uiSamplerCount = 0;
        const SEResourceDesc ShadowMapSampler = { "ShadowMapSampler", eShaderResourceType_SamplerComparison,
                                                  0, 0, m_uiOffset++, m_uiSamplerCount++, 1 };
        const SEResourceDesc gFixed           = { "gFixed", eShaderResourceType_ConstantBuffer,
                                                  0, sizeof( SEFixedConstants ), m_uiOffset++, m_uiCBVCount++, 1 };
    };
    // ...three more groups
};

The layout is a function-local static, so the whole walk runs exactly once, the first time anything asks for it, in declaration order. That is the honest description of the cost: not a compile-time constant, but one pass over a few dozen integer increments per layout, with no strings compared and no compiled shader consulted. The name string in each descriptor exists for debugging and is never read by the binding code. The implementation is deliberately plain: no templates, no constexpr, no standard library, no string processing.

Both sides walk the same list in the same order, so they land on the same slots by construction, not by agreement. There is no shader reflection anywhere in the engine's binding path.

The shared groups are shared on purpose

Almost every layout in the project opens the same way:

How a .srl.h starts
SE_DECL_SHADER_RESOURCE_LAYOUT_BEGIN( SlugText )

    SE_DECL_RESOURCE_GROUP_BEGIN( Fixed )
        #include "SEFixedGroup.srg.h"
    SE_DECL_RESOURCE_GROUP_END( Fixed )

    SE_DECL_RESOURCE_GROUP_BEGIN( PerFrame )
        #include "SEPerFrameGroup.srg.h"
    SE_DECL_RESOURCE_GROUP_END( PerFrame )

    // ...then whatever this shader alone needs

Fixed and PerFrame are never retyped per shader. Nearly every layout in the engine opens with the same two fragments, in the same order, before anything of its own. The few exceptions are self-contained, such as the GPU physics solver, which touches no shared engine state and declares only its own group. So every shader that uses engine state agrees on where the camera, the lights and the samplers sit, and the engine binds those groups once for all of them.

On D3D11 this is structural, not a convention somebody has to remember. Shader Model 5 has no register spaces, so the counters run continuously and never restart per group. The identical prefix is the only reason the numbering lines up across shaders. There is a hard ceiling behind it too: D3D11 gives a shader stage 14 constant buffer slots, and the shared prefix spends a good share of them before a shader declares a single buffer of its own.

D3D11 constant buffer registers for the text shader, a snapshot
b0  gFixed              b3  gGlobalLights
b1  gShadowSettings     b4  gSky
b2  gPerFrame           b5  gShadowPS
b6  gSlugText           // the first register this shader owns

Because that prefix is the same everywhere, the remaining slots are countable. If every shader invented its own prefix, nobody could say how many were left.

Update frequency is the invariant

The five APIs disagree about almost everything: register spaces, how counters run, whether samplers get their own heap. They agree on exactly one thing, and none of them says it out loud. Resources should be grouped by how often they change while a command buffer is being built.

THE ENGINE'S FOUR GROUPS WHAT EACH API CALLS THE SAME IDEA SHARED BY EVERY .srl.h IN THE PROJECT Fixed set once at init PerFrame camera, lights, shadows DECLARED PER FAMILY OF SHADERS PerBatch material and its textures PerDraw world transform one concept, five names D3D11constant-buffer slots, no grouping primitive at all D3D12descriptor tables, one per register space Vulkandescriptor sets Metalargument buffers, one per SPIR-V set WebGPUbind groups WebGPU caps bind groups at four. Vulkan guarantees only four. The engine already had exactly four frequencies. Nothing was changed to fit.
Not a lowest common denominator. It is the grouping every one of these APIs was already doing. The top two groups are shared engine-wide, and the bottom two belong to whichever system declares them.

SETech has four groups: Fixed, PerFrame, PerBatch and PerDraw. A register space on D3D12, a descriptor set on Vulkan, a bind group on WebGPU: three names for the same idea, and the declaration mentions none of them.

One declaration, five outputs

Here is the text renderer's per-batch group, declared once. It holds the entity constants, a texture of Bezier control points and a band acceleration texture.

SESlugText.srl.h
SE_DECL_RESOURCE_GROUP_BEGIN( PerBatch )
    SE_DECL_CONSTANT_BUFFER( SESlugTextConstants, gSlugText )
    SE_DECL_TEXTURE_2D( SlugCurveTexture )   // RGBA32F Bezier control points
    SE_DECL_TEXTURE_2D( SlugBandTexture )    // RGBA8, uint16 pairs per band
SE_DECL_RESOURCE_GROUP_END( PerBatch )

And here is what the backends compile, with the shared PerFrame buffer alongside so the spaces are visible. Watch the two constant buffers.

D3D11: one flat namespace, no spaces exist
cbuffer _cb_gPerFrame : register( b2 )  { ... };
cbuffer _cb_gSlugText : register( b6 )  { ... };
Texture2D SlugCurveTexture : register( t11 );
Texture2D SlugBandTexture  : register( t12 );
D3D12: b, t and u share one counter inside each space
cbuffer _cb_gPerFrame : register( b0, space1 ) { ... };
cbuffer _cb_gSlugText : register( b0, space2 ) { ... };
Texture2D SlugCurveTexture : register( t1, space2 );
Texture2D SlugBandTexture  : register( t2, space2 );
Vulkan: every register bank counts from zero per space
cbuffer _cb_gPerFrame : register( b0, space1 ) { ... };
cbuffer _cb_gSlugText : register( b0, space2 ) { ... };
Texture2D SlugCurveTexture : register( t0, space2 );
Texture2D SlugBandTexture  : register( t1, space2 );
WebGPU: WGSL out of Tint, one bind group per frequency
@group(2u) @binding(0u) var<uniform> _cb_gSlugText : S_1;
@group(2u) @binding(1u) var SlugCurveTexture : texture_2d<f32>;
@group(2u) @binding(2u) var SlugBandTexture  : texture_2d<f32>;

PerFrame and PerBatch are both b0 on D3D12 and Vulkan, and only the space tells them apart. D3D11 has no spaces, so the same two buffers land on two different flat registers, and the textures are pushed up the t range by every shader resource view declared ahead of them in the shared groups. The exact numbers in these blocks are a snapshot and will drift as the shared groups grow. The shape will not. Metal goes through SPIRV-Cross, where the cooker assigns flat buffer, texture and sampler slots from the same per-group numbering. Moving that to one argument buffer per group is the natural next step, and the declarations would not change.

What the frequencies buy on D3D12

A root signature used to be a per-effect object in our D3D12 backend. Every shader effect built its own, and the command buffer bound and switched them as it walked the frame. SetGraphicsRootSignature invalidates every root parameter, so each switch forced a full rebind of every table, whether or not anything in them had changed.

Once every shader in the engine agrees on the same four groups, none of that is needed. There are now two root signatures for all raster and compute work, one graphics and one compute, created once at init and set once when a command buffer begins. No draw ever changes them, so root parameters stay valid for the whole command buffer and only the tables that actually changed get rewritten.

rebinds every draw bound once per frame param 0 PerDraw CBV / SRV / UAV table space3 cheapest root slot param 1 PerBatch CBV / SRV / UAV table space2 param 2 PerFrame CBV / SRV / UAV table space1 param 3 Fixed CBV / SRV / UAV table space0 param 4 Fixed sampler table space0 8 static samplers baked into the signature space5 zero root-param cost
The engine-wide root signature. The root parameter for a group is simply ( groupCount - 1 - registerSpace ), one formula shared by the builder and the per-draw binding code.

Root parameters are ordered by descending update frequency, so the table that rebinds most often sits in the cheapest root slot. That ordering is only possible because the declaration already says how often a resource changes. The compute signature strips the graphics-only deny flags and is otherwise identical. The one exception is ray tracing: a DXR pipeline carries its own global root signature, and the command buffer restores the shared compute signature after DispatchRays.

Resources were only the first thing worth sharing

Once one file is compiled by both compilers, anything declared in the same shape stops being able to drift. Two more things moved in.

Constant buffer bodies

Declared once, in the layout header
struct SESlugTextConstants
{
    float4x4 matWorld;
    float4   SelectionId;
    float4   EngraveParams;
    float4   TextParams;
};

The HLSL side wraps it in a cbuffer. The C++ side compiles the identical body as a plain struct, because the shared header maps the HLSL type names onto the engine's math types:

SEShaderMacros.h, the C++ branch
typedef SEVec4       float4;
typedef SEMat4       float4x4;
typedef unsigned int uint;

So the buffer's size comes from the declaration too. Nobody writes a sizeof that has to match a shader by hand.

Vertex layouts

SEVertexStructs.h
SE_VERTEX_LAYOUT_BEGIN(VSTexturedColorIn)
    SE_VERTEX_ELEMENT(SE_FLOAT4, m_Position,  SE_POSITION)
    SE_VERTEX_ELEMENT(SE_FLOAT2, m_TexCoord0, SE_TEXCOORD0)
    SE_VERTEX_ELEMENT(SE_UINT,   m_Color,     SE_COLOR0)
SE_VERTEX_LAYOUT_END
To HLSL, an input struct
struct VSTexturedColorIn {
    float4 m_Position  : LOCALPOS;
    float2 m_TexCoord0 : TEXCOORD0;
    uint   m_Color     : COLOR0;
};
To C++, a layout descriptor
{ 0, SETECH_DATATYPE_FLOAT, 4, SETECH_DATAUSAGE_POSITION, 0, false },
{ 0, SETECH_DATATYPE_FLOAT, 2, SETECH_DATAUSAGE_TEXCOORD, 0, false },
{ 0, SETECH_DATATYPE_UINT8, 4, SETECH_DATAUSAGE_COLOR,    0, false },

SE_FLOAT4 is float4 to one compiler and a datatype plus a count to the other. SE_POSITION is a semantic on one side and a usage enum on the other. A vertex layout that disagrees with its shader's input struct fails the same way a mis-numbered register does: it builds, it runs, and the geometry is wrong. Same file, same fix.

What did not move in is anything whose body needs real HLSL semantics or intrinsics, which is why the shader helper header is guarded with #if !defined( __cplusplus ). The rule is simple: only declarations whose bodies are expressible in both languages can be shared.

Where this actually runs

Runtime shader compilation is a development convenience, available on the D3D and Vulkan backends. Everything else is cooked ahead of time, per backend, by the engine's offline asset tool: the WebGPU build, for one, can only load cooked shaders. The same layout preprocessor feeds all five targets, with the backend as a parameter, and the in-process compile path calls that same preprocessor.

OFFLINE · SETechAssetTool · THIS IS WHAT SHIPS .srl.h + .srg.h .fx includes flattened SRL preprocessor target is a parameter D3D11 FXC HLSL → bytecode DXBC D3D12 DXC HLSL → DXIL DXIL Vulkan DXC -spirv SPIR-V Metal DXC → SPIR-V patches → SPIRV-Cross MSL → default.metallib WebGPU DXC → SPIR-V patches → Tint WGSL all five land in the .pkg the app mounts IN PROCESS · DEVELOPMENT ITERATION ONLY FXC / DXC, in process on first use of a variant cache/shaders/*.shader local, disposable, never shipped one preprocessor, two callers. the cook decides what ships; the in-process path exists so a shader edit is visible without a cook.
Five targets, five toolchains, one set of declarations. Only Metal and WebGPU need the SPIR-V detour.

This is why the register rules had to be data and not convention. A slot is decided in one function, from one declaration list, and then either baked before release or resolved in process while someone edits a shader. A shader compiled in process and a shader cooked offline cannot disagree about a binding, because neither path is deciding one.

The SPIR-V detour has its own small lessons. DXC emits names that SPIRV-Cross needs stripped, SPIRV-Cross caches input locations at construction so they have to be rewritten first, and Tint asserts on non-finite constants that DXC is happy to fold out of dead code. Three tools, three ideas of what valid SPIR-V is.

The receipt

The test of a design like this is what happens when a new backend arrives after it has shipped. WebGPU was the fifth. Adding it to the layout system was a small, self-contained change to the layout preprocessor. It changed none of the layout files and none of the C++ call sites.

We first shipped this idea as the Shader Resource Table in a commercial cross-platform rendering framework, and SETech is where we took it further: to constant buffer bodies, vertex layouts and five backends.

You can watch it work. World Alone streams a city from OpenStreetMap around any GPS point you give it: terrain, facades, water, text, sky, clouds and shadows. Every one of those shaders declares its resources in a .srl.h, and the same declarations run on D3D11, D3D12, Vulkan, Metal and, in your browser, WebGPU.