A shader's resource interface is normally written down in three places and held together by hand: once in the shader, once in C++, and once in a reflection table that tries to reconcile the two after the shader has compiled. In SETech, our in-house engine, it is written once. One header is compiled twice, as shader code and as C++, and the engine binds resources on D3D11, D3D12, Vulkan, Metal and WebGPU without ever looking a name up.
This article walks through how that works, why the grouping underneath it is the part that matters, and what it bought us on D3D12. Everything here is from the engine itself. You can see it running in the World Alone demo, which is this same code compiled to WebAssembly.
The usual way: reconcile the two sides at runtime
The traditional flow has three steps. Declare the resource in the shader. Declare it a second time on the C++ side. Then match the two, usually by name, through a reflection table built when the shader compiles.
This is a reasonable trade, and the cost is not performance: the lookup resolves once at load, and the runtime carries integers from then on. The cost is where the check can happen. The table cannot exist until the shader has compiled, so nothing checks the two declarations against each other at the point where you write them. Rename a texture on one side and the C++ compiles, the shader compiles, and the texture is simply never bound. You get an unbound parameter, not a build error.
There is a real reason that cannot simply be turned into an error. An unused shader parameter is legitimately optimized out during shader compilation, so "this name is not in the parameter map" cannot be told apart from "the shader stopped using it". Once binding is by name against a compiled artifact, that ambiguity is unavoidable. Removing the name is what removes the ambiguity.
One pass, three files, one declaration
Screen-space ambient occlusion is a good small example. The pass needs a noise texture, a point sampler and some constants. In SETech the declaration lives in a shared group header, the shader just uses it, and the C++ points at the declaration instead of repeating it.
SE_DECL_STATIC_SAMPLER( gSamplerPointWrap )
SE_DECL_TEXTURE_2D( NoiseTexture )#include "SEBaseEngine.srl.h"
float3 rand = SE_SAMPLE_2D( NoiseTexture, gSamplerPointWrap, noiseUV ).xyz;aResources[ 5 ].m_pDesc = SE_RESOURCE_DESC( BaseEngine, Fixed, NoiseTexture );
aResources[ 5 ].m_ppTextures = &m_pNoiseTexture;
m_pFixedResourceGroup->Update( m_uiFixedRingSlot, aResources, 12 );SE_RESOURCE_DESC is not a string lookup. It expands to a struct member access into the layout that the header itself declared. Rename NoiseTexture in the header and this line stops compiling. The mistake that would silently unbind a texture under name matching does not build here.
One declaration, two compilers
A layout is declared in a .srl.h file (shader resource layout), built out of .srg.h fragments (shader resource groups). The same file is consumed by both toolchains, and the fork between them is one #ifdef __cplusplus in SEShaderMacros.h.
On the shader side, a small preprocessor walks the declaration list before the HLSL compiler sees it and rewrites each two-argument declaration into one that carries its register and space. The macros then expand to ordinary HLSL: a cbuffer, a Texture2D, a SamplerState. The target backend is a parameter to that preprocessor, which matters later.
On the C++ side, the same macros expand to a struct. Each group becomes a nested struct, and each resource becomes a member whose default initializer bumps a per-kind counter:
struct SELayout_BaseEngine {
struct _SEGroup_Fixed {
uint32 m_uiOffset = 0, m_uiCBVCount = 0, m_uiSRVCount = 0, m_uiSamplerCount = 0;
const SEResourceDesc ShadowMapSampler = { "ShadowMapSampler", eShaderResourceType_SamplerComparison,
0, 0, m_uiOffset++, m_uiSamplerCount++, 1 };
const SEResourceDesc gFixed = { "gFixed", eShaderResourceType_ConstantBuffer,
0, sizeof( SEFixedConstants ), m_uiOffset++, m_uiCBVCount++, 1 };
};
// ...three more groups
};The layout is a function-local static, so the whole walk runs exactly once, the first time anything asks for it, in declaration order. That is the honest description of the cost: not a compile-time constant, but one pass over a few dozen integer increments per layout, with no strings compared and no compiled shader consulted. The name string in each descriptor exists for debugging and is never read by the binding code. The implementation is deliberately plain: no templates, no constexpr, no standard library, no string processing.
Both sides walk the same list in the same order, so they land on the same slots by construction, not by agreement. There is no shader reflection anywhere in the engine's binding path.
The shared groups are shared on purpose
Almost every layout in the project opens the same way:
SE_DECL_SHADER_RESOURCE_LAYOUT_BEGIN( SlugText )
SE_DECL_RESOURCE_GROUP_BEGIN( Fixed )
#include "SEFixedGroup.srg.h"
SE_DECL_RESOURCE_GROUP_END( Fixed )
SE_DECL_RESOURCE_GROUP_BEGIN( PerFrame )
#include "SEPerFrameGroup.srg.h"
SE_DECL_RESOURCE_GROUP_END( PerFrame )
// ...then whatever this shader alone needsFixed and PerFrame are never retyped per shader. Nearly every layout in the engine opens with the same two fragments, in the same order, before anything of its own. The few exceptions are self-contained, such as the GPU physics solver, which touches no shared engine state and declares only its own group. So every shader that uses engine state agrees on where the camera, the lights and the samplers sit, and the engine binds those groups once for all of them.
On D3D11 this is structural, not a convention somebody has to remember. Shader Model 5 has no register spaces, so the counters run continuously and never restart per group. The identical prefix is the only reason the numbering lines up across shaders. There is a hard ceiling behind it too: D3D11 gives a shader stage 14 constant buffer slots, and the shared prefix spends a good share of them before a shader declares a single buffer of its own.
b0 gFixed b3 gGlobalLights
b1 gShadowSettings b4 gSky
b2 gPerFrame b5 gShadowPS
b6 gSlugText // the first register this shader ownsBecause that prefix is the same everywhere, the remaining slots are countable. If every shader invented its own prefix, nobody could say how many were left.
Update frequency is the invariant
The five APIs disagree about almost everything: register spaces, how counters run, whether samplers get their own heap. They agree on exactly one thing, and none of them says it out loud. Resources should be grouped by how often they change while a command buffer is being built.
SETech has four groups: Fixed, PerFrame, PerBatch and PerDraw. A register space on D3D12, a descriptor set on Vulkan, a bind group on WebGPU: three names for the same idea, and the declaration mentions none of them.
One declaration, five outputs
Here is the text renderer's per-batch group, declared once. It holds the entity constants, a texture of Bezier control points and a band acceleration texture.
SE_DECL_RESOURCE_GROUP_BEGIN( PerBatch )
SE_DECL_CONSTANT_BUFFER( SESlugTextConstants, gSlugText )
SE_DECL_TEXTURE_2D( SlugCurveTexture ) // RGBA32F Bezier control points
SE_DECL_TEXTURE_2D( SlugBandTexture ) // RGBA8, uint16 pairs per band
SE_DECL_RESOURCE_GROUP_END( PerBatch )And here is what the backends compile, with the shared PerFrame buffer alongside so the spaces are visible. Watch the two constant buffers.
cbuffer _cb_gPerFrame : register( b2 ) { ... };
cbuffer _cb_gSlugText : register( b6 ) { ... };
Texture2D SlugCurveTexture : register( t11 );
Texture2D SlugBandTexture : register( t12 );cbuffer _cb_gPerFrame : register( b0, space1 ) { ... };
cbuffer _cb_gSlugText : register( b0, space2 ) { ... };
Texture2D SlugCurveTexture : register( t1, space2 );
Texture2D SlugBandTexture : register( t2, space2 );cbuffer _cb_gPerFrame : register( b0, space1 ) { ... };
cbuffer _cb_gSlugText : register( b0, space2 ) { ... };
Texture2D SlugCurveTexture : register( t0, space2 );
Texture2D SlugBandTexture : register( t1, space2 );@group(2u) @binding(0u) var<uniform> _cb_gSlugText : S_1;
@group(2u) @binding(1u) var SlugCurveTexture : texture_2d<f32>;
@group(2u) @binding(2u) var SlugBandTexture : texture_2d<f32>;PerFrame and PerBatch are both b0 on D3D12 and Vulkan, and only the space tells them apart. D3D11 has no spaces, so the same two buffers land on two different flat registers, and the textures are pushed up the t range by every shader resource view declared ahead of them in the shared groups. The exact numbers in these blocks are a snapshot and will drift as the shared groups grow. The shape will not. Metal goes through SPIRV-Cross, where the cooker assigns flat buffer, texture and sampler slots from the same per-group numbering. Moving that to one argument buffer per group is the natural next step, and the declarations would not change.
What the frequencies buy on D3D12
A root signature used to be a per-effect object in our D3D12 backend. Every shader effect built its own, and the command buffer bound and switched them as it walked the frame. SetGraphicsRootSignature invalidates every root parameter, so each switch forced a full rebind of every table, whether or not anything in them had changed.
Once every shader in the engine agrees on the same four groups, none of that is needed. There are now two root signatures for all raster and compute work, one graphics and one compute, created once at init and set once when a command buffer begins. No draw ever changes them, so root parameters stay valid for the whole command buffer and only the tables that actually changed get rewritten.
Root parameters are ordered by descending update frequency, so the table that rebinds most often sits in the cheapest root slot. That ordering is only possible because the declaration already says how often a resource changes. The compute signature strips the graphics-only deny flags and is otherwise identical. The one exception is ray tracing: a DXR pipeline carries its own global root signature, and the command buffer restores the shared compute signature after DispatchRays.
Resources were only the first thing worth sharing
Once one file is compiled by both compilers, anything declared in the same shape stops being able to drift. Two more things moved in.
Constant buffer bodies
struct SESlugTextConstants
{
float4x4 matWorld;
float4 SelectionId;
float4 EngraveParams;
float4 TextParams;
};The HLSL side wraps it in a cbuffer. The C++ side compiles the identical body as a plain struct, because the shared header maps the HLSL type names onto the engine's math types:
typedef SEVec4 float4;
typedef SEMat4 float4x4;
typedef unsigned int uint;So the buffer's size comes from the declaration too. Nobody writes a sizeof that has to match a shader by hand.
Vertex layouts
SE_VERTEX_LAYOUT_BEGIN(VSTexturedColorIn)
SE_VERTEX_ELEMENT(SE_FLOAT4, m_Position, SE_POSITION)
SE_VERTEX_ELEMENT(SE_FLOAT2, m_TexCoord0, SE_TEXCOORD0)
SE_VERTEX_ELEMENT(SE_UINT, m_Color, SE_COLOR0)
SE_VERTEX_LAYOUT_ENDstruct VSTexturedColorIn {
float4 m_Position : LOCALPOS;
float2 m_TexCoord0 : TEXCOORD0;
uint m_Color : COLOR0;
};{ 0, SETECH_DATATYPE_FLOAT, 4, SETECH_DATAUSAGE_POSITION, 0, false },
{ 0, SETECH_DATATYPE_FLOAT, 2, SETECH_DATAUSAGE_TEXCOORD, 0, false },
{ 0, SETECH_DATATYPE_UINT8, 4, SETECH_DATAUSAGE_COLOR, 0, false },SE_FLOAT4 is float4 to one compiler and a datatype plus a count to the other. SE_POSITION is a semantic on one side and a usage enum on the other. A vertex layout that disagrees with its shader's input struct fails the same way a mis-numbered register does: it builds, it runs, and the geometry is wrong. Same file, same fix.
What did not move in is anything whose body needs real HLSL semantics or intrinsics, which is why the shader helper header is guarded with #if !defined( __cplusplus ). The rule is simple: only declarations whose bodies are expressible in both languages can be shared.
Where this actually runs
Runtime shader compilation is a development convenience, available on the D3D and Vulkan backends. Everything else is cooked ahead of time, per backend, by the engine's offline asset tool: the WebGPU build, for one, can only load cooked shaders. The same layout preprocessor feeds all five targets, with the backend as a parameter, and the in-process compile path calls that same preprocessor.
This is why the register rules had to be data and not convention. A slot is decided in one function, from one declaration list, and then either baked before release or resolved in process while someone edits a shader. A shader compiled in process and a shader cooked offline cannot disagree about a binding, because neither path is deciding one.
The SPIR-V detour has its own small lessons. DXC emits names that SPIRV-Cross needs stripped, SPIRV-Cross caches input locations at construction so they have to be rewritten first, and Tint asserts on non-finite constants that DXC is happy to fold out of dead code. Three tools, three ideas of what valid SPIR-V is.
The receipt
The test of a design like this is what happens when a new backend arrives after it has shipped. WebGPU was the fifth. Adding it to the layout system was a small, self-contained change to the layout preprocessor. It changed none of the layout files and none of the C++ call sites.
We first shipped this idea as the Shader Resource Table in a commercial cross-platform rendering framework, and SETech is where we took it further: to constant buffer bodies, vertex layouts and five backends.
You can watch it work. World Alone streams a city from OpenStreetMap around any GPS point you give it: terrain, facades, water, text, sky, clouds and shadows. Every one of those shaders declares its resources in a .srl.h, and the same declarations run on D3D11, D3D12, Vulkan, Metal and, in your browser, WebGPU.