We deleted our OpenGL backend, and with it the WebGL build of the engine. For a while the web entry point was a single assert saying the platform was no longer supported. Then the browser build came back on WebGPU, and it could do things the old one never could. This article is about why we made that trade, what compute shaders bought us in the browser, and what broke on the way.
Everything here is from SETech, our in-house C++ engine, and you can run the result: the strand hair demo, World Alone and Studio Lite are all this engine compiled to WebAssembly and rendering through the browser's native WebGPU.
What the WebGL build was
The old web build ran on our OpenGL ES 2 subsystem, which in a browser means WebGL. Shaders were written in HLSL like everywhere else in the engine, compiled to SPIR-V with DXC, and translated to GLSL with SPIRV-Cross. It worked, it was small, and it ran almost anywhere.
So this was not a broken backend. It was a backend that had stopped being able to describe the engine.
Why we removed it
The short answer is compute. Close to half of the engine's shaders are compute shaders now, and they are not an optional layer. They are how the frame is built:
- Clustered light culling, one thread per froxel, builds the light lists every lit pixel reads.
- Skinning runs in a compute pass.
- Particles are simulated, emitted and sorted on the GPU.
- Triangle filtering feeds the visibility buffer, and a similar kernel culls geometry per tile for the spot light shadow atlas.
- Physics and deformables: cloth, ropes, soft bodies, fluids and hair are dozens of small kernels (predict, hash, solve, apply).
OpenGL ES 2 has no compute stage and no storage buffers. Keeping WebGL alive would have meant one of two things. Either every one of those passes grows a CPU fallback that has to be written, tested and kept matching, which is a second engine. Or the web build freezes at the feature set of the day compute arrived, which is a demo of an engine we no longer ship.
There was a second, quieter reason. Around the same time the engine moved to its shader resource layout: every shader's resources grouped into four update frequencies and bound as four sets. That model maps directly onto register spaces, descriptor sets and bind groups. It has nowhere to go in an API that has none of them.
So the backend went, and the GLSL translation path went with it. One fewer shading language to target, one fewer binding model to special-case, and a web build that was honest about being missing until it could be real.
What compute buys in the browser
On a desktop build, compute is one source of parallelism among several. In our browser build it is the only one. The WebAssembly build is single-threaded: without a shared-memory worker pool, creating a thread throws, so the engine's job system runs each job inline.
#if defined( __EMSCRIPTEN__ ) && !defined( __EMSCRIPTEN_PTHREADS__ )
/// single-threaded wasm: std::thread's constructor THROWS ( no
/// SharedArrayBuffer worker pool ). Run the job inline
ThreadEntry( this );
#else
std::thread worker( &SEThread::ThreadEntry, this );
worker.detach();
#endifOne CPU core, running WebAssembly, is a hard budget. Anything that scales with content has to live on the GPU or it does not ship on the web. Compute is what makes that possible.
The hair demo is the clearest case. Its guide strands are simulated in a compute pass with one thread per strand. Those guides are expanded on the GPU into tens of thousands of rendered strands, each a camera-facing ribbon built in the vertex shader directly from the solver's node buffer. There is no vertex data for the hair, no CPU work per strand, and nothing is read back. The single CPU thread is left to run the page.
The same holds across the other demos. World Alone culls its lights per froxel on the GPU while the CPU thread is busy turning OpenStreetMap data into geometry. None of this was expressible in the WebGL build, at any frame rate.
One solver, five backends
Portable compute needs some care. Our constraint solver lets many threads push corrections into the same node, which wants an atomic add, and GPU atomics are integer-only. So corrections are accumulated in fixed point, along with a contribution count:
// InterlockedAdd ( portable across all backends, no float atomics )
InterlockedAdd( SimDeltas[ uiA * 4 + 0 ], ( int )( vDxA.x * SETECH_SIM_DELTA_SCALE ) );
InterlockedAdd( SimDeltas[ uiA * 4 + 1 ], ( int )( vDxA.y * SETECH_SIM_DELTA_SCALE ) );
InterlockedAdd( SimDeltas[ uiA * 4 + 2 ], ( int )( vDxA.z * SETECH_SIM_DELTA_SCALE ) );
InterlockedAdd( SimDeltas[ uiA * 4 + 3 ], 1 );A separate kernel then takes the Jacobi average, applies it and clears the accumulators. Nothing in that is specific to the web, which is the point: the kernel that runs on D3D12 and Metal is the kernel that runs in the browser, compiled from the same HLSL.
What broke, and how quietly
The port itself was one more backend behind the engine's graphics interface. The interesting part was how the failures presented, because almost none of them announced themselves.
WebGPU has no shader reflection
There is no API to ask a shader module what it binds. Our shaders reach the browser as WGSL produced by Tint, so the backend reads that text: it anchors on each module-scope var, scans a bounded window for the @group and @binding attributes, which Tint can emit in either order, and classifies the type by scanning forward to the semicolon.
The same scan answers a subtler question. A 32-bit float texture cannot be bound as filterable unless the device offers that feature, and some do not, mobile Safari among them. So the backend checks whether a texture is ever the first argument of a textureSample or textureGather call. If it is only ever loaded, it is declared unfilterable and works everywhere.
One storage buffer over the floor is a black screen
WebGPU guarantees only eight storage buffers per shader stage. Our shared per-frame group alone needs more than that in the fragment stage: the scene lights, the froxel grid and its index list, and the spot shadow tile records. One over the floor invalidates the pipeline layout, so every pipeline built from it fails, and the only message is that something is invalid due to a previous error. The screen stays black.
Desktop adapters offer far more than the guaranteed floor. The fix is to ask the adapter for what it actually has instead of accepting the defaults, and the same goes for the storage buffer binding size, where the default is smaller than our light cluster buffer at higher render resolutions.
limits.maxStorageBufferBindingSize = adapter.limits.maxStorageBufferBindingSize;
limits.maxBufferSize = adapter.limits.maxBufferSize;
limits.maxStorageBuffersPerShaderStage = adapter.limits.maxStorageBuffersPerShaderStage;A cache that only grew
Bind groups are immutable, so we cache them by a fingerprint of the resources they hold. The fingerprint keys on each resource's unique creation id and not its pointer, because a render target recreated on resize often lands on the freed object's address, and a pointer key would keep serving a bind group that pins the dead texture. The cost of unique ids showed up later. When World Alone streams a new location, every resource is new, so every cached entry describing the old world becomes unreachable. They were capped per slot but never evicted, and we watched the count of live bind groups climb into the hundreds of thousands over a session of location changes. Under Dawn on Windows, tens of thousands of live bind groups are tens of thousands of D3D12 descriptor heap entries, which is a finite resource.
A slot legitimately cycles through a small set of fingerprints, so the cache now keeps only a handful per slot. The lesson is general: a cache keyed on identity needs an eviction story the day the identities start changing.
Four bind groups, exactly
One limit worked in our favour. WebGPU allows four bind groups, and the engine's layout has exactly four update frequencies. The grouping we had chosen for D3D12 and Vulkan fit the browser without a special case, which is the best evidence we have that grouping by update frequency is the right invariant and not an accident of one API.
No shader compiler in the page
Shaders for the web are cooked offline only. HLSL goes through DXC to SPIR-V, two small binary patches fix things the next tool rejects, and Tint produces WGSL. One of those patches replaces non-finite float constants, which DXC can fold out of dead code and which are valid SPIR-V, because Tint asserts on them. If a cooked shader is missing, the web build says so and stops. There is nothing to fall back to, and that is deliberate.
What it cost
Reach. WebGL ran nearly everywhere. WebGPU needs a recent Chrome, Edge or Safari, and a visitor without it gets a message instead of a demo. We decided that showing the real engine to most people was worth more than showing an old one to everyone, and that trade will not suit every project.
Size is the other cost, and it is smaller than it looks if the server is set up properly. Our WebAssembly modules are a few megabytes each, and they compress to roughly a third of that once the server sends them as application/wasm with compression enabled. Ours was doing neither until we checked real downloads instead of trusting the response headers of a HEAD request. It is worth checking yours.
The point
We did not remove WebGL because it was slow. We removed it because the engine had become a compute engine, and a backend without compute could only ever show a different, older engine. WebGPU gave the browser build the same frame as the native builds, from the same shaders and the same resource layout, on a single CPU thread. That is what made it worth going without a web build for a while.