Correct gl_FragCoord Y orientation on D3D, Metal and Vulkan - #1840
Conversation
There was a problem hiding this comment.
Pull request overview
This PR normalizes gl_FragCoord.y to OpenGL/WebGL’s bottom-left-origin convention on the non-OpenGL backends (D3D, Metal, Vulkan) by injecting an AST rewrite during shader compilation and supplying the render-target dimensions at draw time.
Changes:
- Add a new shader-compiler traverser to rewrite every fragment-stage
gl_FragCoordread to a Y-flipped equivalent using a new target-size uniform. - Populate the injected target-size uniform from the currently bound framebuffer dimensions in
NativeEngine::DrawInternal. - Add two render-and-readback unit tests to pin down
gl_FragCoord.yorientation and its composition with sampler coordinate flips.
Reviewed changes
Copilot reviewed 12 out of 12 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
| Plugins/ShaderCompiler/Source/ShaderCompilerVulkan.cpp | Runs FlipFragCoordY in the Vulkan compilation pipeline before uniform transforms. |
| Plugins/ShaderCompiler/Source/ShaderCompilerMetal.cpp | Runs FlipFragCoordY in the Metal compilation pipeline before uniform transforms. |
| Plugins/ShaderCompiler/Source/ShaderCompilerDXIL.cpp | Runs FlipFragCoordY in the DXIL compilation pipeline before uniform transforms. |
| Plugins/ShaderCompiler/Source/ShaderCompilerDXBC.cpp | Runs FlipFragCoordY in the DXBC compilation pipeline before uniform transforms. |
| Plugins/ShaderCompiler/Source/ShaderCompilerTraversers.h | Declares and documents the new FlipFragCoordY traverser API. |
| Plugins/ShaderCompiler/Source/ShaderCompilerTraversers.cpp | Implements FragCoordYFlipTraverser, declares bnFragCoordTargetSize, and rewrites gl_FragCoord reads. |
| Core/Graphics/InternalInclude/Babylon/Graphics/BgfxShaderInfo.h | Introduces FRAGCOORD_TARGET_SIZE_UNIFORM_NAME constant for the injected uniform name. |
| Plugins/NativeEngine/Source/Program.h | Adds cached lookup accessor for the injected uniform’s UniformInfo. |
| Plugins/NativeEngine/Source/Program.cpp | Caches bnFragCoordTargetSize uniform info during program initialization. |
| Plugins/NativeEngine/Source/NativeEngine.cpp | Sets bnFragCoordTargetSize each draw based on the bound framebuffer size (when present). |
| Apps/UnitTests/Source/Tests.ShaderCompilation.FragCoord.cpp | Adds render/readback tests validating gl_FragCoord.y orientation and UV-vs-fragcoord addressing equivalence. |
| Apps/UnitTests/CMakeLists.txt | Adds the new test source file to the UnitTests build. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
CI caught a real problem with the test, though not with the fix. Pushed 666e1ad. What failed: Why: the test, not the traverser. Worth stating plainly: Fix: the test now writes normalized Re-verified the negative control — with the and with it restored: So the reworked assertion still fails by the full range of the ramp when the fix is absent — it did not become weaker by becoming portable. |
|
Short answer: no — and I can now show why rather than just reporting that the candidates I tried didn't flip. I instrumented Result: the traverser is invoked constantly and matches nothing. Across a 37-test spread sampling the whole 720-test catalog, plus ~40 hand-picked tests covering every feature whose shaders reference The reason is that in
and the features that do use it unconditionally — OIT, TAA, curvature, volumetric, clustered lighting, the FrameGraph tests — abort before any shader is compiled, e.g.: Those depend on engine APIs added to Babylon.js after 9.15.0, so they fail during scene construction regardless of this change. Two consequences worth stating explicitly:
This is precisely why the change ships with the two GPU-readback unit tests: they exercise Once the pinned Babylon.js moves past 9.15.0 the OIT / TAA / curvature / clustered-lighting tests become the natural integration coverage for this, and I'm happy to follow up with a PR enabling them at that point. |
|
Pushed three follow-up commits: the Babylon.js 9.21.2 bump, an engine fix the bump exposed, and the validation tests it unlocks. The bump exposed a real bug
Babylon.js 9.21 draws thin instances with render-self motion blur using Regression checkAll 302 previously-enabled tests, one process per test, on Win32 D3D11: 299 pass, 0 fail. The only non-passing are 53-55 (scissor), which crash identically on 9.15.0 — pre-existing on Windows/D3D11 and unrelated to this change. 9 tests enabledSeveral of these exclusions describe order-dependent behaviour, which a per-test sweep structurally cannot reproduce, so I also ran a single sequential process over indices 56-719: ran=256 passed=256 failed=0.
Two things I want to flag honestly1. Three of these can only be judged by CI. Their exclusion reasons are backend-specific and I have no way to reproduce them on a D3D11 host: 137 "fails on Linux (large diff)", 287 "fails to compile on desktop GL", and 299 OpenGL 2. Tests 321 and 323 are now marginal — 99.1% and 94.2% of their error budget (2.478% and 2.355% against 2.5%). Bit-identical across three runs, so not flaky, but that margin is unlikely to survive a different backend. The cause is visible in the render: Babylon Native leaves a soft motion-blur halo around objects where Babylon.js converges to zero velocity, so there is a residual gap in the motion-blur path beyond the instance-limit bug. I ruled out stale reference images (substituting Babylon.js's own PNGs gives identical diffs). Worth a follow-up issue; I did not want to hide it behind a raised Note that none of the newly enabled tests exercise |
e5e1799 to
b6bc698
Compare
|
Correction to my previous comment: the instance-limit commit I pushed here duplicated work that already exists in #1839, so I have dropped it and force-pushed. This PR now depends on #1839. The Babylon.js 9.21.2 bump regresses three tests that #1839 fixes:
Until #1839 merges, those three will fail here. They should not be worked around in this PR. For the record, my dropped commit and #1839 converged on the same guard independently ( #1839 also explains the residual halo I flagged, and it is a third, separate bug: The rest of this PR is unchanged: the |
|
Pushed What was failingThe three Ubuntu jobs aborted at test 287 "Prepass SSAO + particles" with Root causebgfx exposes every non-sampler uniform as a float
uniform vec4 samples;
...
mediump int _42 = -samples.x;
for (mediump int i = _42; i < samples.x; i += 2)HLSL tolerated this because it converts implicitly, which is why it only ever showed up on the OpenGL and OpenGLES backends. FixShape the value to float first, then ask glslang for a real conversion node via ValidationI stood up an ANGLE/GLES build on Windows (
The eleven are the whole prepass SSAO family - 287, 289, 293, 296, 299, 302, 304, 305, 306 - plus 176 "GUI Slate" and 177 "GUI Near Menu", which are still excluded on master but were hitting the same bug. Still redThe three Win32 failures (321/322/323, the motion blur trio) are fixed by #1839, which is still awaiting review. They should go green once that lands and this branch rebases. Out of scope, noted for laterOn the ANGLE build test 363 "Screen Space Reflections 2" trips a different, pre-existing bgfx bug: |
|
The uniform fix worked - the Ubuntu jobs no longer hit All four Ubuntu jobs stop at byte-identical positions (159 completed comparisons, then this). The program being linked at the crash carries Since that is a Mesa/LLVM bug rather than anything Babylon Native emits, I have re-excluded just test 287 in The other eight tests enabled by the 9.21.2 bump stay enabled, and a full sequential ANGLE/GLES run confirms all of them now render: 289, 293, 296, 299, 302, 304, 305, 306 pass, with zero |
|
Correction to my earlier comment: I claimed the Progress so far on this branch:
The bgfx code in question is: const GLenum drawBuffer = GL_COLOR_ATTACHMENT0 + colorIdx;
GL_CHECK(glDrawBuffers(1, &drawBuffer) );GLES requires I am reproducing the sequential run locally on ANGLE against this exact config to pin down which test triggers it - in per-test isolation every test in that range passes, so it looks like it depends on run order rather than on one scene. Will report back. |
|
Found it, and the branch should now be as green as it can get without #1839. The
|
| ran | passed | failed | |
|---|---|---|---|
| ANGLE/GLES, full sequential | 286 | 282 | 4 |
| D3D11, full sequential | 300 | 297 | 3 |
Zero asserts and zero BGFX FATAL on either. The three shared failures are the motion blur trio (321, 322, 323), which #1839 fixes; the fourth on ANGLE is an ANGLE-only pixel difference in MeshDebugPluginMaterial that does not reproduce on D3D11.
So the expected CI result is Win32 and Ubuntu both red on exactly those three tests, and green once #1839 lands and this rebases.
Summary of what changed on this branch for CI
2eaaf94d- restore the original basic type when narrowing widened uniforms. Fixesintandbooluniforms on OpenGL/GLES. Eleven tests go fromFatal::InvalidShaderto passing on ANGLE.e6c431e1,e5490d2f- exclude the prepass SSAO family; llvmpipe on the runner aborts withLLVM ERROR: Cannot emit physreg copy instruction, a Mesa register allocator bug. They pass on D3D11 and ANGLE.68508845- enable GUI Near Menu on OpenGL, exclude Screen Space Reflections 2 for the bgfxglDrawBuffersbug above.
Root cause of the
|
| Chakra | V8 | |
|---|---|---|
scene.isReady() |
false forever | true @ frame 200 |
glowLayer.isReady(subMesh) |
false forever | true |
drawWrapper.effect.isReady() |
true @ frame 50 | true |
| post-processes ready | true @ frame 50 | true |
_shadersLoaded |
false forever | true |
isLayerReady() |
false forever | true |
Everything is ready except _shadersLoaded, and the one promise that never settles is ThinGlowLayer._importShadersAsync() — it neither resolves nor rejects.
The actual bug: super inside an arrow function on Chakra
Chakra resolves super.x inside an arrow function nested in a class method to the derived class's own method instead of the base one:
class A { foo() { return "BASE"; } }
class B extends A {
foo() {
const s = Object.create(null, { foo: { get: () => super.foo } });
return s.foo.call(this);
}
}
new B().foo();| Chakra | V8 | |
|---|---|---|
resolved fn === B.prototype.foo |
true | false |
resolved fn === A.prototype.foo |
false | true |
| calling through it | Error: Out of stack space |
"BASE" |
TypeScript emits precisely that Object.create(null, { get: () => super.x }) helper for a super call inside an async method, and Babylon.js started shipping it in the UMD bundle in 9.16.0 — exactly where this PR's bump from 9.15.0 to 9.21.2 crosses. Bisected by swapping babylon.max.js: 9.15.0 passes, 9.16.0 hangs, and there are no patch releases in between.
Called synchronously it dies with Out of stack space. Called from a promise chain — which is what _importShadersAsync is — every level is a fresh microtask, so it recurses forever without overflowing the stack: never settles, never throws, burns CPU and heap. That is the runner signature.
Fix
The repo already has the remedy: Apps/scripts/downlevelNativeScripts.mjs, added in #1789 for this exact reason —
Babylon Native's Chakra engine consumes ES5-level script, so the bundle must be down-leveled before it runs.
TypeScript's ES5 emit rewrites super.x to _super.prototype.x and removes the arrow entirely, so the bug cannot be expressed. The script was only ever wired into getNightly, so every build that takes Babylon.js from npm — i.e. every normal build — ran the un-downleveled ES2015 bundle. Running it from postinstall closes that gap for npm install, npm ci, CI and local builds alike; the nightly path is untouched (getNightly.js still downlevels the files it refills from the CDN).
Validation
Windows / D3D11, Debug, tests 0-52 and 56-719 (53-55 crash locally in Debug regardless of this change):
| engine | before | after |
|---|---|---|
| Chakra | hangs at test 23 | 297/300 |
| V8 | 297/300 | 297/300 |
Byte-identical on V8 — no regression from the ES5 emit — and Chakra now matches it. The three remaining failures are the motion-blur trio (Thin instances + dynamic buffer resize, Instances + render self motion blur, Thin instances + render self motion blur) that #1839 fixes.
|
The 9 red jobs here are all the same three validation failures, and they are not caused by the
I reproduced this locally on Win32 D3D11 (RelWithDebInfo), running only the three tests:
So this PR is blocked on #1839 rather than needing a change of its own. Once #1839 lands, rebasing here should turn all 9 jobs green. One thing worth flagging while that is in flight: even with #1839 the margins are thin — 2.478% against a 2.5% budget on |
f172dbb to
b3a447b
Compare
b3a447b to
4bca7b4
Compare
9aba372 to
6c85611
Compare
Babylon Native's shader model is "shader-visible coordinates are GL-logical; convert to physical at each sampler access". This is implemented by FlipSamplerCoordinatesTraverser (texture() v -> 1-v, texelFetch y -> h-1-y) and InvertYDerivativeOperandsTraverser (negate dFdy), which run for DXBC/DXIL/Metal/Vulkan but not OpenGL. gl_FragCoord was the one shader input left in physical space. D3D, Metal and Vulkan rasterize with a top-left origin while GL uses bottom-left, and BN does not flip geometry (ProcessShaderCoordinates only remaps depth). So for GL row y the hardware yields height - y - 0.5 instead of y + 0.5, i.e. gl_FragCoord.y arrives mirrored. Shaders using the symmetric "sample at my own position" pattern are unaffected because the physical/physical pairing is self-consistent. The mismatch only shows up where the row index itself is meaningful: prefix sums (iblCdfy), neighbour offsets, and copies into a differently-oriented target (copyTexture3DLayerToTexture). That is why 39 shaders reference gl_FragCoord but only a handful render incorrectly. Add FragCoordYFlipTraverser, which rewrites every gl_FragCoord read in the fragment stage to vec4(fc.x, targetHeight - fc.y, fc.z, fc.w). The correction is exactly `height - y` with no -1 term (see derivation above). Shaders that never read gl_FragCoord are left byte-for-byte unchanged. The target height comes from a new vec4 uniform, bnFragCoordTargetSize, declared as a linker object so MoveNonSamplerUniformsIntoStruct sweeps it into the "Frame" struct like every other uniform and it is emitted by name into the bgfx uniform table. NativeEngine sets it in DrawInternal from the bound framebuffer's dimensions. bgfx's predefined u_viewRect is deliberately not used: it is narrowed to the viewport by FrameBuffer::SetBgfxViewPortAndScissor whenever one is set, whereas gl_FragCoord is relative to the whole render target. FlipFragCoordY must run before ChangeUniformTypes / MoveNonSamplerUniformsIntoStruct so the uniform is collected with the rest. A fresh replacement subtree is built per occurrence rather than reusing MakeReplacements, which maps one node per symbol name and would give that node multiple parents - something later traversers do not expect. OpenGL is intentionally left alone, as with the other flip traversers. Validated on D3D11: 149 tests validated with 0 pixel-diff failures. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 60c2ec68-6de1-445d-9fc9-b699db737eae
Two render-and-readback tests in UnitTests, both gated off where the existing render tests are (D3D12, noop Metal device). FragCoordYIncreasesUpwards writes gl_FragCoord.y / height into a render target and checks the ramp is brightest at the top row, matching GL's bottom-left origin. Values come out as 253 / 126 / 2 for the top, middle and bottom rows of a 64-row target, exactly the (height - row - 0.5) / height ramp the correction is derived from. FragCoordAndUVAddressATextureIdentically samples one texture twice, once through the interpolated UVs of a full-screen quad and once through gl_FragCoord.xy / targetSize, and requires the two images to match. That is the addressing pattern used by order-independent transparency, TAA and screen space curvature, and it only holds if the gl_FragCoord correction and FlipSamplerCoordinatesTraverser compose to a no-op. Comparing the two addressing modes against each other rather than against the source pixels keeps the test independent of how createRawTexture orients its upload. Both fail without FlipFragCoordY: the first ramp inverts to 2 / 129 / 253 and the second renders vertically mirrored (255..3 against 3..255). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 60c2ec68-6de1-445d-9fc9-b699db737eae
Matches Tests.ShaderCompilation.cpp and Tests.UniformPadding.cpp, and avoids relying on the caps property being configurable.
Helpers::ReadPixels is a plain glReadPixels on OpenGL, which returns the bottom scanline first, whereas the D3D11 path returns the top scanline first. The test asserted an absolute ramp direction over readback rows, so it encoded the D3D11 readback convention and failed on Linux even though gl_FragCoord.y was correct there. Compare normalized gl_FragCoord.y against the interpolated vUV.y written by the same fragment invocation instead. The quad maps uv.y to clip y, so the two ramps must agree on every backend regardless of readback row order, and a flipped gl_FragCoord.y still misses by the full range of the ramp. Renamed to FragCoordYMatchesInterpolatedUV to match what it now checks.
6158d48 to
eff3a08
Compare
Bump the shader cache version so serialized v5 entries recompile with FlipFragCoordY, wait on render with get() like the other tests, replace gl_FragCoord when it is a TIntermBranch expression, and add a direct-return orientation regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: da74bc94-a7dc-4817-bd81-59b5c6b123fc
Keeps shotgun JsRuntimeHost/SPIRV wiring and existing FlipFragCoordY, PCF sampler-flip, compute Program::InitializeCompute, and inter-stage varying location assignment. Drops the duplicate FragCoord uniform upload introduced by the auto-merge of BabylonJS#1840. Takes upstream Playground validation cleanup (BabylonJS#1863) and shader-info cache bits. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 60c2ec68-6de1-445d-9fc9-b699db737eae
gl_FragCoord.ywas never flipped on the backends that rasterize with a top-left origin (D3D, Metal, Vulkan), so shaders reading it saw a vertically mirrored coordinate compared to WebGL/OpenGL. This breaks any shader that addresses a screen-sized texture by fragment position.ShaderCompilerTraversers::FlipFragCoordYrewrites eachgl_FragCoordread tovec4(x, targetHeight - y, z, w)on those backends. The height comes from a newbnFragCoordTargetSizeuniform thatNativeEnginesets per draw from the bound framebuffer; bgfx'su_viewRectcannot be used because it is narrowed to the viewport, whilegl_FragCoordis relative to the whole render target.Also switches parallel shader compile off with
nullinstead ofdelete, matching the other C++-hosted shader compilation tests.Three unit tests pin the orientation without depending on a backend's readback order: one checks
gl_FragCoord.yagainst the interpolated UV ramp of a full-screen quad, one that a directreturn gl_FragCoord;compiles and matches that ramp, and one thatgl_FragCoordand UV address a screen-sized texture identically.Additive only — no existing behavior changes on OpenGL.