Efficient Validation for
LLM-Generated Rendering Optimizations

Mikhail Dereviannykh Dmitrii Klepikov Hendrik Brucker Dominik Wüst Benedikt Krimmel Reiner Dolp Carsten Dachsbacher
Karlsruhe Institute of Technology
Paper (soon) Video Scenes Supplementary (326 MB) BibTeX
OriginalOptimizedFLIP error
1.07–2.19×faster Shadertoy shaders at FLIP threshold 0.01
1.89× / 1.70×MaterialX material and a Godot engine render pass
47–85%fewer validation frames than conservative validation

Abstract

LLM-driven evolutionary search can propose many rendering-code optimizations, but each approximate candidate must still be validated across dynamic viewing, lighting, material, and temporal conditions. We use a backend-specific Primary Render Parameters Space (PRPS) to sample rendering conditions, and formulate candidate acceptance as a Bayesian sequential test over the rate at which FLIP error exceeds a user threshold. To reduce rendering work, the validator adapts an empirical SAFE decision threshold: globally from the observed distribution of rejected candidates, and per candidate from LLM-derived mutation embeddings compared with previously validated neighbors. Rendered evidence remains the acceptance basis; previous candidates only modulate how strict the sampled-evidence threshold must be before acceptance. Across Shadertoy shaders, MaterialX-generated GLSL materials, and two Godot shader settings, the pipeline finds 1.12×–2.28× speedups under stated error thresholds, while replay and live-pipeline studies reduce validation frames by 47–85% relative to conservative validation with few disagreements against conservative replay labels. The resulting contribution is an efficient sampled-validation mechanism for open-ended rendering optimizers, with guarantees defined over sampled PRPS evidence.

Method

The search loop follows FunSearch and AlphaEvolve: an LLM proposes patches, and an island-based evolution keeps the fast ones. For rendering, correctness is perceptual, depends on viewing conditions, and can only be sampled. So we build the system around the validator. Rendering runs on our own distributed render system, scaled to up to 64 RTX 3060 GPUs rented on Vast.ai.

Render program shader / pass Evolutionary search islands, parent + prior programs LLM mutator proposes a code patch Bayesian sequential validator (ours) sample r ∈ [0,1]D reference candidate FLIP error > ε ? exceedance τ posterior over rate p undecided: sample and render the next batch SAFE → accept P(p<τ) ≥ γsafe(i) BAD → reject P(p<τ) ≤ 0.01 budget out → reject Profiler timing for SAFE candidates only Trust-guided threshold (ours) diff LLM views vs. safe / bad low trust · 99% high trust · early γsafe(i) sets γsafe(i) per candidate labels, rates patch accepted program with its render time and error → population

1Render-parameter space

Every dynamic input the application exposes (camera, time, lights, material inputs, environment) becomes one coordinate of a unit cube r ∈ [0,1]D. The validator samples the cube; a small adapter maps points to renderer state. D ranges from 1 to 34.

2Sequential Bayesian test

A frame is an exceedance if its FLIP error exceeds ε. With k exceedances in n frames, the rate p ~ Beta(1+k, 1+n−k). Accept when P(p < τ) ≥ γsafe, reject when it falls below 1−γhigh.

3Trust-guided threshold

How much clean evidence acceptance needs is set per candidate. Mutations whose LLM embeddings resemble validated safe edits may stop early; risky ones stay at the conservative 99%. Rejection is never relaxed.

Buoy render OriginalOurs|Error| ×20

Where validation time goes

Clearly bad candidates are rejected within tens of frames. A few borderline ones, failing on 1–5% of conditions, need hundreds — and under one fixed 99% threshold every safe candidate pays that price: 916 renders.

Candidate diff - for (i<200; i++) + for (i<180; i++) float t = ...; - c = a / b; + c = a * inv_b; ... new mutation i LLM describes it mechanismwhat computation changed guardwhich conditions are at risk visualwhat a viewer would see mechcalcalibrated on validated neighbors one embedding per view Contrast with history nearest safe vs. 3 nearest bad (rate-weighted) RCv = bad_scorev − safe_maxv Fuse and rank s(i) = Σv wv RCv KTA-learned weights rank among N candidates trust = exp(−rank / N) SAFE threshold γhigh = 0.99 trust 0: 916 renders γlow (learned) trust 1: few renders γeff = γhigh + (γlow − γhigh) · trustp
Try itChange the trust or the candidate.

Posterior over the exceedance rate p

P(p < τ) as frames arrive

Results

Eleven Shadertoy shaders, three MaterialX materials compiled to GLSL, and two Godot workloads, one of them an engine-internal render pass. All results use τ = 1%.

(a) Speedup of the best accepted program

(b) Validation frames vs. conservative, session replay

(c) Validation samples per candidate, live runs

Level of detail

Optimized render
optimized
FLIP error map
FLIP
// 4ltfDr (1.69×): audio-reactive second stage-  intersectStage(surface, pp,-      rotateMat(angle.x, angle.y, angle.z) * rot, scale);+  if (beat >= 32.0) {+      intersectStage(surface, pp,+          rotateMat(angle.x, angle.y, angle.z) * rot, scale);+  }

Domain reasoning: the second stage is invisible before beat 32, so it is skipped there. The same patch hoists loop invariants out of the fractal.

Hover a dot to see its lineage.

Per-scene results

Best accepted program per scene, with what the whole optimization session cost: LLM API calls and rented GPU time for rendering and validation.

Acknowledgements

We thank Johannes Schudeiske for taking part in the discussions that shaped this project, our student Ilya for the early studies, Kirill Bazhenov for his help in connecting us with industry, and Angelo Pesce (Roblox) and Brian Karis (Epic Games) for their feedback from the industry perspective.

We are especially grateful to Inigo Quilez, who kindly granted us permission to use his shaders in this work, and to all Shadertoy, Godot and MaterialX authors whose work we build on. All shaders remain the property of their authors; their licenses are listed below.

Shadertoy

Shadertoy shaders without an explicit license fall under Shadertoy's default CC BY-NC-SA 3.0 Unported.

Godot

  • SSIL render pass: Godot Engine (MIT, © Godot Engine contributors), ported to Godot by Clay John from Intel's ASSAO (MIT, © 2016 Intel Corporation). Scene: Godot TPS demo.
  • Black hole: Black Hole Extended by HyperJragon, godotshaders.com (MIT).

MaterialX

  • Linen, Bricks, Car Paint: Standard Surface example materials from MaterialX (Apache 2.0, © Contributors to the MaterialX Project); the generated GLSL includes OSL-derived code (BSD-3-Clause, © Sony Pictures Imageworks).

Metric

  • NVIDIA FLIP (BSD-3-Clause) for all image-error measurements.

BibTeX

@misc{dereviannykh2026renderopt,
  title  = {Efficient Validation for LLM-Generated Rendering Optimizations},
  author = {Dereviannykh, Mikhail and Klepikov, Dmitrii and Brucker, Hendrik and
            W{\"u}st, Dominik and Krimmel, Benedikt and Dolp, Reiner and
            Dachsbacher, Carsten},
  year   = {2026}
}