Depth-Assisted Video Relighting in the Browser with WebGPU and WebGL
I added an editable video-relighting effect to Timeline Studio , my open-source browser video editor. It lets you move a virtual light around a scene, adjust its color and softness, and compare the result with the original while keeping the effect editable on the timeline. The useful part is the division of work: a depth model analyzes the footage, then a small GPU shader handles lighting…
A new feature has been added to Timeline Studio, an open-source browser video editor. This feature allows users to interactively relight video footage directly in the browser using WebGPU and WebGL. By moving a virtual light source around the scene and adjusting its color, softness, and position, users can see the changes in real-time without having to re-run the entire lighting analysis.
The system breaks down the process into two main components. First, a depth model analyzes the video footage, determining the distances and orientations of various elements within the scene. This depth information is then used by a small GPU shader to make lighting adjustments. Importantly, moving the light source does not require an additional inference pass, making the process more efficient.
The editor already incorporated depth estimation for cinematic depth-of-field and grayscale depth-map videos. Relighting builds upon this foundation by utilizing the same scene-depth data. The pipeline for this feature includes decoding sampled frames, running Depth Anything V2 Small for depth analysis using Transformers.js and a WebGPU inference path, normalizing and refining the depth maps, and rendering the original frame and depth textures through a WebGL lighting shader.
WebGPU handles the computationally intensive model, while WebGL takes care of the relighting effect, allowing for interactive adjustments while the expensive analysis can be reused. By estimating surface normals from the depth gradients, the shader creates an approximate surface representation. The lighting effect is then applied, taking into account factors such as color, strength, softness, and range.
The user interface features a draggable light-position sphere with controls for position ( azimuth and elevation ), distance, strength, color, range, and softness. Presets are available for common lighting styles, such as warm, neutral, and cool colors. A hold-to-view-original button allows users to compare the effect with the original footage without permanently applying the changes.
Since video requires precise alignment of source time, the system stores additional information such as trim settings, duration, playback rate, speed curve, quality, and model revision. This ensures that the correct depth map is used for each frame, even when the video is trimmed or played at different speeds. While the effect is generally smooth, it is not motion-compensated and may result in some flicker on certain clips.
Blending neighboring depth samples helps to smooth out abrupt changes in the depth map, although it does not perform true motion compensation. The depth refinement process uses a simple edge-guided smoothing pass to prevent mixing across boundaries and preserve the separation between foreground and background elements.
One important consideration is that seeking (changing the current position in the video) is an asynchronous operation. If the preview canvas is updated too early during a seek, it may display the previous frame, leading to a "seek bug." To mitigate this issue, the preview now avoids updating while the source is seeking and refreshes only after decoding events have occurred.
The main video also utilizes video-frame callbacks to synchronize the preview with the actual media frame, providing a more accurate representation of the current state.
It is important to note that the relighting effect is an approximation based on visible scene depth. The system does not reconstruct a complete 3D scene, recover material properties, or simulate physically accurate shadows and reflections. Its primary purpose is to add a directional fill, color accent, or restrained edge-lighting effect to existing footage. The algorithm is best suited for scenarios with strong motion or difficult depth boundaries, where more advanced techniques may not be feasible.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.