Files
Glint-Runtime/docs/en/architecture/02-performance.md

2.5 KiB
Raw Permalink Blame History

Performance Analysis

Measurements with the --perf flag on desktop.glbc.

Measurement Results

Steady state (no events, idle):

VDOM: 6.6ms | Style: 0.3ms | Render: 7.2ms | Total: 14.1ms

Under events (slider drag, peak values):

VDOM: 17.4ms | Style: 0.8ms | Render: 15.2ms | Total: 33.5ms ⚠️

60 FPS frame budget: 16ms. At idle we fit (14ms), under events — not (up to 33ms).

Bottleneck Analysis

Style matching — NOT a bottleneck (0.3-0.8ms)

Style matching takes less than 1ms even at peak. This is the result of:

  • Phase 4 (StyleIndex) — O(K) instead of O(N×M)
  • Phase 8 (StyleCache) — computed style memoization

VDOM eval — main consumer (6-17ms)

At idle ~6.6ms — a full traversal of the Element tree. Under events up to 17ms:

  • evaluate_vdom_incr recalculates dirty elements (Phase 3)
  • Each event can dirty entire subtrees
  • Internally: resolve_string, resolve_prop, compute_cached, recursive traversal

Render — second consumer (7-15ms)

This is where the main optimization headroom lies. render_element creates ALL Iced widgets from scratch every frame, even if the Element hasn't changed. Iced then diffs the new tree against the old one — but the tree construction itself costs ~7ms.

Scenario: slider

  1. SliderChangedage and volume_level change
  2. tracker.on_variable_changed → dirty_set for dependent elements
  3. evaluate_vdom_incr recalculates dirty elements and their children
  4. view()render_element for ALL elements (full re-render)
  5. Result: VDOM 13ms + Render 14ms = 27ms — frame drop

The first few frames after an event are the heaviest (VDOM ~13ms), then stabilize (~7ms) as dirty_set gradually clears.

Recommendations

  1. Phase 9.3 — widget cache — would reduce Render from 7ms to ~0ms for unchanged elements. If 1 out of 50 elements changes, only that one needs re-rendering. This would lower total time from 14ms to ~7ms at idle.

  2. Phase 1 — Value enum in stylesComputedStyle::compute() and parse_*() take &str and parse each property. Passing &Value instead would skip parsing. Potentially speeds up both style matching and VDOM eval.

  3. Phase 2 — InternedStr — string comparisons (type_name == "Button", key == "padding-top") happen thousands of times per frame. Replacing with u32 comparison would give 1.5-2× in VDOM and render paths.

  4. Phase 0.3 — flamegraph — confirm hypotheses with profiler measurements (perf record) before investing in optimization.