Skip to content

Zed GPUI — one draw(&Scene), eight primitives a shader can evaluate

Category: primitive scene. Last reviewed: August 23, 2026. Pinned at d71f1461.

GPUI's renderer seam is a single method taking a single value: a Scene of eight owned primitive kinds, each shaped so a fragment shader can evaluate it in one draw call. Text is not in that vocabulary at all — measurement and shaping are a separate platform service, and by the time a glyph reaches the scene it is an atlas tile.

FieldValue
LanguageRust
LicenseApache-2.0 (the gpui crate; crates/gpui/Cargo.toml)
Repositoryzed-industries/zed, crate crates/gpui
Documentationcrates/gpui/README.md; declared homepage gpui.rs
Categoryprimitive scene
Pinned revisiond71f1461045c098dc6ca6b1b5adcf1b8949722e8 (origin/main, 2026-08-11)
Target rangedesktop GPU only — macOS, Windows, Linux, and wasm. No CPU or character-cell target
Backends shippedMetal (metal_renderer.rs), Direct3D 11 (gpui_windows), wgpu (wgpu_renderer.rs, used by Linux and web)
Seam declarationPlatformWindow::draw(&self, scene: &Scene) in platform.rs

Overview

What it solves

GPUI is the framework Zed is written in, and its constraint is a text editor's: tens of thousands of glyphs per frame, with the whole UI — not only the buffer — on the same path. Its README opens:

GPUI is a hybrid immediate and retained mode, GPU accelerated, UI framework for Rust, designed to support a wide variety of applications.

crates/gpui/README.md

The "hybrid" is load-bearing here. Elements are re-run each frame (immediate), but the painted output is retained: Window::reuse_paint copies an unchanged element's primitives out of the previous frame's scene by index range (window.rs). That one feature dictates most of the seam's shape and is this survey's strongest evidence for friction §7.

Design philosophy

The scene is not a picture of what the widget layer meant but a description of what the GPU will be asked to do, in the smallest vocabulary that covers a real editor. Two rules follow, both visible throughout scene.rs:

  • A primitive is what one shader can evaluate. A rounded, bordered, gradient-filled rectangle is one Quad, because fs_quad in shaders.wgsl evaluates a signed-distance field with corner radii and per-edge border widths. Anything a shader cannot do — an arbitrary path — is tessellated to triangles by the framework before it reaches the scene.

  • A primitive is GPU memory. The structs are #[repr(C)], written straight into instance buffers — which is why the file defines its own bool:

    rust
    /// A boolean stored as a `u32` so that GPU-facing structs contain no
    /// compiler-inserted padding bytes, which would be undefined behavior to
    /// reinterpret as `&[u8]` when writing instance buffers. Guaranteed to be
    /// `0` or `1` by construction; shaders read it as a `u32`/`uint`.
    #[repr(transparent)]
    pub struct PaddedBool32(u32);

    scene.rs

How it works

The backend contract is three methods on PlatformWindow (platform.rs), of which only the first draws:

rust
fn draw(&self, scene: &Scene);
fn sprite_atlas(&self) -> Arc<dyn PlatformAtlas>;
fn is_subpixel_rendering_supported(&self) -> bool;

The value draw receives is a sum type of eight variants:

rust
pub enum Primitive {
    Shadow(Shadow),
    Quad(Quad),
    Path(Path<ScaledPixels>),
    Underline(Underline),
    MonochromeSprite(MonochromeSprite),
    SubpixelSprite(SubpixelSprite),
    PolychromeSprite(PolychromeSprite),
    Surface(PaintSurface),
}

scene.rs

Scene itself is not a Vec<Primitive>. It is one public vector per kind (quads, shadows, paths, underlines, three sprite vectors, surfaces), plus a private BoundsTree<ScaledPixels>, a layer_stack, and a Vec<PaintOperation> insertion history.

The enum is the insertion API (Scene::insert_primitive(impl Into<Primitive>)); the storage is struct-of-arrays because the renderer consumes Scene::batches(), an iterator of PrimitiveBatch values naming a kind and a contiguous index range — Quads(Range<usize>), MonochromeSprites { texture_id, range }. Every shipped renderer is one match over that iterator (metal_renderer.rs, wgpu_renderer.rs, directx_renderer.rs).

Sorting per-kind vectors destroys insertion order, so painter's order is reconstructed from an explicit key: insert_primitive takes one from a spatial index — BoundsTree::insert returns "one greater than the maximum ordering of any" intersecting bounds (bounds_tree.rs) — and Scene::finish sorts each vector by that order: DrawOrder. Non-overlapping primitives share an order and batch together; overlapping ones cannot.

Clipping is not an operation. Every primitive carries a content_mask: ContentMask<ScaledPixels> — a single Bounds — which the framework keeps as a stack in Window::with_content_mask, intersecting on push (window.rs). Scene::push_layer/pop_layer exist but are an ordering device: "a batch of geometry that are non-overlapping and have the same draw order."

Q1 — measurement units, and who answers

Measurement is a first-class service on two levels, and neither of them is the renderer.

  • PlatformTextSystem (platform.rs) is the per-OS trait — CoreText, DirectWrite, cosmic-text — with font_metrics, typographic_bounds, advance, glyph_for_char, glyph_raster_bounds, rasterize_glyph and layout_line(&self, text, font_size, runs) -> LineLayout.
  • TextSystem and WindowTextSystem (text_system.rs) are the framework's caching facade over it: em_advance, ch_advance, cap_height, x_height, ascent, descent, baseline_offset, shape_line, line_wrapper, plus an Arc-shared LineLayoutCache.

layout_line returns a LineLayout { font_size, width, ascent, descent, runs, len } where each ShapedGlyph carries id, position and index (line_layout.rs). There is no monospace assumption anywhere: a caller advances a pen by ShapedLine::width(), documented as "the glyph advance width computed by the text shaping system" (line.rs).

By the time the renderer is involved there is no text. Window::paint_glyph resolves a GlyphId to a raster, inserts it into the sprite atlas, and emits a MonochromeSprite/SubpixelSprite holding an AtlasTile (window.rs).

This is the fifth surveyed subject to keep measurement off the painter, so F1 holds. What it settles is only the placement question, and F2 is the reminder that placement is one decision of six. On the unit decision GPUI is the opposite of Slint: Slint lets each backend answer in its own Font::Length, while GPUI fixes one unit — Pixels, a newtype over f32 — for the framework, the text system and the scene alike, and converts to ScaledPixels exactly once, at paint time, by multiplying by the window's scale factor. Backend-chosen units are not the consensus; not putting measurement on the renderer is.

Q2 — is the contract stated in one place?

Yes, and this is the cleanest answer in the survey. The drawing contract is fn draw(&self, scene: &Scene). A backend author reads one enum with eight variants and writes eight match arms; Rust's exhaustiveness check enforces the surface, so there is nothing to probe for and nothing to discover at a call site. Compare isCanvas, which advertises five methods while the interpreter probes for four more with __traits(compiles) at each call site — five methods, eight kinds (friction §2).

Capability variation is handled by two mechanisms outside the primitive vocabulary:

  • A boolean query, consulted by the framework.PlatformWindow::is_subpixel_rendering_supported() returns false on macOS (gpui_macos/src/window.rs) and true on Windows (gpui_windows/src/window.rs). Window::should_use_subpixel_rendering also folds in window background opacity and the user's TextRenderingMode before deciding, so the Metal backend simply never receives that variant — its batch arm is PrimitiveBatch::SubpixelSprites { .. } => unreachable!().
  • Conditional compilation. PaintSurface's payload field is #[cfg(target_os = "macos")] pub image_buffer: CVPixelBuffer — on other platforms the variant exists but carries nothing.

NOTE

That is Qt's hasFeature idea with the negotiation moved up a level: the framework asks, once, and lowers to a different primitive; the backend is never asked to degrade anything. It supports F5's "floor" tier — seven of eight variants are mandatory — while showing that the refusable tier can be one query, not a bitmask.

Q3 — semantic widgets, or primitives?

Neither, and the axis this survey has been using does not cut here.

Nothing widget-shaped reaches the scene: no scrollbar, no text input, no button — Zed's scrollbars are elements in a separate ui crate that emit quads through Window::paint_quad. But the primitives are far from raw geometry:

rust
pub struct Quad {
    pub order: DrawOrder,
    pub border_style: BorderStyle,       // Solid | Dashed
    pub bounds: Bounds<ScaledPixels>,
    pub content_mask: ContentMask<ScaledPixels>,
    pub background: Background,          // solid | linear gradient | slash | checkerboard
    pub border_color: Hsla,
    pub corner_radii: Corners<ScaledPixels>,
    pub border_widths: Edges<ScaledPixels>,
}

scene.rs

Shadow carries blur_radius, corner_radii, element_bounds, element_corner_radii and an inset flag ("0 = drop shadow (rendered outside the element), 1 = inset shadow"). Underline carries thickness and wavy: PaddedBool32 — the same squiggle our LineStyle.wavy names.

The organising principle is what a fragment shader can evaluate in one pass. fs_quad computes a rounded-rect SDF from corner_radii and shades borders from border_widths (shaders.wgsl); a dashed border is a shader branch, not a different primitive. Where the shader runs out, the framework lowers: PathBuilder tessellates through lyon into PathVertex triangles (path_builder.rs), and paint_svg rasterises into the atlas.

So GPUI answers F4's "where does the lowering live" with nowhere: it picks a vocabulary that needs no lowering at the seam and pre-lowers the rest in the framework. That option exists only because every target is a GPU, so it is unavailable to sparkles:ui — our optional scrollbar primitive exists because a cell backend and a pixel backend genuinely disagree about what a scrollbar looks like, a situation GPUI designed itself out of rather than solved.

Q4 — command shape

A sum type at the API — Primitive, plus PaintOperation for the layer brackets:

rust
pub(crate) enum PaintOperation {
    Primitive(Primitive),
    StartLayer(Bounds<ScaledPixels>),
    EndLayer,
}

This is the second reifying subject after egui, and the one that reifies for a reason we share: Scene::replay(range, prev_scene) (Q7) needs commands to be values.

IMPORTANT

GPUI prices both sides of F3's live trade at once, and shows that the choice is per-layer rather than global. Its authoring type is a closed sum, like DrawOp. Its GPU-facing payload is the flat tag-plus-dead-fields record a closed sum is supposed to make unnecessary: Background is a #[repr(C)] record of tag: BackgroundTag, color_space, solid, gradient_angle_or_pattern_height, colors: [LinearColorStop; 2] and pad: u32 (color.rs), where a solid colour leaves five fields dead. That is deliberate: the struct is memcpy'd into an instance buffer for a shader that reads the tag and branches, and a discriminated union with per-variant payloads cannot be that. The two encodings are not competitors — the seam's SumType buys the comparable, walkable values a recorder needs, and a flat record buys a byte layout a shader can read. A Skia or Graphite path should expect one at whatever boundary faces a GPU, beneath the sum rather than instead of it.

The width side of that trade is where the two seams differ. GPUI's struct-of-arrays storage gives each kind its own vector, so a Quad costs a quad; a DrawOp costs the widest payload — TextRun — under a static assert(DrawOp.sizeof <= 64) budget, which is what makes a uniform DrawOp[] walkable and pairwise-comparable in the first place (friction §4).

Two second-order notes: reification plus batching costs an explicit order key (our DrawOp[] is ordered implicitly by array position — cheaper, and strictly unable to be reordered for batching); and Scene::len() is paint_operations.len(), not a primitive count, so that index ranges into the history stay meaningful across frames.

Q5 — sub-unit placement

Coordinates are continuous (Pixels, an f32 newtype), so GPUI never spells a position as a compass direction the way our rule op names a RuleEdge. What it does have — and this is the transferable part — is an explicit, role-differentiated rounding policy at the moment continuous logical units meet discrete device pixels (window.rs):

HelperRuleUsed for
snap_boundsround each edge to a device pixel, clamping so right >= leftquad bounds
snap_stroke"Rounds half-to-zero but clamps any non-zero input up to 1 dp so thin strokes do not disappear"border widths
cover_boundsfloor near edges, ceil far edges — "a strict superset of the raw region"content masks, shadow bounds
glyph quantisationSUBPIXEL_VARIANTS_X = 4, SUBPIXEL_VARIANTS_Y = 1glyph origin → atlas cache key

snap_stroke's clamp is the hairline rule stated as policy: a stroke the caller asked for never rounds away to nothing, only down to the thinnest thing the device has. That is exactly the guarantee our rule op provides by naming an edge — and it needs no enumerators, only a stroke concept plus a rounding rule per role.

GPUI is therefore F6's case in point rather than a counter-example: a float seam does not dissolve the sub-unit problem, it relocates it into a snapping policy that has to be written down. And it sharpens the second half of F6's answer — the fidelity to name is a rounding policy attached to the kind of thing being placed, and the device unit it rounds against is the scale factor the window already knows.

Q6 — resolved appearance, semantic role, or both?

Resolved only. Scene primitives carry Hsla and Background; element opacity is folded in at emission (quad.background.opacity(opacity), color.opacity(element_opacity)), and the scale factor is applied on the way in. There is no slot, role, class name or theme key anywhere in scene.rs — Zed's themes resolve entirely above GPUI, in separate crates.

Nothing re-resolves downstream because nothing could: every backend is a shader pipeline. That is the fifth subject to pay for one representation rather than two, and one more instance of F9's count — nobody else carries a resolved appearance and a semantic role together. Ours does: six of the eight payloads store a Slot beside the resolved colours their primitive paints from, which is friction §6. Reconstructing a Visual on demand instead of storing one makes that hedge cheaper without making it a decision. What GPUI confirms is why we pay it at all: the HTML interpreter re-resolves from the role to emit class names, and a re-resolving consumer is something none of these projects has.

Q7 — payload ownership, and outliving the frame

The strongest result in this deep-dive. Primitive has no lifetime parameter. Every variant owns its payload:

  • Quad, Shadow, Underline and all three sprite kinds are Copy plain data.
  • Path owns Vec<PathVertex<ScaledPixels>>.
  • PaintSurface owns a CVPixelBuffer.
  • Glyphs and images are not carried at all: a sprite holds an AtlasTile (texture_id, tile_id, padding, bounds) into a backend-owned atlas, populated through PlatformAtlas::get_or_insert_with(key, build) (platform.rs) — Slint's draw_cached_pixmap bargain, generalised to every raster payload and keyed by RenderGlyphParams, a hashable description of font, glyph, size, subpixel variant, scale factor and dilation.

The payoff is not thread-safety but frame reuse. An element whose input did not change is not re-painted; Window::reuse_paint splices its previous output forward:

rust
self.next_frame.scene.replay(
    range.start.scene_index..range.end.scene_index,
    &self.rendered_frame.scene,
);

window.rs, alongside self.text_system.reuse_layouts(...) for the matching shaped lines.

That is impossible with a frame-scoped payload, and it sharpens F8 in a way the count alone does not. Our seam is already on F8's side of the line: CmdBuffer.textRun copies the run into a frame arena, so nothing borrows from the caller and a scope source is safe. The arena is one of the three mechanisms F8 enumerates. What the copy does not buy is a retain boundary — the rule stated on the type is that an operation is valid while the buffer that built it is alive and unreset (friction §7), and reset() lands once a frame. So the cost is not hygiene, which the arena already pays for; it is that the biggest paint-time optimisation available to a retained-output frame loop sits on the far side of that reset. UI-O4 is open on exactly this question.

Q8 — can a backend ask the scene its extent?

No, and deliberately. Scene's public surface is clear, len, push_layer, pop_layer, insert_primitive, replay, finish, batches and the per-kind vectors — no bounds accessor. The BoundsTree that knows every primitive's bounds is private and exists only to assign draw order. Extent comes from the surface:

  • On-screen: PlatformWindow::draw(&Scene) — the window already knows its content_size() and scale_factor().
  • Offscreen, including golden-image tests: the size is a separate parameter on PlatformHeadlessRenderer::render_scene_to_image, whose signature is (&mut self, scene: &Scene, size: Size<DevicePixels>) (platform.rs) — and TestWindow computes it from the window bounds, not the scene (platform/test/window.rs).

GPUI is therefore one of the minority F7 identifies: it answers all three extent questions — surface, layout and ink — from somewhere other than the scene, and declines both sides of F7's maintained-versus-scanned axis by never asking the question. The consumer friction §8 names as needing a scene-derived extent — an offscreen golden — is exactly the case GPUI serves by passing the size in, which is worth weighing against the majority: our skia-canvas-render.d scans every operation's rect for that number, and the scan works only because a TextRun's rect.width happens to be its advance in cells.

Strengths

  • The contract is one method and one exhaustively-matched enum. A new backend is a match the compiler completes for you; there is no optional surface and no probing.
  • The vocabulary is chosen mechanically — evaluable by one shader — not by taste, which is why eight primitives cover a full editor UI with no escape hatch.
  • Commands own their payloads, buying cross-frame replay, per-kind batching and freedom from lifetime plumbing at once.
  • Raster ownership sits with the backend (PlatformAtlas), keyed by a hashable description, so the party that knows a payload's lifetime holds it.
  • The coarse/fine boundary is named explicitly (snap_bounds, snap_stroke, cover_bounds) rather than left to each backend's rounding.

Weaknesses

  • It cannot reach a non-GPU target. The vocabulary is what a shader can do; there is no software or cell backend and no obvious path to one.
  • The unit is fixed at Pixels framework-wide — a backend cannot answer in its own unit the way a Slint AbstractFont can.
  • Correctness of z-order depends on a spatial index. Two overlapping primitives of different kinds are ordered by an R-tree query; a bug there is a z-order bug with no local explanation.
  • The GPU ABI leaks into the APIPaddedBool32, explicit pad: u32 fields, a tag-plus-dead-fields Background.
  • Capability negotiation is ad hoc: one boolean plus cfg gating, with no declared feature set — the gap Qt fills with PaintEngineFeature.

Key design decisions and trade-offs

DecisionRationaleTrade-off
One seam method, draw(&self, scene: &Scene)Backend surface is a total function of one value; exhaustive matching enforces itThe scene type is the API; adding a primitive breaks every backend at compile time
Primitives = "what one shader can evaluate"No backend ever degrades; a rounded bordered gradient quad is one draw callRules out any non-GPU target
Sum type for insertion, struct-of-arrays for storageBatching by kind and by atlas texture is the point of the whole sceneRequires an explicit DrawOrder and a spatial index to preserve painter's order
GPU-facing structs are #[repr(C)] with padding fieldsInstance buffers are written as &[u8]; compiler padding would be UBThe tag-plus-dead-fields encoding a sum type was supposed to eliminate reappears
Primitives own their payloads; no lifetimesScene::replay can splice last frame's output forwardPath vertices are cloned per frame; the scene is heavier than a borrowed one
Rasters live in a backend-owned atlas, keyed by paramsThe party that knows a payload's lifetime owns it; glyph raster cost paid onceCache keys (RenderGlyphParams) must enumerate every rendering-relevant parameter
Clip travels as a content_mask field, not an opNo clip stack to mis-bracket; every primitive is self-describingOne rectangle only — no clip shapes, and the field is paid for on every primitive
Extent is never asked of the sceneThe surface chose its own size; offscreen callers pass it inAn offscreen "size to content" consumer must compute bounds itself

Bearing on the proposal

  1. Move the retain boundary, not the copy — the argument is frame reuse. The frame arena already satisfies F8; what friction §7 records is that the borrow expires at reset(). Scene::replay names the price of that expiry: no splicing an unchanged subtree's operations forward from the previous frame. That is the case UI-O4 should be decided on.
  2. Keep the sum type — and expect a flat record at the GPU boundary. This prices one side of F3's live trade: Background is a tagged struct with dead fields on purpose, being instance-buffer memory. DrawOp stays a closed sum for authoring, recording and comparison; whatever a Skia or Graphite path uploads is a separate, deliberately flat type underneath it, not a replacement for it.
  3. State the contract in one place, not five methods plus four probes.Friction §2 is confirmed by the cleanest counter-example here: eight variants, one match, compiler-enforced. Whatever we keep of the optional-primitive bargain belongs outside the drawing vocabulary — a query the framework consults, like is_subpixel_rendering_supported, not a __traits(compiles) at a call site.
  4. Replace RuleEdge with a stroke plus a per-role rounding policy.F6 asks for a named fidelity and a queried device unit; GPUI names three fidelities (snap_bounds/snap_stroke/cover_bounds) and gives the one that matters a rule we can copy verbatim: a non-zero stroke never rounds to zero. That is the whole content of ruleEndpoints' hairline degradation, without six enumerators.
  5. Adopt a backend-owned, key-addressed raster cache. PlatformAtlas::get_or_insert_with with a hashable params key generalises Slint's draw_cached_pixmap. Our Glyph op carries a dchar and leaves the raster backend-side already; the key idea is what F8 recommends for whatever image payload a Skia or Graphite path grows.
  6. Do not read GPUI as evidence that semantic ops are wrong. It has none, but only because it has designed away the situation that produces them: with every target a GPU, no primitive ever needs degrading. That does not generalise to a seam that must also paint into character cells, so it neither supports nor undermines friction §3 — it narrows the claim in F4 to "when all targets share a rendering model, the lowering can live nowhere, because a vocabulary exists that needs none."
  7. Note what reification costs once you want batching. If sparkles:ui ever sorts ops by kind (a plausible Skia/Graphite optimisation), it inherits GPUI's problem: an explicit draw order derived from overlap. Our implicit array order is a feature; it is worth recording as a constraint rather than rediscovering it.

WARNING

One nuance to carry into the proposal: GPUI keeps measurement off the renderer but fixes the unit at Pixels for the entire framework. The unanimity behind F1 is therefore about placement alone — "move measure off the canvas" is unanimous, while "with a backend-chosen unit" is Slint's answer to one of the five further decisions F2 enumerates, not the field's.

Sources

All pinned at d71f1461045c098dc6ca6b1b5adcf1b8949722e8 and verified to exist at that revision with git cat-file -e:

Related: the umbrella (Q1-Q8), the synthesis (F1-F12), the peers argued with — Slint, egui, Qt, Notcurses — the seam itself (canvas.d) and the friction log.