EUI applications without a browser

10 — Budgets

Status: measured except where marked. cargo run --release -p xtask -- bench produces §2, §3 and §4 and exits non-zero when a budget is missed; CI runs it.

Soli's own benchmark document publishes the rows it loses. This one does the same: where a number comes from a run, it says so, and where it is still a goal, it says that instead. A budget that was never measured is a slogan.

1. Runtime — targets, enforced by cargo xtask bench once the client lands

MeasureBudget
Launch → first pixel< 80 ms
Idle, 200-node application0 % CPU, zero wakeups
Idle RSS, 200-node application< 25 MB
10 000-row virtualised table, RSS< 45 MB
10 000-row virtualised table, scroll60 fps, < 2 ms CPU per frame
10 000-row virtualised table, drag60 fps, < 2 ms CPU per frame
Client binary, stripped, 2 variable fonts included< 12 MB
MAX_ISLANDS — sessions one page may open (01 §2.7)8

MAX_ISLANDS is a ceiling and not a target: most pages want none and a page that wants one wants one. It is here because a tree is data — a view that derived an island per row would open a socket per row, and the reader whose machine pays for that is not the one who wrote the view. Past the ceiling a client opens no more sessions and leaves those islands showing what the page was rendered with, which is the same thing that happens when an island's session cannot be opened at all.

An edited field is what an input arms a deadline for: 06 §2's idle change, one per burst of typing, disarmed by the event it produces and by the blur or Enter that would have produced it first, and 03 §3's caret blink, twice a second while somebody is plainly there. Both stop — the change when it fires, the blink after ten seconds with nothing moving the caret — so the zero-wakeup line above holds with a field on the screen and a caret in it, which is the state a window left open on a form is actually in. A field no handler is listening to arms no change at all.

A drag frame costs what a wheel notch costs: one hit test, one search for the slot, at most one scroll step, and a relayout of the rows that have boxes — not of the ten thousand that do not. The thing that can be slow is the server's answer, and that is the honest limit here. A live reorder is a render and a diff per boundary crossed, which the one-answer-at-a-time rule of 06-events.md §2 caps at a few dozen a second; past that the client goes on following the hand, because nothing about the hand waits on the server, and the rows arrive behind it. That degrades; it does not break.

Measured on 2026-09-15, cargo build --release -p eui-client, x86-64 Linux, LTO, stripped: 16.28 MB (16 284 888 bytes) with the default features. The same build was 15.64 MB on 2026-09-08 and 12.13 MB on 2026-09-06; with --no-default-features it was 12.66 MB on 2026-09-08, and that one has not been re-measured since. Of that 2026-09-08 difference, 0.25 MB is the two pictures beside PNG (03 §1) — JPEG 157 KB (zune-jpeg) and WebP 64 KB (image-webp), each behind its own feature, both on by default. The rest is the accessibility stack — AccessKit and, on Linux, the AT-SPI bus client it needs (zbus), which brings its own async runtime beside tokio.

The 0.64 MB added since 2026-09-08 is scene: eui-shader, the scene pipeline and its targets. It is smaller than a whole language front end would suggest, because naga was already paid for — wgpu links it to translate the client's own shaders — and the verifier reuses that parser rather than bringing one of its own. A second verifier for a second kind of executable content is two and a half times what the two extra picture formats cost and under a quarter of what the accessibility adapter costs — which is the comparison that matters, because those two are levers and this one is not.

It was one per cent over in September's first week. It is a third over now, and the budget stays: what grew is what the client learned to do — sound and moving pictures decoded in the worker (03 §7, §8: symphonia's WAV, FLAC, MP3 and Vorbis, GIF and animated WebP), still pictures in the two other formats the web uses (03 §1), a symbols fallback face at 227 KB beside the two variable ones, the desktop theme's watcher, and the worker boundary itself. Naming the miss is the point of this document; the levers are known, measured, and none of them is free: the accessibility adapter (2.72 MB, by building without it), OGG and Vorbis (0.72 MB — symphonia reserves FFT twiddle tables up to 65 536 points, half a megabyte of zeroes in .data because they are Lazy statics rather than .bss), the four embedded faces (1.33 MB of .rodata), and the shader translator wgpu needs at runtime (naga, 876 KB of .text).

A face an application supplies (02 §5.1) does not move this line at all: it is an asset, it travels once per session and lives in the asset budget of 01 §2.2, and it is not in the binary. The lever it offers points the other way — a client that could rely on the application to send a face would need fewer of its own — but that is not a trade this version makes: the embedded faces are what makes an application that asks for nothing still draw, and what makes the same text shape the same way on every machine.

Scenes — measured where it can be, named where it cannot

A scene (03 §1.2) is the one node that can ask the client for an unbounded amount of work, so what it may ask for is written down rather than left to the application's good manners.

MeasureBudget
Shader source64 KiB (= MAX_CHUNK_BYTES)
Static steps per fragment invocation4 096 (= the VM's fuel)
Mesh vertices / indices65 536 / 196 608
Render targets alive per session8
Compiled modules held per window process32 (gpu::MAX_SCENE_MODULES), least recently drawn dropped first with its pipelines; its source is kept and compiled again when a scene names it
Target edge, before the device's own limit trims it4 096 device px (scene::MAX_EDGE), rounded up to a multiple of 64
Target memory, per scene4 B/px colour, +4 B/px per sample when msaa > 1, +4 B/px per sample for depth
Author's uniforms8 floats
Verifying a module< 2 ms — measured 81.5 µs
Decoding and checking a 60 000-vertex mesh (2.5 MB)< 20 ms — measured 739 µs
Scene frame, CPU0 beyond the list that was already painted — measured 0 ns, 29 ns against 39 ns for the same tree without the scene
Scene frame, driver processzero wakeups while the clock is the only thing moving
A scene that is not playing0 CPU, no pass encoded, nothing scheduled

The target quota is the one with a concrete attacker behind it: a virtualised list of ten thousand rows with a scene in each is cheap to send and would otherwise ask for ten thousand render targets. Past the quota a scene draws its own background, which is what one whose module has not compiled yet already does — so the application has no new failure to handle.

What the quota is worth, arithmetically. The edge row said 2 048 until 2026-09-20, when it was read against crates/eui-render/src/scene.rs:104 and found to be half the real number — which is a quarter of the real area, and area is what a texture costs. The formula is make_scene_target's, read off the three allocations it makes: a colour texture always, an MSAA colour texture only when samples > 1, and a depth texture only when the module wants one, the last two at sample_count samples each. So a scene at the default msaa: 4 with depth is 36 B/px, one without MSAA is 8 B/px, and colour alone is 4 B/px.

Eight targets at the full 4 096 × 4 096, with MSAA and depth, is therefore 4.8 GB and not the 1.2 GB the old row implied. That is the ceiling a device's own max_texture_dimension_2d is expected to cut into long before an application reaches it, and it is written here because a budget nobody can multiply out is not a budget. For scale, the demo's own scene node — 260 × 200 logical at 2×, quantised to 576 × 448 — is 9.3 MB with MSAA and depth, and 2.1 MB without.

The two zero lines are the ones that keep the paragraph below honest, and they are structural, not tuned: an animating scene is the same draw list redrawn with a later clock, so the frame costs the window a uniform write and costs the process that reads the server's bytes nothing at all. If a scene frame ever costs that process work, the design is wrong rather than the budget.

The three measured lines are the ones the argument rests on. The first two say a module and a mesh can be checked in the worker on the frame they arrive rather than over several: at 81 µs and 739 µs, both fit inside one frame with room to spare, so there is no need for either to be done in pieces. The third says what it looks like: a frame with a turning scene in it costs the paint nothing, because everything moving in it moves on the GPU from a clock, and the list is the same list.

Measured 2026-09-14, release, on the machine of §"Where the machine is". Debug numbers are not these numbers — the mesh decode reads 40 ms there, fifty-five times the release figure — which is why bench refuses to be quiet about the profile it was built in.

Still not measured by xtask bench, and named here rather than left implied: the per-frame GPU cost of a scene, the memory its targets hold, and what the feature adds to the binary. The first two want an adapter, which the bench does not assume.

The zero-wakeup line is an architectural consequence, not a tuning parameter: winit runs in ControlFlow::Wait and the client redraws only when a frame, an input event, or a window event asked it to. There is no render loop in the code to accidentally leave running.

What the process boundary costs — measured

The configuration that ships puts the driver in a confined worker process (08 §10), so the scroll budget above is met, or not, over a pipe. cargo run --release -p xtask -- bench measures both sides on the same 10 000-row table, x86-64 Linux, 2026-09-08:

MeasureIn this processThrough a worker
First paint (shaping)5.9 ms5.8 ms
Scroll step (median of ten, four runs)180–200 µs355–500 µs
of which the input round trip—77–104 µs
what the boundary adds to a paint—96–218 µs
bytes over the pipe per scroll step—27 B out, 47.1 KB back

Two round trips a frame — the input, then the paint — are most of the difference, and the draw list is 132 bytes a quad — 112 until the quad took on both ends of a transition (03 §5), 128 until it carried a bounce's height (03 §5, version 6; a gradient, 02 §5.3, rides in fields a solid box leaves empty) — sent in the shape the renderer uploads (the gallery is about 760 quads, its documentation dialog about 3 900). A frame owed to a spin, a transition or a glide alone is not sent at all: the window draws the last list again, and answers the driver's ticks itself meanwhile. xtask bench holds those frames to their budget: a repeated frame, a transition frame, a page transition frame and a glide frame each under 0.2 ms of driver time — they measure at tens of nanoseconds — and a glide with no layout after its first. A page transition is measured with both pages on screen and no layout at all: the one arriving was laid out once where it lands, and the one leaving is a picture. The rest is layout and paint, which the boundary does not change; the in-process figure is steady and the worker's is not, because it includes two process wake-ups. Both sides are inside the 2 ms budget, and folding an input into the paint that follows it would make it one wake-up a frame.

Going somewhere (03 §5.1, 06 §5). A page that leaves keeps its painting, and these are what stop that from being unbounded:

LimitValueWhy
Kept paintings per session1a second supersedes the first, so the memory is a constant and not a function of how fast a person can tap
Pairs resolved per change8a shared element is a slot in a small fixed table, as a scroll in flight is — and the table is four bits wide, with the arriving and departing pages holding a slot each
Leading-edge strip~20 logical pxa bezel's width: wide enough to find without looking, narrow enough that what it takes from the application is an edge nobody puts a control against

What is kept is the quads and not the nodes, so the cost is what was on screen rather than what the page was made of: a departing ten-thousand-row table costs its forty visible rows. A client MUST let a kept painting go rather than draw it wrong when the glyph atlas, the frame size or the viewer's palette has moved under it, and a page mid-transition owes the window no more than a transition frame does — the list does not change while it runs, so nothing is walked, repainted, or sent across a worker's pipe.

A scroll step was 1.75 ms in this process and 2.03 ms through a worker until 2026-09-08, when the row tops of a virtualised list stopped being rebuilt on every frame (04 §7): the tops are the rows' heights added up, and a scroll changes neither. set_scroll marks a node dirty::SCROLL rather than dirty::SELF for exactly that reason.

What a blurred frame costs

A blur (03 §2.1) is the one thing in the client that is not a single pass, so it is worth saying what it is allowed to cost. A frame with no blurred node in it — every frame of the applications measured above — takes none of this: no texture is allocated, no pipeline is built, and the frame is the one pass and one draw per scissor run it always was. The rows above are therefore unaffected, and are meant to stay that way.

The shaped-run cache

The text engine keeps shaped runs in a cache of 16 384 entries, keyed by the run, the face, the size and the width it was measured at. The size is a budget of its own: a page whose working set does not fit is not merely a cache miss, it is a guaranteed one, because an entry evicted before its next use will be shaped again — and layout measures the same run at several widths on the way to a line break.

A 630-line markdown document — 4 400 nodes, one per word, because a paragraph that flows across lines cannot be one text node (03 §1) — asks for 23 324 shapes against a 4 096-entry cache and 7 785 against this one: the difference is thrash, and it was 334 ms of layout against 137 ms in a release build. Eviction is first-in, first-out; with a cache that holds a page that is enough, and a page larger than this one would thrash again.

Two more things were found on the same page. A run that fits on one line is the same run at every width that holds it, so the engine answers a bounded request from the run's natural shape whenever that fits (Stats::reused): 7 785 shapes became 5 713. And a text leaf with no height of its own is as tall as its lines whatever room it was offered, so layout memoises every non-Exact height constraint as one: 53 477 uncached measures became 21 617. Together, release layout of that page went 334 → 108 ms; a debug build, which shapes text at a twentieth of the speed, 6.9 s → 1.9 s.

What remains is the shape of the work, not its cost: a paragraph that flows around a styled word is one node per word (03 §1), and a wrapping row measures every word for its line, its minimum and its cross size, once per pass of each ancestor. A page that long wants the protocol's own answer to long content — a windowed list (04 §7.1) laying out only the rows in view — rather than a faster full layout.

Measured 2026-09-09: RSS growth: session + layout, 10k rows 25.8 MB and driver RSS growth, table-10k with real text 17.7 MB, both against the 45 MB line above — the cache only fills as runs are shaped, so an idle application holds a few hundred entries, not sixteen thousand.

A frame that does carry one adds, per distinct radius, three passes over a snapshot of the region that asked — the union of the blurred rects grown by three standard deviations, clipped to the window — and one pass to take that snapshot. The snapshot is a texture the size of that region, rounded up to a multiple of 128 px a side, not of the framebuffer: a 360 px dialog is a few hundred kilobytes where a 4K window would be 33 MB. The rounding is what lets a region that slides or grows keep its textures, and the views and bind groups made with them, rather than allocating all of them again every frame of the movement; the region is drawn into the texture's corner and the reduction reads no further than its edge (gpu::tests::a_blur_region_that_moves_or_grows_a_little_keeps_its_textures, and the frosted-pane pixel tests of 09 §8). It is freed when a frame stops asking, so a closed dialog costs nothing, which is what keeps the idle RSS line above honest.

The reduction is chosen so the kernel is about fifteen taps whatever radius was asked for, so a wider blur buys a smaller texture rather than more samples and the cost does not grow with the radius. What it does grow with is the area, and a scrim covers the window: that is the case to measure before this line can stop saying "goal".

2. Wire — measured

Source: crates/eui-proto/tests/size_budget.rs, run with --nocapture. The subject is a 50-row × 4-column invoice table: 256 nodes, rows keyed, six shared style records. The HTML side is the same table with the Tailwind classes such a table really carries, unindented — the favourable case for HTML.

Bytes
Cell and header text (identical both ways)2 337
EUI total4 619
EUI, with the status column interned4 209
EUI, of which style records (once per session)396
EUI structure only2 282
HTML total14 362
HTML structure only12 025

Total 3.1×. Structure 5.3×. 8.9 bytes of structure per node.

The text is the data; neither side can compress it away, so the honest headline is the structural ratio. The regression budget is set at 4×, below the measured 5.3×, so a real regression trips it and ordinary drift does not.

OperationMeasuredBudget
Single-cell update22 B< 40 B
Reversing 50 keyed rows201 B< 300 B
Same reversal by re-mounting4 619 B—

That last pair is the case for MoveChild. Soli's current LiveView diff works on lines of an HTML string, so reordering a table has no cheap spelling at all; here it is 49 moves and 201 bytes.

What this does not claim

Bytes on the wire were never the main prize, and under compression they are not even a win: gzip the same table both ways and the HTML is smaller, 910 B to 1 491 B, because its bulk is one tag repeated fifty times and that is what a compressor is for. A 3.1× raw reduction matters on a link with no compression and nowhere else.

What does not change is the update — 22 B against a page HTML must send again — and the reason EUI exists, which is what the client does not do with those bytes: no tolerant parse, no selector matching, no cascade resolution, no reflow of an untyped tree, no JIT. Those are the numbers in §1, and they are the ones worth judging the project on once the client can be measured.

3. Decoder and session — measured

MeasureMeasuredBudget
Decode a 4.4 KB batch18 µs< 50 µs
Decode the 10 000-row mount, 770 KB3.1 ms< 20 ms
Apply that mount into a session, 50 002 nodes6.4 ms< 30 ms
Layout of the table, virtualised0.14 ms< 5 ms

4. Wire — the table, text included

The §2 figure is structure; a real mount carries the cells too. The 10 000-row invoice table mounts in 769.5 KB, 15.8 bytes per node with the text, of which about 9 are structure and the rest the data itself. Budget: 18 B per node.

5. Files

MeasureBudget
One transfer chunk256 KiB
One upload, ceiling64 MiB — a node's pick may ask for less
One upload, default16 MiB, when pick names no ceiling
One save256 MiB, whatever the server sends
Abort reason256 B
Memory held while a file uploadsat most twelve chunks (3 MiB), whatever the file weighs
Frames held while the socket is down256, and 1 MiB
Dialogs open per node1

The upload ceiling is the client's; the tree's is whatever max says, and it may only be lower. The save ceiling exists because the person chose where a file goes and not how much of their disk it may take.

Nothing here is a stream: a transfer belongs to its session and does not survive it, and the memory it costs is fixed by the chunk size rather than by the file. The disk runs two chunks ahead of the window, and the window takes a chunk only while the socket's unwritten backlog is under eight chunks — none at all while there is no socket — so the twelve are two read ahead, one being read, and at most nine written to the socket and not yet on the wire. It used to say two, counting only the first of those: the window drained the reader straight into an unbounded channel, and on a slow link the whole file waited there.

Where the machine is, and what it is held against

MeasureBudget
Nodes asking to be placed at once2, the rest ignored
Fastest interval1 s, whatever locate asks for
Coordinate resolution reported0.001°, about 110 m
Accuracy reported, floor100 m
Scans open per node1
Tags per scan1
Records kept off one tag64

A fix costs a radio rather than a timer, which is battery on the only two platforms that have one — so the floor is a second and not the hundred milliseconds a clock gets, and the count is two and not four. Neither is negotiable by the tree: a locate of 10 is a locate of 1000, silently, because a server that could argue about it would.

The resolution and the accuracy floor are in this table rather than in the client's judgement for the same reason every other number here is: a budget a reader cannot check is a promise, and this one is a privacy promise.

Assets

What 01 §2.2 calls the session's asset budget.

MeasureBudget
One asset16 MiB (MAX_ASSET_BYTES)
Asset store, per session: files kept and pictures decoded, together128 MiB (MAX_STORE_BYTES)
A picture's pixels, as heldno more than 1024 on a side (ATLAS_EDGE), 4 MiB
Fetches in flight, per origin4 (FETCHES_PER_ORIGIN)
Connecting, TLS handshake, sending the request10 s each (CONNECT_TIMEOUT)
Waiting for a byte of the response15 s without one (READ_IDLE)

A picture is held once: its pixels, shrunk to what the sheet takes, and its natural size, which is what the layout measures it by. The file of a PNG or a JPEG goes as soon as its pixels are held; a WebP's is kept, because a video node may ask for the same bytes as a moving picture. Fonts, chunks, meshes, modules, sounds and moving pictures are held as files.

Past the budget the store lets go of what the live tree does not name, least recently used first, and keeps the natural size of each picture it lets go of, so a page that names it again does not change shape while it is fetched again. What the tree does name is never let go — a page that shows more than the budget at once holds more than it — and it is what "remaining" is measured against: a fetch may bring back the budget less what the named assets already hold, and one whose Content-Length says more is abandoned before its body is read. Until 2026-09-24 the store had no budget at all, held every picture twice — file and full-size pixels — and a feed of 1024 × 768 photographs grew by 3.2 MB a picture for as long as it was open.

A fetch is queued, not started: four workers per origin take them in turn, each on one runtime it keeps, so a page of two hundred thumbnails is four connections at a time rather than two hundred threads and two hundred TLS handshakes at once. A server that stops talking is given up on — the limit is on silence, not on the whole transfer, since sixteen megabytes over a slow link take as long as they take — and a timeout is a failure like any other, tried twice more before it is final. A body that fits is read into a buffer reserved at its declared length.

Pictures, sounds and moving pictures are decoded on two threads the process keeps for it, never on the one that paints: a large JPEG, a long MP3 or a GIF used to hold the window — and, on a desktop, the lock the audio callback takes — for as long as it took, and a sound already playing ran dry. What a decode makes lands at the next tick or paint, which a driver with a decode in flight asks to be woken for every 8 ms and at no other time, and only the nodes that name it are measured again.

Pinned in crates/eui-client/tests/assets.rs: a_picture_is_decoded_off_the_painting_thread, a_server_that_stops_talking_is_given_up_on, a_page_of_pictures_is_fetched_by_a_few_workers, the_store_counts_what_it_holds_and_keeps_one_copy_of_a_still, the_store_lets_go_of_what_nothing_names_least_recently_used_first, the_driver_lets_go_of_a_picture_the_page_stopped_showing and a_fetch_past_the_room_left_is_abandoned.

Sound

MeasureBudget
Sources playing at once8, a ninth refused
Frames buffered ahead of the device200 ms
CPU with nothing loaded0 — the device is closed
Decoded samples a sound128 MiB — 5 min 50 s of stereo at 48 kHz
Decoded samples held, every sound together, per session256 MiB

Mixing is a multiply and an add per sample per source, with one linear interpolation for the rate; the cost is in the decode, which happens once per sound.

A sound is held decoded — f32, interleaved — because that is what makes mixing it cost nothing, and so the bound is on what the samples weigh rather than on how long the sound says it is. Frames alone were the bound until 2026-09-24: an hour of stereo, which let the largest asset there is (16 MB) expand to 1.38 GB before it was refused. A sound past 128 MiB is refused, not truncated (eui_audio::MAX_BYTES; pinned by a_sound_past_its_decoded_budget_is_refused in crates/eui-audio/tests/audio.rs).

A decoded sound is held while an audio node names it and not after: the mixer let go of a source when its node went, and the samples behind it stayed, so a playlist kept every track it had played. What the nodes of one tree name at once may weigh 256 MiB together; the sound that would take them past it is refused like one past its own ceiling (decoded_media_goes_when_no_node_names_it, decoded_media_past_the_sessions_room_is_refused, both in crates/eui-client/tests/driver.rs).

Moving pictures

MeasureBudget
Pixels a frame1920 × 1080
Frames a picture3 600
Decoded frames a picture96 MB
Decoded frames held, every picture together, per session192 MiB
Uploads per frame shown1, and none while the frame does not change
CPU while paused0 — nothing is scheduled

A picture is held, frames and size both, while a video node names it, and the pictures the nodes of one tree name may weigh 192 MiB together: a feed of fifty animated avatars used to keep every one it had shown, uncompressed, for as long as the tab was open. The picture that would take the total past it is refused, as one past its own 96 MB is (pinned by the same two tests as sound's).

Rendered from spec/10-budgets.md in the repository. The normative text is spec/; where it and a prose page disagree the spec wins, and that is a bug worth reporting. Edit this page