Files
ai-app/app/androidApp/src/main/kotlin/com/example/aiapp/BenchFixture.kt
T
irisandClaude Opus 5 77cee6a8fa The bench fixture streams a reply shaped like a real one, and keeps the run-on as stress
Iris, on the two findings from the incremental-text investigation:
"let's switch to new lines for the test, and also let's keep the single
line around for stress + could be something to try to optimize later."

The streamed tail now takes a blank line every 4-12 deltas, so it is 53
markdown blocks with a longest of 502 characters instead of one block of
14,888 -- against a measured p50 of 147 and a largest-ever 1,580 over
7,706 blocks of real assistant messages. Layer 1's streaming frame went
from p50 3.86ms / p90 8.65ms / worst 10.95ms to p50 2.20 / p90 5.90 /
worst 8.78.

The run-on message is kept as the first two backlog events, 14,824
characters in one block, just under text_cap's 16 KiB so it draws in
full. The *streaming* pathology stays in frame_profile.rs rather than the
fixture: it needs a growing block, and iterating on it there costs a
second instead of a two-minute phone run.

Adding it is purely additive -- the random state is saved and restored
around those two events, so every other backlog event is byte-identical.
That is not tidiness: the first attempt shifted the backlog and broke
`a_long_press_and_drag_selects_text`, which replays a real recording at
(300, 1000) and needs the content it was recorded against to still be
there. BACKLOG_COUNT is 3202 now, in generate.py, fixture.rs and
BenchFixture.kt, which split the file by line index.

And the answer to Iris's question, which the code already had: the newest
message does *not* cap. `build_row`'s `cap` is false for the live tail
because a row that grew while capped would appear to stop growing, and a
reply growing past the cap is never caught either since it grows through
apply_delta. So a streamed block's shaping cost has no ceiling -- ~29ms
per delta at 50k characters, ~58ms at 100k.

Recorded but not chased: the emulator's `stream: build p50` did not move
(10.4 -> 10.5ms) while layer 1's frame nearly halved, so most of a
streaming frame on a GPU path is the whole-arena primitive re-upload
layer 1 never performs -- 11,568 primitives rewritten per delta, with the
fling phase as the control at 0.4ms for the same primitives moved
through move_offsets.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 01:32:59 -04:00

109 lines
4.6 KiB
Kotlin

package com.example.aiapp
import android.content.Context
import java.util.concurrent.CopyOnWriteArrayList
/**
* P0's benchmark gate (see docs/RUST.md and the 2026-09-05 decision): an in-process fake of the
* backend, so the `bench` build type can drive a real session screen -- the real
* [TranscriptSource], the real fold, the real paging -- with no server and no network permission.
*
* Only ever installed when [BuildConfig.FIXTURE_MODE] is true (see [MainActivity]); everything else
* in this build compiles it in but never calls it, since Kotlin has no per-build-type source set
* that both [MainActivity] (which every variant compiles) and this can share without one.
*
* The design: [requestFromServer] and [Sse] talk to `https://$FIXTURE_HOST:$FIXTURE_PORT` through
* ordinary `java.net.URL`, exactly as they would talk to a real server. A
* [java.net.URLStreamHandlerFactory] registered once for the whole process intercepts every
* `https://` connection to that host and answers from this object's in-memory event log instead of
* opening a socket -- see BenchNetwork.kt. Everything above that (TranscriptSource, SessionScreen,
* the fold, uniqueItems) never learns the difference.
*/
object BenchFixture {
const val FIXTURE_HOST = "bench.fixture.invalid"
const val FIXTURE_PORT = 1
/** How many of the fixture's events are the opening backlog; see bench-fixture/README.md. */
private const val BACKLOG_COUNT = 3202
val settings = ServerSettings(FIXTURE_HOST, FIXTURE_PORT, "bench")
/** The session id every bench run opens; nothing else in this build ever mints one. */
const val SESSION_ID = "bench-fixture-session"
/**
* The whole transcript, seq order, growing as [pushLive] is called during the streaming phase.
* Read by both the REST page handler and the SSE handler, so a page requested mid- stream and a
* live frame agree on what has "already happened" -- the same thing a real server's own
* transcript file guarantees.
*/
private val log = CopyOnWriteArrayList<Pair<String, SeqEvent>>()
/** The events not yet appended to [log] -- the streaming phase's own source. */
private var streamTail: List<Pair<String, SeqEvent>> = emptyList()
private val images = mutableMapOf<String, ByteArray>()
@Volatile private var loaded = false
/**
* Parses the bundled fixture once. Safe to call more than once; only the first does anything.
*/
@Synchronized
fun ensureLoaded(context: Context) {
if (loaded) return
val lines =
context.assets.open("transcript.jsonl").bufferedReader().readLines().filter {
it.isNotBlank()
}
val parsed = lines.map { it to parseSeqEvent(it) }
log.addAll(parsed.take(BACKLOG_COUNT))
streamTail = parsed.drop(BACKLOG_COUNT)
for (name in listOf("bench1.png", "bench2.png")) {
images[name] = context.assets.open(name).readBytes()
}
loaded = true
}
/** The events the streaming phase has left to send. */
fun remainingStreamEvents(): Int = streamTail.size
/** Sends the next fixture event onto the live log, as a real SSE frame would arrive. */
fun pushNextLiveEvent(): Boolean {
val next = streamTail.firstOrNull() ?: return false
streamTail = streamTail.drop(1)
log.add(next)
return true
}
/** Undoes [pushNextLiveEvent] and reloads the opening backlog, for running the bench twice. */
@Synchronized
fun resetToBacklog(context: Context) {
loaded = false
log.clear()
ensureLoaded(context)
}
fun fileBytes(name: String): ByteArray? = images[name]
/**
* Raw JSON lines with seq > [after], in order -- what an `/events?after=` connection replays.
*/
fun linesAfter(after: Long): List<String> =
log.filter { it.second.seq > after }.map { it.first }
/**
* One REST page: [fetchTranscript]'s `before`/`limit`/`after`, against the growing log. Ignores
* `coalesce` -- the fixture's own deltas are already split the way a real reply streams, and
* what the benchmark exercises is the fold and the paging, not the server's row-joining, which
* client-core's own port tracks separately (CLIENT_CORE.md).
*/
fun page(before: Long?, limit: Int, after: Long?): List<String> {
val upper = before ?: (log.lastOrNull()?.second?.seq?.plus(1) ?: 1L)
val candidates = log.filter {
it.second.seq < upper && (after == null || it.second.seq > after)
}
return candidates.takeLast(limit).map { it.first }
}
}