Compare commits
4
Commits
bfe93c4188
...
68b48cfd14
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
68b48cfd14 | ||
|
|
e6c884a0cd | ||
|
|
a6cb9a9082 | ||
|
|
0be6a571c4 |
No files matched your search
@@ -96,7 +96,11 @@ two icon buttons the same width without either being given one — and why
|
||||
:androidApp:compileDebugKotlin :androidApp:lintDebug
|
||||
:androidApp:testDebugUnitTest`. The unit tests are JVM-only and cover the
|
||||
syntax highlighter, the ANSI parser and the transcript cache — the app's
|
||||
pure logic with no Android in it.
|
||||
pure logic with no Android in it. Touching anything under `BenchFixture.kt`,
|
||||
`BenchNetwork.kt`, `BenchRun.kt` or the `bench` build type also needs
|
||||
`:androidApp:compileBenchKotlin :androidApp:lintBench` — a second build
|
||||
type compiles separately and lint has caught real bugs debug alone never
|
||||
would (see "Android Lint" below).
|
||||
- **Android Lint is not optional and is not run by a build.** It found a
|
||||
crash that had been shipping (`java.time` on a minSdk-24 app with
|
||||
desugaring off) and later a permission check that silently dropped every
|
||||
@@ -156,6 +160,27 @@ two icon buttons the same width without either being given one — and why
|
||||
|
||||
Each exists because something was invisible without it.
|
||||
|
||||
- **The `bench` build type and `app/bench-fixture/`** exist for P0 (RUST.md
|
||||
and DECISIONS.md's 2026-09-05 entries), the phone benchmark gate Iris
|
||||
asked for before porting continues: a deterministic, checked-in synthetic
|
||||
transcript (`app/bench-fixture/generate.py`, never a real one) that both
|
||||
this app and iris open with no server, so a frame-time comparison
|
||||
measures the renderer rather than the data. `./build-apk.sh bench` builds
|
||||
it — own application id (`com.example.aiapp.bench`) and label ("AI
|
||||
Sessions bench") so it installs beside a real enrollment rather than
|
||||
replacing it. Opening it goes straight to a session screen holding the
|
||||
fixture (no enrollment, no permission prompts) with a "Run benchmark"
|
||||
control beside "Copy" in session settings: it drives the same scroll loop
|
||||
and streaming phase `transcript-bench.sh`/`stream-bench.sh` drive over
|
||||
`ui-trace`, but in-process, since a real phone has no usable system
|
||||
tracing and no agent can drive one (this-machine-android's skill).
|
||||
`BenchFixture.kt`/`BenchNetwork.kt` fake the backend by installing a
|
||||
`URLStreamHandlerFactory` that answers `TranscriptSource`/`EventStream`'s
|
||||
requests from an in-memory copy of the fixture instead of opening a
|
||||
socket — so the fold, the paging and `uniqueItems` under test are the
|
||||
screen's real ones, never a shortcut built just for this. The report
|
||||
gains a `bench:` section (process CPU time, peak RSS, battery current) on
|
||||
every build, empty except when `BenchRun.kt` filled it in.
|
||||
- **`app/ui-sandbox.sh`** — a second `ai-server` with its own `$HOME`, config
|
||||
and data directory, holding eight invented Claude Code transcripts and a
|
||||
`claude` that is two lines of shell. **That isolation is the point**: the
|
||||
|
||||
@@ -102,6 +102,16 @@ android {
|
||||
targetSdk = 37
|
||||
versionCode = 1
|
||||
versionName = "1.0"
|
||||
// Read by MainActivity to decide, at startup, whether this is the P0 benchmark build
|
||||
// (docs/RUST.md's P0 box) rather than the app somebody enrolled. False everywhere except
|
||||
// the `bench` build type below, which overrides it.
|
||||
buildConfigField("boolean", "FIXTURE_MODE", "false")
|
||||
}
|
||||
buildFeatures {
|
||||
// Only for FIXTURE_MODE above; nothing else here reaches for generated BuildConfig fields.
|
||||
buildConfig = true
|
||||
// Only for the bench build type's resValue("string", "app_name", ...) below.
|
||||
resValues = true
|
||||
}
|
||||
packaging {
|
||||
resources { excludes += "/META-INF/{AL2.0,LGPL2.1}" }
|
||||
@@ -131,6 +141,35 @@ android {
|
||||
isMinifyEnabled = false
|
||||
if (keystore != null) signingConfig = signingConfigs.getByName("release")
|
||||
}
|
||||
// P0's benchmark build (docs/RUST.md, docs/DECISIONS.md's 2026-09-05 entry): release
|
||||
// optimisations so a frame time measured here means what release means everywhere else in
|
||||
// this project, its own application id so it installs beside a real enrollment rather than
|
||||
// replacing it, and FIXTURE_MODE so MainActivity opens straight onto the fixture session
|
||||
// instead of asking to be enrolled. Signed with the same key as release -- it never talks
|
||||
// to a real backend, so there is no CA of its own to mismatch, and a second keystore would
|
||||
// be one more secret to keep off this machine's shared mount for no benefit.
|
||||
create("bench") {
|
||||
initWith(getByName("release"))
|
||||
// :link (wg-app-link) has no "bench" build type of its own -- it is a library shared
|
||||
// with dev-updater and has no reason to know this project invented one -- so this says
|
||||
// which of its build types to link against instead.
|
||||
matchingFallbacks += listOf("release")
|
||||
applicationIdSuffix = ".bench"
|
||||
// "AI Sessions bench" everywhere the OS shows the app's name (launcher, recents,
|
||||
// Settings): this resValue overrides res/values/strings.xml's app_name for this
|
||||
// build type alone, and AndroidManifest.xml's android:label reads @string/app_name
|
||||
// rather than a literal so a build type can override it without touching the
|
||||
// manifest.
|
||||
resValue("string", "app_name", "AI Sessions bench")
|
||||
buildConfigField("boolean", "FIXTURE_MODE", "true")
|
||||
if (keystore != null) signingConfig = signingConfigs.getByName("release")
|
||||
}
|
||||
}
|
||||
sourceSets {
|
||||
// The fixture both bench builds (this one and iris's) open with; see
|
||||
// app/bench-fixture/README.md. Read directly from its own directory rather than copied
|
||||
// into androidApp/src -- one file to keep in sync with the generator, not two.
|
||||
getByName("bench").assets.directories.add("../bench-fixture/assets")
|
||||
}
|
||||
compileOptions {
|
||||
sourceCompatibility = JavaVersion.VERSION_21
|
||||
|
||||
@@ -25,7 +25,7 @@
|
||||
the fix is a judgement about how this app should look. Drop this
|
||||
suppression when a real icon lands. -->
|
||||
<application
|
||||
android:label="AI Sessions"
|
||||
android:label="@string/app_name"
|
||||
android:allowBackup="true"
|
||||
android:theme="@android:style/Theme.Material.Light.NoActionBar"
|
||||
tools:ignore="MissingApplicationIcon">
|
||||
|
||||
@@ -0,0 +1,108 @@
|
||||
package com.example.aiapp
|
||||
|
||||
import android.content.Context
|
||||
import java.util.concurrent.CopyOnWriteArrayList
|
||||
|
||||
/**
|
||||
* P0's benchmark gate (see docs/RUST.md and docs/DECISIONS.md's 2026-09-05 entry): an in-process
|
||||
* fake of the backend, so the `bench` build type can drive a real session screen -- the real
|
||||
* [TranscriptSource], the real fold, the real paging -- with no server and no network permission.
|
||||
*
|
||||
* Only ever installed when [BuildConfig.FIXTURE_MODE] is true (see [MainActivity]); everything else
|
||||
* in this build compiles it in but never calls it, since Kotlin has no per-build-type source set
|
||||
* that both [MainActivity] (which every variant compiles) and this can share without one.
|
||||
*
|
||||
* The design: [requestFromServer] and [Sse] talk to `https://$FIXTURE_HOST:$FIXTURE_PORT` through
|
||||
* ordinary `java.net.URL`, exactly as they would talk to a real server. A
|
||||
* [java.net.URLStreamHandlerFactory] registered once for the whole process intercepts every
|
||||
* `https://` connection to that host and answers from this object's in-memory event log instead of
|
||||
* opening a socket -- see BenchNetwork.kt. Everything above that (TranscriptSource, SessionScreen,
|
||||
* the fold, uniqueItems) never learns the difference.
|
||||
*/
|
||||
object BenchFixture {
|
||||
const val FIXTURE_HOST = "bench.fixture.invalid"
|
||||
const val FIXTURE_PORT = 1
|
||||
|
||||
/** How many of the fixture's events are the opening backlog; see bench-fixture/README.md. */
|
||||
private const val BACKLOG_COUNT = 3200
|
||||
|
||||
val settings = ServerSettings(FIXTURE_HOST, FIXTURE_PORT, "bench")
|
||||
|
||||
/** The session id every bench run opens; nothing else in this build ever mints one. */
|
||||
const val SESSION_ID = "bench-fixture-session"
|
||||
|
||||
/**
|
||||
* The whole transcript, seq order, growing as [pushLive] is called during the streaming phase.
|
||||
* Read by both the REST page handler and the SSE handler, so a page requested mid- stream and a
|
||||
* live frame agree on what has "already happened" -- the same thing a real server's own
|
||||
* transcript file guarantees.
|
||||
*/
|
||||
private val log = CopyOnWriteArrayList<Pair<String, SeqEvent>>()
|
||||
|
||||
/** The events not yet appended to [log] -- the streaming phase's own source. */
|
||||
private var streamTail: List<Pair<String, SeqEvent>> = emptyList()
|
||||
|
||||
private val images = mutableMapOf<String, ByteArray>()
|
||||
|
||||
@Volatile private var loaded = false
|
||||
|
||||
/**
|
||||
* Parses the bundled fixture once. Safe to call more than once; only the first does anything.
|
||||
*/
|
||||
@Synchronized
|
||||
fun ensureLoaded(context: Context) {
|
||||
if (loaded) return
|
||||
val lines =
|
||||
context.assets.open("transcript.jsonl").bufferedReader().readLines().filter {
|
||||
it.isNotBlank()
|
||||
}
|
||||
val parsed = lines.map { it to parseSeqEvent(it) }
|
||||
log.addAll(parsed.take(BACKLOG_COUNT))
|
||||
streamTail = parsed.drop(BACKLOG_COUNT)
|
||||
for (name in listOf("bench1.png", "bench2.png")) {
|
||||
images[name] = context.assets.open(name).readBytes()
|
||||
}
|
||||
loaded = true
|
||||
}
|
||||
|
||||
/** The events the streaming phase has left to send. */
|
||||
fun remainingStreamEvents(): Int = streamTail.size
|
||||
|
||||
/** Sends the next fixture event onto the live log, as a real SSE frame would arrive. */
|
||||
fun pushNextLiveEvent(): Boolean {
|
||||
val next = streamTail.firstOrNull() ?: return false
|
||||
streamTail = streamTail.drop(1)
|
||||
log.add(next)
|
||||
return true
|
||||
}
|
||||
|
||||
/** Undoes [pushNextLiveEvent] and reloads the opening backlog, for running the bench twice. */
|
||||
@Synchronized
|
||||
fun resetToBacklog(context: Context) {
|
||||
loaded = false
|
||||
log.clear()
|
||||
ensureLoaded(context)
|
||||
}
|
||||
|
||||
fun fileBytes(name: String): ByteArray? = images[name]
|
||||
|
||||
/**
|
||||
* Raw JSON lines with seq > [after], in order -- what an `/events?after=` connection replays.
|
||||
*/
|
||||
fun linesAfter(after: Long): List<String> =
|
||||
log.filter { it.second.seq > after }.map { it.first }
|
||||
|
||||
/**
|
||||
* One REST page: [fetchTranscript]'s `before`/`limit`/`after`, against the growing log. Ignores
|
||||
* `coalesce` -- the fixture's own deltas are already split the way a real reply streams, and
|
||||
* what the benchmark exercises is the fold and the paging, not the server's row-joining, which
|
||||
* client-core's own port tracks separately (CLIENT_CORE.md).
|
||||
*/
|
||||
fun page(before: Long?, limit: Int, after: Long?): List<String> {
|
||||
val upper = before ?: (log.lastOrNull()?.second?.seq?.plus(1) ?: 1L)
|
||||
val candidates = log.filter {
|
||||
it.second.seq < upper && (after == null || it.second.seq > after)
|
||||
}
|
||||
return candidates.takeLast(limit).map { it.first }
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,181 @@
|
||||
package com.example.aiapp
|
||||
|
||||
import java.io.ByteArrayInputStream
|
||||
import java.io.IOException
|
||||
import java.io.InputStream
|
||||
import java.io.PipedInputStream
|
||||
import java.io.PipedOutputStream
|
||||
import java.net.HttpURLConnection
|
||||
import java.net.URL
|
||||
import java.net.URLStreamHandler
|
||||
import java.net.URLStreamHandlerFactory
|
||||
import java.security.Principal
|
||||
import java.security.cert.Certificate
|
||||
import javax.net.ssl.HttpsURLConnection
|
||||
import javax.net.ssl.SSLPeerUnverifiedException
|
||||
import org.json.JSONArray
|
||||
|
||||
/**
|
||||
* Installs the process-wide interception [BenchFixture] needs. Idempotent and safe to call more
|
||||
* than once; the JDK only allows [URL.setURLStreamHandlerFactory] to be called successfully once
|
||||
* per process, and a second real call throws -- so this guards it rather than relying on every
|
||||
* caller to remember.
|
||||
*
|
||||
* Scoped to [BenchFixture.FIXTURE_HOST]: any other `https://` URL falls through to the platform's
|
||||
* ordinary handler, so this only ever changes behaviour for the one host the bench build invents.
|
||||
*/
|
||||
@Synchronized
|
||||
fun installFixtureNetworkOnce() {
|
||||
if (installed) return
|
||||
installed = true
|
||||
URL.setURLStreamHandlerFactory(
|
||||
URLStreamHandlerFactory { protocol ->
|
||||
if (protocol != "https") null
|
||||
else
|
||||
object : URLStreamHandler() {
|
||||
override fun openConnection(url: URL): HttpURLConnection =
|
||||
if (url.host == BenchFixture.FIXTURE_HOST) FixtureConnection(url)
|
||||
else
|
||||
// The bench build makes no other https call -- this factory is
|
||||
// installed only in FIXTURE_MODE (MainActivity) -- so there is
|
||||
// deliberately no delegate to a platform handler here: once a
|
||||
// URLStreamHandlerFactory is installed there is no supported way to
|
||||
// ask the JDK for its own default handler back, and re-entering this
|
||||
// same factory for the fallback would recurse forever rather than
|
||||
// reach one.
|
||||
throw java.io.IOException(
|
||||
"bench build's fixture network has no route to https host " +
|
||||
"${url.host} -- only ${BenchFixture.FIXTURE_HOST} is served"
|
||||
)
|
||||
}
|
||||
}
|
||||
)
|
||||
}
|
||||
|
||||
private var installed = false
|
||||
|
||||
/**
|
||||
* Answers one request against [BenchFixture] instead of opening a socket. Implements just enough of
|
||||
* [HttpsURLConnection] for [requestFromServer] and [Sse] to work unmodified: both only call
|
||||
* `connect`/`disconnect`, set a handful of request properties they never need answered, and read
|
||||
* `responseCode` and `inputStream`.
|
||||
*/
|
||||
private class FixtureConnection(url: URL) : HttpsURLConnection(url) {
|
||||
private var input: InputStream? = null
|
||||
private var writer: Thread? = null
|
||||
|
||||
override fun connect() {
|
||||
if (input != null) return
|
||||
input = route(url.path, url.query)
|
||||
}
|
||||
|
||||
override fun disconnect() {
|
||||
writer?.interrupt()
|
||||
try {
|
||||
input?.close()
|
||||
} catch (_: IOException) {}
|
||||
}
|
||||
|
||||
override fun usingProxy() = false
|
||||
|
||||
override fun getResponseCode(): Int {
|
||||
connect()
|
||||
return 200
|
||||
}
|
||||
|
||||
override fun getInputStream(): InputStream {
|
||||
connect()
|
||||
return input!!
|
||||
}
|
||||
|
||||
override fun getErrorStream(): InputStream? = null
|
||||
|
||||
// Nothing here reads any of these; implemented only because HttpsURLConnection declares them
|
||||
// abstract. A fixture never negotiates real TLS, so each says exactly that rather than
|
||||
// fabricating a plausible-looking certificate.
|
||||
override fun getCipherSuite() = "none (bench fixture, no TLS)"
|
||||
|
||||
override fun getLocalCertificates(): Array<Certificate>? = null
|
||||
|
||||
override fun getServerCertificates(): Array<Certificate> =
|
||||
throw SSLPeerUnverifiedException("bench fixture connection presents no certificate")
|
||||
|
||||
override fun getPeerPrincipal(): Principal =
|
||||
throw SSLPeerUnverifiedException("bench fixture connection presents no certificate")
|
||||
|
||||
override fun getLocalPrincipal(): Principal? = null
|
||||
|
||||
/**
|
||||
* [path] is `/sessions/{id}/...`; everything else this build's fixture is asked for is a bug.
|
||||
*/
|
||||
private fun route(path: String, query: String?): InputStream {
|
||||
val params =
|
||||
(query ?: "")
|
||||
.split("&")
|
||||
.filter { it.contains('=') }
|
||||
.associate {
|
||||
val (k, v) = it.split("=", limit = 2)
|
||||
k to java.net.URLDecoder.decode(v, "UTF-8")
|
||||
}
|
||||
return when {
|
||||
path.endsWith("/transcript") -> {
|
||||
val lines =
|
||||
BenchFixture.page(
|
||||
before = params["before"]?.toLongOrNull(),
|
||||
limit = params["limit"]?.toIntOrNull() ?: 80,
|
||||
after = params["after"]?.toLongOrNull(),
|
||||
)
|
||||
val body = JSONArray(lines.map { org.json.JSONObject(it) })
|
||||
ByteArrayInputStream(body.toString().toByteArray())
|
||||
}
|
||||
path.endsWith("/events") -> openEventsStream(params["after"]?.toLongOrNull() ?: 0L)
|
||||
path.contains("/files/") -> {
|
||||
val name = path.substringAfterLast("/files/")
|
||||
val bytes =
|
||||
BenchFixture.fileBytes(name)
|
||||
?: throw IOException("bench fixture has no file named $name")
|
||||
ByteArrayInputStream(bytes)
|
||||
}
|
||||
else -> throw IOException("bench fixture has no route for $path")
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* A live SSE body: [BenchFixture.linesAfter] replayed immediately, then polled every 50ms for
|
||||
* anything [BenchFixture.pushNextLiveEvent] has added since -- the same shape a real backend's
|
||||
* backlog-then-follow gives [Sse], just polled instead of woken, which is a fixture's business
|
||||
* rather than something worth a condition variable for.
|
||||
*/
|
||||
private fun openEventsStream(after: Long): InputStream {
|
||||
val pipeIn = PipedInputStream(1 shl 16)
|
||||
val pipeOut = PipedOutputStream(pipeIn)
|
||||
var sent = after
|
||||
val thread = Thread {
|
||||
try {
|
||||
while (!Thread.currentThread().isInterrupted) {
|
||||
val fresh = BenchFixture.linesAfter(sent)
|
||||
for (line in fresh) {
|
||||
pipeOut.write("data: $line\n\n".toByteArray())
|
||||
pipeOut.flush()
|
||||
sent = org.json.JSONObject(line).getLong("seq")
|
||||
}
|
||||
Thread.sleep(50)
|
||||
}
|
||||
} catch (_: InterruptedException) {
|
||||
// disconnect() -- the ordinary way this ends.
|
||||
} catch (_: IOException) {
|
||||
// The reader side (Sse) closed its end.
|
||||
} finally {
|
||||
try {
|
||||
pipeOut.close()
|
||||
} catch (_: IOException) {}
|
||||
}
|
||||
}
|
||||
.also {
|
||||
it.isDaemon = true
|
||||
it.start()
|
||||
}
|
||||
writer = thread
|
||||
return pipeIn
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,143 @@
|
||||
package com.example.aiapp
|
||||
|
||||
import android.content.Context
|
||||
import android.os.BatteryManager
|
||||
import android.os.Process
|
||||
import androidx.compose.animation.core.tween
|
||||
import androidx.compose.foundation.gestures.animateScrollBy
|
||||
import androidx.compose.foundation.lazy.LazyListState
|
||||
import java.io.File
|
||||
import kotlinx.coroutines.CoroutineScope
|
||||
import kotlinx.coroutines.delay
|
||||
import kotlinx.coroutines.isActive
|
||||
import kotlinx.coroutines.launch
|
||||
|
||||
/**
|
||||
* P0's scripted benchmark, run in-process instead of by a shell script: the phone has no usable
|
||||
* system tracing (this-machine-android's skill) and no agent can drive it, so the same scroll loop
|
||||
* and streaming phase `transcript-bench.sh`/`stream-bench.sh` drive over `ui-trace` are reproduced
|
||||
* here against [LazyListState] and [BenchFixture] directly. Only reachable from the `bench` build
|
||||
* (see [SessionSettingsDialog]'s `onRunBenchmark`), but compiled into every build for the reason
|
||||
* [BenchFixture]'s doc comment gives.
|
||||
*/
|
||||
object BenchRun {
|
||||
/** transcript-bench.sh's default: 6 cycles of 4 swipes each, 900px over 200ms, 500ms apart. */
|
||||
private const val CYCLES = 6
|
||||
private const val SWIPE_PX = 900f
|
||||
private const val SWIPE_MS = 200
|
||||
private const val SWIPE_PAUSE_MS = 500L
|
||||
|
||||
/** stream-bench.sh's shape: a real reply arrives as many small deltas, not one big write. */
|
||||
private const val STREAM_EVENTS_PER_SEC = 20
|
||||
private const val STREAM_SECONDS = 20
|
||||
|
||||
/**
|
||||
* Scrolls, then streams, then returns the extra report lines P0 asked for (CPU time, peak RSS,
|
||||
* battery current) -- [FrameStats] and [DebugStats] are reset first, exactly as
|
||||
* `copyRenderReport` resets them, so the two accountings cover the same stretch of work.
|
||||
*/
|
||||
suspend fun run(
|
||||
context: Context,
|
||||
scope: CoroutineScope,
|
||||
listState: LazyListState,
|
||||
): List<String> {
|
||||
FrameStats.reset()
|
||||
DebugStats.reset()
|
||||
val cpuStartMs = Process.getElapsedCpuTime()
|
||||
|
||||
val battery = BatterySampler(context)
|
||||
// Launched in the caller's scope rather than a fresh coroutineScope{} here, which would
|
||||
// suspend this function until the sampler job ended -- and it only ends when told to.
|
||||
val samplerJob = scope.launch {
|
||||
while (isActive) {
|
||||
battery.sample()
|
||||
delay(1000)
|
||||
}
|
||||
}
|
||||
|
||||
// The swipe loop: transcript-bench.sh's four swipes per cycle are two drags toward newer
|
||||
// content and two back, so a cycle returns to where it started and the whole loop measures
|
||||
// steady-state scrolling rather than travelling somewhere new each time.
|
||||
repeat(CYCLES) {
|
||||
repeat(2) {
|
||||
listState.animateScrollBy(SWIPE_PX, tween(SWIPE_MS))
|
||||
delay(SWIPE_PAUSE_MS)
|
||||
}
|
||||
repeat(2) {
|
||||
listState.animateScrollBy(-SWIPE_PX, tween(SWIPE_MS))
|
||||
delay(SWIPE_PAUSE_MS)
|
||||
}
|
||||
}
|
||||
|
||||
// Pinned to the newest end before streaming starts, the way stream-bench.sh's "Jump to
|
||||
// latest" tap is -- a reply streamed into a list parked further back arrives off-screen and
|
||||
// the report would show nothing happened.
|
||||
listState.scrollToItem(0)
|
||||
|
||||
var sent = 0
|
||||
val total = STREAM_EVENTS_PER_SEC * STREAM_SECONDS
|
||||
while (sent < total && BenchFixture.remainingStreamEvents() > 0) {
|
||||
BenchFixture.pushNextLiveEvent()
|
||||
sent++
|
||||
delay(1000L / STREAM_EVENTS_PER_SEC)
|
||||
}
|
||||
// Lets the last few deltas land and draw before the report is read.
|
||||
delay(300)
|
||||
|
||||
samplerJob.cancel()
|
||||
val cpuMs = Process.getElapsedCpuTime() - cpuStartMs
|
||||
val rssLine = peakRssLine()
|
||||
val batteryLine = battery.finish()
|
||||
|
||||
return listOf(
|
||||
" scroll: $CYCLES cycles (${CYCLES * 4} swipes), streamed $sent/$total fixture events",
|
||||
" process CPU time over this run: ${cpuMs}ms",
|
||||
rssLine,
|
||||
batteryLine,
|
||||
)
|
||||
}
|
||||
|
||||
/** VmHWM from /proc/self/status: the process's high-water mark, in kB, since it started. */
|
||||
private fun peakRssLine(): String {
|
||||
val kb =
|
||||
try {
|
||||
File("/proc/self/status")
|
||||
.readLines()
|
||||
.firstOrNull { it.startsWith("VmHWM:") }
|
||||
?.trim()
|
||||
?.removePrefix("VmHWM:")
|
||||
?.trim()
|
||||
?.removeSuffix("kB")
|
||||
?.trim()
|
||||
?.toLongOrNull()
|
||||
} catch (_: Exception) {
|
||||
null
|
||||
}
|
||||
return " peak RSS: " +
|
||||
(kb?.let { "${it}kB" } ?: "unavailable (/proc/self/status unreadable)")
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Samples [BatteryManager.BATTERY_PROPERTY_CURRENT_NOW] (microamps) once a second for the length of
|
||||
* a run. The property returns `Int.MIN_VALUE` on hardware that does not support it -- most
|
||||
* emulators -- and that is reported as "unavailable" rather than folded into an average with the
|
||||
* real samples, which would silently understate every number after it. See UI_RULES: never present
|
||||
* an inferred value as a measured one.
|
||||
*/
|
||||
private class BatterySampler(context: Context) {
|
||||
private val manager = context.getSystemService(BatteryManager::class.java)
|
||||
private val samples = mutableListOf<Int>()
|
||||
|
||||
fun sample() {
|
||||
val value = manager?.getIntProperty(BatteryManager.BATTERY_PROPERTY_CURRENT_NOW)
|
||||
if (value != null && value != Int.MIN_VALUE) samples.add(value)
|
||||
}
|
||||
|
||||
fun finish(): String {
|
||||
if (samples.isEmpty()) return " battery current: unavailable on this device"
|
||||
val meanUa = samples.sum() / samples.size
|
||||
return " battery current: mean ${meanUa}µA over ${samples.size} samples" +
|
||||
" (min ${samples.min()}, max ${samples.max()})"
|
||||
}
|
||||
}
|
||||
@@ -127,6 +127,13 @@ fun debugReport(
|
||||
frames: List<String>,
|
||||
accounting: List<String>,
|
||||
crash: String?,
|
||||
/**
|
||||
* P0's benchmark-only measurements (process CPU time, peak RSS, battery current) -- empty on
|
||||
* every path but [BenchRun.runP0Benchmark], which is the only caller that has them. A section
|
||||
* heading only appears when there is something to put under it, so an ordinary copy from the
|
||||
* render-report button reads exactly as it did before this existed.
|
||||
*/
|
||||
extra: List<String> = emptyList(),
|
||||
): String = buildString {
|
||||
appendLine("ai-app render report")
|
||||
appendLine(device)
|
||||
@@ -152,6 +159,11 @@ fun debugReport(
|
||||
appendLine("work since this was last copied:")
|
||||
val work = DebugStats.lines()
|
||||
if (work.isEmpty()) appendLine(" nothing recorded") else work.forEach { appendLine(it) }
|
||||
if (extra.isNotEmpty()) {
|
||||
appendLine()
|
||||
appendLine("bench:")
|
||||
extra.forEach { appendLine(it) }
|
||||
}
|
||||
}
|
||||
|
||||
/** Puts [text] on the clipboard under [label], which is what the system offers as its name. */
|
||||
|
||||
@@ -66,6 +66,35 @@ class MainActivity : ComponentActivity() {
|
||||
// Transparent status bar on every version; the Surface below paints through underneath it
|
||||
// and content insets itself. Same reasoning as dev-updater's MainActivity.
|
||||
enableEdgeToEdge()
|
||||
|
||||
// The `bench` build's entire purpose (P0, docs/RUST.md): open straight onto the session
|
||||
// screen against BenchFixture's in-process fake backend, with no enrollment, no network
|
||||
// permission, and no notification prompt -- none of them mean anything with no server and
|
||||
// no real device to notify. See BenchFixture.kt and BenchNetwork.kt for how a screen built
|
||||
// to talk to a real backend is made to talk to this instead. Still needs the same
|
||||
// status/navigation-bar padding the ordinary flow below applies: edge-to-edge is the
|
||||
// platform's own default from Android 15 on this app's targetSdk, with or without the call
|
||||
// above, so skipping the padding here put the header's own buttons under the status bar --
|
||||
// there to look at, but not there for `ui-trace`'s tap-by-label to land on.
|
||||
if (BuildConfig.FIXTURE_MODE) {
|
||||
installFixtureNetworkOnce()
|
||||
BenchFixture.ensureLoaded(this)
|
||||
setContent {
|
||||
MaterialTheme(colorScheme = AiAppColors) {
|
||||
Surface(modifier = Modifier.fillMaxSize()) {
|
||||
Box(Modifier.fillMaxSize().statusBarsPadding().navigationBarsPadding()) {
|
||||
SessionScreen(
|
||||
settings = BenchFixture.settings,
|
||||
summary = benchSessionSummary(),
|
||||
onBack = { finish() },
|
||||
onFiles = {},
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
return
|
||||
}
|
||||
// Dark status-bar icons only over a light background, decided from the scheme rather than
|
||||
// fixed. It was hardcoded to `true`, which was right against the default light surface and
|
||||
// became unreadable the moment the app wore Catppuccin Mocha.
|
||||
@@ -147,6 +176,26 @@ class MainActivity : ComponentActivity() {
|
||||
}
|
||||
}
|
||||
|
||||
/** The one session the `bench` build ever shows -- BenchFixture's session id, nothing else. */
|
||||
private fun benchSessionSummary() =
|
||||
SessionSummary(
|
||||
id = BenchFixture.SESSION_ID,
|
||||
setup = "bench",
|
||||
setupName = "bench",
|
||||
provider = "bench",
|
||||
title = "P0 benchmark",
|
||||
model = null,
|
||||
keepsOwnTranscript = false,
|
||||
permissionMode = null,
|
||||
imported = false,
|
||||
notify = false,
|
||||
cwd = null,
|
||||
contextTokens = null,
|
||||
maxImageEdge = null,
|
||||
status = "idle",
|
||||
lastActivity = 0.0,
|
||||
)
|
||||
|
||||
// launchMode="singleTop": an enrollment scan, or a notification tapped while the app is open,
|
||||
// lands here rather than in a second activity instance.
|
||||
override fun onNewIntent(intent: Intent) {
|
||||
|
||||
@@ -1191,7 +1191,12 @@ fun SessionScreen(
|
||||
// the bench scripts keep working when this moves again. They pressed it at a hand-measured
|
||||
// coordinate until 2026-09-03, and anything that moved the header made that tap land on
|
||||
// whatever now sat there -- reporting a number that was never measured.
|
||||
val copyRenderReport = {
|
||||
// Shared by the ordinary "Copy" button and (bench build only) "Run benchmark": what differs
|
||||
// between them is only whether there is a [extra] section, built by BenchRun.run beforehand --
|
||||
// everything about assembling, copying and logging the report is exactly the same act either
|
||||
// way, and a second copy of it beside `onRunBenchmark` below would be the two silently
|
||||
// disagreeing about what "the report" contains the first time either one changed.
|
||||
fun buildAndCopyReport(extra: List<String> = emptyList()) {
|
||||
val report =
|
||||
debugReport(
|
||||
device =
|
||||
@@ -1213,6 +1218,7 @@ fun SessionScreen(
|
||||
accounting =
|
||||
FrameStats.drawPhase().let { (nanos, count) -> drawAccounting(nanos, count) },
|
||||
crash = lastCrash(context),
|
||||
extra = extra,
|
||||
)
|
||||
context.copyToClipboard("ai-app render report", report)
|
||||
// Also to the log, so a session driving the app over adb can read the same report the
|
||||
@@ -1226,6 +1232,20 @@ fun SessionScreen(
|
||||
DebugStats.reset()
|
||||
Toast.makeText(context, "Copied render report", Toast.LENGTH_SHORT).show()
|
||||
}
|
||||
val copyRenderReport = { buildAndCopyReport() }
|
||||
// Bench build only: P0's scripted scroll-and-stream benchmark (BenchRun.kt), against the
|
||||
// fixture session opened below instead of a real server. Null everywhere else -- see
|
||||
// [SessionSettingsDialog]'s onRunBenchmark.
|
||||
val runBenchmark: (() -> Unit)? =
|
||||
if (BuildConfig.FIXTURE_MODE) {
|
||||
{
|
||||
settingsOpen = false
|
||||
scope.launch {
|
||||
val extra = BenchRun.run(context, scope, listState)
|
||||
buildAndCopyReport(extra)
|
||||
}
|
||||
}
|
||||
} else null
|
||||
Box(Modifier.fillMaxSize()) {
|
||||
Column(Modifier.fillMaxSize()) {
|
||||
Row(
|
||||
@@ -1869,6 +1889,7 @@ fun SessionScreen(
|
||||
},
|
||||
onDismiss = { settingsOpen = false },
|
||||
onCopyRenderReport = copyRenderReport,
|
||||
onRunBenchmark = runBenchmark,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -66,6 +66,13 @@ fun SessionSettingsDialog(
|
||||
* measures is that screen's own state.
|
||||
*/
|
||||
onCopyRenderReport: () -> Unit,
|
||||
/**
|
||||
* Runs P0's scripted scroll-and-stream benchmark and copies the extended report, or null on
|
||||
* every build but `bench` -- see [BuildConfig.FIXTURE_MODE] and BenchRun.kt. Null rather than
|
||||
* always-present-but-disabled: this has no meaning at all outside the bench build, and a
|
||||
* control with nothing behind it on every other build is not a state worth drawing.
|
||||
*/
|
||||
onRunBenchmark: (() -> Unit)? = null,
|
||||
) {
|
||||
val scope = rememberCoroutineScope()
|
||||
var name by remember(sessionId) { mutableStateOf(title) }
|
||||
@@ -317,6 +324,21 @@ fun SessionSettingsDialog(
|
||||
Text("Render timings", modifier = Modifier.weight(1f))
|
||||
TextButton(onClick = onCopyRenderReport) { Text("Copy") }
|
||||
}
|
||||
// Bench-build only: see [onRunBenchmark]. Named exactly "Run benchmark" because
|
||||
// ui-trace and the emulator smoke run find it by that label, the same way every
|
||||
// other control here is found -- see AGENTS.md's "Driving the UI".
|
||||
onRunBenchmark?.let { run ->
|
||||
Spacer(Modifier.height(8.dp))
|
||||
Row(
|
||||
verticalAlignment = Alignment.CenterVertically,
|
||||
modifier = Modifier.fillMaxWidth(),
|
||||
) {
|
||||
Glyph(SPEED_GLYPH, colour = MaterialTheme.colorScheme.onSurface)
|
||||
Spacer(Modifier.width(8.dp))
|
||||
Text("P0 benchmark", modifier = Modifier.weight(1f))
|
||||
TextButton(onClick = run) { Text("Run benchmark") }
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
// Disabled rather than absent while there is nothing to save: a button that comes and goes
|
||||
|
||||
@@ -0,0 +1,6 @@
|
||||
<?xml version="1.0" encoding="utf-8"?>
|
||||
<resources>
|
||||
<!-- Overridden by the `bench` build type's resValue (build.gradle.kts) to "AI Sessions bench",
|
||||
so the two are never mistaken for each other in the launcher or in Settings. -->
|
||||
<string name="app_name">AI Sessions</string>
|
||||
</resources>
|
||||
@@ -0,0 +1,28 @@
|
||||
# The P0 benchmark fixture
|
||||
|
||||
`transcript.jsonl` is a synthetic transcript in the app's own event model (the JSON lines
|
||||
`GET /sessions/{id}/transcript` returns; see `Events.kt`'s `parseSeqEvent` and
|
||||
`server/src/session/driver.rs`) -- never a real one. It is what both the Compose `bench` build
|
||||
and iris's bench build open with no server, so the two apps draw exactly the same content and a
|
||||
frame-time comparison is measuring the renderer rather than the data.
|
||||
|
||||
Generated by `./generate.py` (Python stdlib only, seeded -- `SEED = 20260905` -- so re-running it
|
||||
reproduces the same file byte for byte). It writes into `assets/` -- a separate directory from this
|
||||
script and README, because the Compose `bench` build type points its own asset source set straight
|
||||
at `assets/` (`app/androidApp/build.gradle.kts`'s `sourceSets { getByName("bench") }`), and a Python
|
||||
script and a markdown file have no business inside an APK:
|
||||
|
||||
- `transcript.jsonl` -- 3,601 events. The first 3,200 (`BACKLOG_COUNT`) are the scrolled-back
|
||||
history the benchmark opens with: user turns, tool calls with kilobyte-scale input/output,
|
||||
assistant replies built from headings, bold/italic/inline code, a link, fenced code blocks that
|
||||
rotate through rust/kotlin/python/sh/json/toml, a markdown table, two embedded images, and
|
||||
periodic `usageDelta`/`compacted` events. The remaining 400 (`STREAM_COUNT`) are not part of the
|
||||
opening window -- both bench harnesses replay them at a fixed rate (20/s) through the same live
|
||||
fold path a real SSE reply arrives on, which is P0's "streaming phase."
|
||||
- `bench1.png`, `bench2.png` -- tiny (8x8) flat-colour PNGs, base64-free on disk but served the
|
||||
same way a real attachment is (`GET /sessions/{id}/files/{name}`), referenced by the two
|
||||
`"type":"image"` events in the transcript.
|
||||
|
||||
Regenerate after changing the shape (a new event type, a different backlog/stream split) with
|
||||
`./generate.py`, and commit the result -- it is checked in rather than generated at build time so
|
||||
both apps' bench builds embed the identical bytes without needing this script at build time.
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 74 B |
Binary file not shown.
|
After Width: | Height: | Size: 74 B |
File diff suppressed because it is too large.
Load diff
Executable
+186
@@ -0,0 +1,186 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Generates transcript.jsonl -- the synthetic fixture P0's benchmark opens in both apps.
|
||||
|
||||
Deterministic (fixed seed), so a Compose bench APK and an iris bench APK draw byte-identical
|
||||
content: the point of the fixture is a like-for-like comparison, not a realistic one.
|
||||
|
||||
Never a real transcript -- see AGENTS.md's ui-sandbox.sh, which this borrows its vocabulary
|
||||
style from (headings, code fences, a table, a link) rather than reusing its Claude-Code JSONL
|
||||
shape. This file's shape is the *app's own event model* instead: one JSON object per line,
|
||||
matching what GET /sessions/{id}/transcript returns and what Events.kt's parseSeqEvent reads
|
||||
(server/src/session/driver.rs is the source of truth for the field names).
|
||||
|
||||
./generate.py writes transcript.jsonl and bench1.png/bench2.png here
|
||||
|
||||
BACKLOG_COUNT events (seq 1..BACKLOG_COUNT) are the scrolled-back history the benchmark opens
|
||||
with. A further STREAM_COUNT events (seq BACKLOG_COUNT+1..) are not part of the opening window;
|
||||
both bench harnesses replay them at a fixed rate as the "streaming reply" phase, appended through
|
||||
the same live path a real SSE reply arrives on. Keeping both halves in one file means one
|
||||
generator and one seed to keep in sync, rather than two fixtures that can drift apart.
|
||||
"""
|
||||
import base64
|
||||
import json
|
||||
import random
|
||||
import struct
|
||||
import zlib
|
||||
from pathlib import Path
|
||||
|
||||
SEED = 20260905
|
||||
BACKLOG_COUNT = 3200
|
||||
STREAM_COUNT = 400
|
||||
HERE = Path(__file__).resolve().parent / "assets"
|
||||
|
||||
random.seed(SEED)
|
||||
|
||||
LANGUAGES = ["rust", "kotlin", "python", "sh", "json", "toml"]
|
||||
|
||||
CODE_SNIPPETS = {
|
||||
"rust": '''fn fold_event(items: Vec<Item>, seq: u64) -> Vec<Item> {
|
||||
// a comment worth keeping: this is the fold the app's own screen runs
|
||||
let mut out = items;
|
||||
out.push(Item::new(seq));
|
||||
out
|
||||
}''',
|
||||
"kotlin": '''fun foldEvent(items: List<TranscriptItem>, entry: SeqEvent): List<TranscriptItem> {
|
||||
// mirrors the server's own event model, one item per line
|
||||
return items + TranscriptItem.from(entry)
|
||||
}''',
|
||||
"python": '''def render_report(frames, cpu_ms, rss_kb):
|
||||
# printed for a human to paste back, so every number carries its unit
|
||||
return f"{frames} frames, {cpu_ms}ms cpu, {rss_kb}kb peak rss"''',
|
||||
"sh": '''#!/bin/sh
|
||||
# scripted scroll loop, the shape transcript-bench.sh drives on a phone
|
||||
for i in $(seq 1 24); do
|
||||
ui-trace record --do "swipe 540 700 540 1600 200"
|
||||
done''',
|
||||
"json": '{"seq": 1, "type": "status", "state": "running"}',
|
||||
"toml": '''[package]
|
||||
name = "bench-fixture"
|
||||
version = "0.1.0"''',
|
||||
}
|
||||
|
||||
HEADINGS = [
|
||||
"## Plan",
|
||||
"## What changed",
|
||||
"## Why this approach",
|
||||
"### Open questions",
|
||||
"## Results",
|
||||
]
|
||||
|
||||
WORDS = (
|
||||
"session render report frame budget scroll transcript fold event cache "
|
||||
"cursor probe stream backlog swipe fixture bench compose iris widget layout "
|
||||
"measure place draw tool call token context window anchor"
|
||||
).split()
|
||||
|
||||
|
||||
def paragraph(n=24):
|
||||
words = [random.choice(WORDS) for _ in range(n)]
|
||||
words[0] = words[0].capitalize()
|
||||
text = " ".join(words) + "."
|
||||
# Sprinkle markdown inline spans so the syntax highlighter/markdown parser sees a real mix.
|
||||
text = text.replace(" fold ", " **fold** ", 1)
|
||||
text = text.replace(" cursor ", " *cursor* ", 1)
|
||||
text = text.replace(" cache ", " `cache` ", 1)
|
||||
if "bench" in text:
|
||||
text = text.replace(
|
||||
" bench ", " [bench](https://example.com/bench) ", 1
|
||||
)
|
||||
return text
|
||||
|
||||
|
||||
def make_png(rgb, size=8):
|
||||
"""A tiny, valid PNG -- flat colour, no external dependency."""
|
||||
|
||||
def chunk(tag, data):
|
||||
c = tag + data
|
||||
return struct.pack(">I", len(data)) + c + struct.pack(">I", zlib.crc32(c))
|
||||
|
||||
sig = b"\x89PNG\r\n\x1a\n"
|
||||
ihdr = struct.pack(">IIBBBBB", size, size, 8, 2, 0, 0, 0)
|
||||
raw = b""
|
||||
for _ in range(size):
|
||||
raw += b"\x00" + bytes(rgb) * size
|
||||
idat = zlib.compress(raw)
|
||||
return sig + chunk(b"IHDR", ihdr) + chunk(b"IDAT", idat) + chunk(b"IEND", b"")
|
||||
|
||||
|
||||
def main():
|
||||
HERE.mkdir(exist_ok=True)
|
||||
lines = []
|
||||
seq = 1
|
||||
ts = 1_788_000_000.0
|
||||
|
||||
def emit(type_, **fields):
|
||||
nonlocal seq, ts
|
||||
obj = {"seq": seq, "ts": round(ts, 3), "type": type_}
|
||||
obj.update(fields)
|
||||
lines.append(json.dumps(obj, separators=(",", ":")))
|
||||
seq += 1
|
||||
ts += random.uniform(0.05, 2.0)
|
||||
|
||||
emit("status", state="running")
|
||||
emit("settings", model="bench-model", permissionMode="auto")
|
||||
|
||||
image_refs = []
|
||||
turn = 0
|
||||
while seq <= BACKLOG_COUNT:
|
||||
turn += 1
|
||||
emit("userMessage", text=f"Turn {turn}: {paragraph(12)}", id=None, attachments=[])
|
||||
|
||||
# A tool call with kilobyte-scale input/output every few turns.
|
||||
if turn % 3 == 0:
|
||||
tool_id = f"tool-{turn}"
|
||||
big_input = json.dumps({"path": f"/repo/file_{turn}.rs", "content": paragraph(400)})
|
||||
emit("toolStart", id=tool_id, tool="Edit", input=big_input)
|
||||
big_output = "\n".join(paragraph(60) for _ in range(20))
|
||||
emit("toolUpdate", id=tool_id, output=big_output[: len(big_output) // 2])
|
||||
emit("toolEnd", id=tool_id, output=big_output)
|
||||
|
||||
# A reply: a heading, prose, a fenced block in a rotating language, a table, then deltas.
|
||||
emit("assistantText", delta=f"{random.choice(HEADINGS)}\n\n")
|
||||
emit("assistantText", delta=paragraph(30) + "\n\n")
|
||||
lang = LANGUAGES[turn % len(LANGUAGES)]
|
||||
emit("assistantText", delta=f"```{lang}\n{CODE_SNIPPETS[lang]}\n```\n\n")
|
||||
if turn % 5 == 0:
|
||||
emit(
|
||||
"assistantText",
|
||||
delta="| column | value |\n|---|---|\n| a | " + paragraph(3) + " |\n\n",
|
||||
)
|
||||
# A run of small deltas -- the shape a live reply actually streams in.
|
||||
for _ in range(random.randint(3, 8)):
|
||||
emit("assistantText", delta=paragraph(6) + " ")
|
||||
|
||||
# A couple of images, base64 PNGs, the way a real transcript embeds a screenshot.
|
||||
if turn in (10, 40):
|
||||
ref = f"bench{len(image_refs) + 1}.png"
|
||||
image_refs.append(ref)
|
||||
emit("image", ref=ref, about=None)
|
||||
|
||||
emit("usageDelta", tokens=random.randint(200, 4000), context=random.randint(2000, 180000))
|
||||
|
||||
if turn % 15 == 0:
|
||||
emit(
|
||||
"compacted",
|
||||
preTokens=180000,
|
||||
postTokens=20000,
|
||||
trigger="auto",
|
||||
)
|
||||
|
||||
# The streaming-phase tail: one long reply, built entirely from text deltas, the shape a
|
||||
# bench harness replays at a fixed events/sec through the live fold path.
|
||||
emit("userMessage", text="One more, streamed live for the benchmark's timing phase.", id=None, attachments=[])
|
||||
while seq <= BACKLOG_COUNT + STREAM_COUNT:
|
||||
emit("assistantText", delta=paragraph(5) + " ")
|
||||
emit("status", state="idle")
|
||||
|
||||
(HERE / "transcript.jsonl").write_text("\n".join(lines) + "\n")
|
||||
|
||||
(HERE / "bench1.png").write_bytes(make_png((220, 90, 90)))
|
||||
(HERE / "bench2.png").write_bytes(make_png((90, 150, 220)))
|
||||
|
||||
print(f"wrote {len(lines)} events ({BACKLOG_COUNT} backlog + {STREAM_COUNT} stream) to transcript.jsonl")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
+9
-3
@@ -4,6 +4,11 @@
|
||||
# ./build-apk.sh the release build, signed (what the phone runs)
|
||||
# ./build-apk.sh debug the debug build, for reproducing something the
|
||||
# emulator scripts would build anyway
|
||||
# ./build-apk.sh bench P0's benchmark build (own app id, "AI Sessions
|
||||
# bench" label, opens straight onto the fixture
|
||||
# session -- see docs/RUST.md's P0 box and
|
||||
# app/bench-fixture/README.md). Signed the same
|
||||
# as release; never touches the CA it pins.
|
||||
#
|
||||
# Dev Updater's `.dev-updater.ron` at the checkout root spells these out as
|
||||
# build modes, one command line each; it passes nothing else, so the word
|
||||
@@ -25,8 +30,9 @@ VARIANT=${1:-release}
|
||||
case "$VARIANT" in
|
||||
release) TASK=assembleRelease ;;
|
||||
debug) TASK=assembleDebug ;;
|
||||
bench) TASK=assembleBench ;;
|
||||
*)
|
||||
echo "build-apk.sh: unknown variant '$VARIANT' (release, debug)" >&2
|
||||
echo "build-apk.sh: unknown variant '$VARIANT' (release, debug, bench)" >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
@@ -81,7 +87,7 @@ fi
|
||||
# uninstalling it first: the signatures differ, and Android refuses to
|
||||
# update across them.
|
||||
KEYSTORE="${AI_APP_KEYSTORE:-${XDG_CONFIG_HOME:-$HOME/.config}/ai-app/release.jks}"
|
||||
if [ "$VARIANT" = release ] && [ ! -f "$KEYSTORE" ]; then
|
||||
if { [ "$VARIANT" = release ] || [ "$VARIANT" = bench ]; } && [ ! -f "$KEYSTORE" ]; then
|
||||
KEYTOOL="${JAVA_HOME:+$JAVA_HOME/bin/keytool}"
|
||||
KEYTOOL="${KEYTOOL:-keytool}"
|
||||
if ! command -v "$KEYTOOL" >/dev/null 2>&1; then
|
||||
@@ -97,7 +103,7 @@ if [ "$VARIANT" = release ] && [ ! -f "$KEYSTORE" ]; then
|
||||
-keyalg RSA -keysize 2048 -validity 10000 \
|
||||
-storepass "$PASSWORD" -keypass "$PASSWORD" -dname "CN=ai-app" >/dev/null 2>&1)
|
||||
fi
|
||||
if [ "$VARIANT" = release ]; then
|
||||
if [ "$VARIANT" = release ] || [ "$VARIANT" = bench ]; then
|
||||
AI_APP_KEYSTORE="$KEYSTORE"
|
||||
AI_APP_KEYSTORE_PASSWORD=$(cat "$KEYSTORE.password")
|
||||
export AI_APP_KEYSTORE AI_APP_KEYSTORE_PASSWORD
|
||||
|
||||
@@ -7,6 +7,49 @@ marked **DEFERRED** are ones the agent chose not to decide alone.
|
||||
|
||||
## 2026-09-05
|
||||
|
||||
- **P0's Compose half is built and smoke-tested on the emulator** — the
|
||||
`bench` build type, the shared `app/bench-fixture/` transcript, and an
|
||||
in-process fake backend (`BenchFixture.kt`/`BenchNetwork.kt`) that
|
||||
answers `TranscriptSource`/`EventStream` from an in-memory event log
|
||||
instead of a real server, so the fold and paging under test are the real
|
||||
ones. Full account, the smoke run's report, and what is deliberately
|
||||
left (the iris half, the real on-phone runs) are in RUST.md's P0 box.
|
||||
Not a decision to review so much as the gate itself now being runnable —
|
||||
flagged here because it is the first half of something Iris explicitly
|
||||
asked to see before P1.
|
||||
|
||||
- **The intermittent touch-scroll dropout is root-caused and fixed: a
|
||||
missed `ACTION_DOWN` hit-test, not the previously-suspected coalesced
|
||||
first `ACTION_MOVE`.** Diagnosed by temporary logcat tracing of every
|
||||
touch event, `DragArbiter` state transition and `Selection::drag`
|
||||
dispatch (removed once confirmed), reproduced on this checkout's own
|
||||
emulator against a real sandbox session. The trace showed the actual
|
||||
mechanism: a gesture's `ACTION_DOWN` lands wherever the finger actually
|
||||
is, which is not guaranteed to fall inside the same row-local sensor
|
||||
region a later `ACTION_MOVE` in the same gesture lands in (a row's own
|
||||
padding/gap, or its non-selectable sender-name header, is
|
||||
pointer-transparent to `iris::sense::CursorSense`). When that happens,
|
||||
the widget that ends up handling the gesture never saw `PressStart`, so
|
||||
`DragArbiter` sits in `Idle` — which answers every subsequent frame with
|
||||
`Undecided` and has no way to tell "no press is happening" from "a press
|
||||
is happening but I missed its start," so it never recovers on its own
|
||||
for the rest of that gesture. One real trace showed exactly this: touch
|
||||
`Down`/`Move`/`Up` all delivered correctly, but zero `PressStart`
|
||||
reaching the arbiter, `state=Idle` unchanged from first frame to last.
|
||||
Fixed at the call site that has the context to recover
|
||||
(`iris::transcript_ui::selection::Selection::drag`,
|
||||
`iris/transcript-ui/src/selection.rs`): a new `DragArbiter::is_idle()`
|
||||
(`iris/src/sense.rs`) lets it notice a `Pressing` frame arriving with the
|
||||
arbiter still `Idle` — which can only mean a missed `PressStart`, since a
|
||||
`Pressing` sense requires the button to genuinely be down — and start the
|
||||
press there instead of where it was missed. Three new unit tests in
|
||||
`sense.rs`'s `drag_arbiter_tests` and one in `transcript-ui`'s
|
||||
`selection::tests` (the latter fails on the code before this fix).
|
||||
Commit follows. Not the same failure the earlier pass's `DECISIONS.md`
|
||||
DEFERRED item speculated about (a coalesced first `ACTION_MOVE` skipping
|
||||
slop detection) — that hypothesis is now ruled out; the arbiter's own
|
||||
slop/long-press logic was never wrong. RUST.md's I5 box,
|
||||
"Touch-scroll dropout root-caused, 2026-09-05" has the full trace.
|
||||
- **P0, a phone benchmark gate before any porting, asked for by Iris
|
||||
2026-09-05**: "before P1 I'd like to see benchmarks & also maybe stress
|
||||
test on my own phone ... If it doesn't match compose reasonably well then
|
||||
|
||||
+216
@@ -36,6 +36,29 @@ session spending an afternoon on them again.
|
||||
|
||||
## Where things stand (2026-09-05)
|
||||
|
||||
- **The intermittent touch-scroll dropout is root-caused and fixed,
|
||||
2026-09-05.** Not the previously-suspected coalesced first
|
||||
`ACTION_MOVE` (ruled out) -- a gesture's `ACTION_DOWN` can land on a
|
||||
row's own padding/gap or its header, which `CursorSense` has no sensor
|
||||
over, so the widget that ends up handling the gesture only ever sees
|
||||
`Pressing` frames and `DragArbiter` never gets `press_start`, leaving it
|
||||
stuck in `Idle` (answers `Undecided` forever) for the rest of that
|
||||
gesture. Fixed in `Selection::drag` (`iris/transcript-ui/src/
|
||||
selection.rs`) via a new `DragArbiter::is_idle()` the caller checks to
|
||||
recover a missed press on the next `Pressing` frame. Four new unit
|
||||
tests (three in `iris/src/sense.rs`'s `drag_arbiter_tests`, one in
|
||||
`transcript-ui`'s `selection::tests`, the latter failing on the
|
||||
pre-fix code). See this box's own "Touch-scroll dropout root-caused,
|
||||
2026-09-05" subsection for the trace and what could and could not be
|
||||
re-verified this pass -- **this checkout's emulator turned out to be
|
||||
concurrently in use by another session's P0 benchmark work partway
|
||||
through verification** (its sandbox server was restarted, wiping this
|
||||
pass's test session, and its Compose `bench` app took window focus),
|
||||
so the "run iris-scroll.sh three times cleanly" and "re-take the
|
||||
host-GPU FrameReport row" pass conditions could not be completed
|
||||
end-to-end this pass. The fix itself is verified by direct, targeted
|
||||
logcat traces taken before that interference began, not by the
|
||||
aggregate script.
|
||||
- **Decided 2026-09-05: iris over Masonry**, by Iris, from the host-GPU
|
||||
numbers in I5's box and E1/E2's findings. See the Recommendation's item
|
||||
3 and `DECISIONS.md`. Next: the remaining screens and the app on iris —
|
||||
@@ -3201,6 +3224,116 @@ silently on real hardware.
|
||||
`e2a1fad`'s own message. `docs/DECISIONS.md`'s DEFERRED item is
|
||||
updated with this section's host-GPU table below.
|
||||
|
||||
**Touch-scroll dropout root-caused, 2026-09-05.** Diagnosed as
|
||||
instructed: temporary `log::info!` tracing on every touch event
|
||||
reaching `IrisViewPeer::on_touch_event` (`iris/src/android/view.rs`),
|
||||
every `DragArbiter` state transition (`press_start`/`update`/
|
||||
`release`, `iris/src/sense.rs`), and every `Selection::drag`
|
||||
dispatch (`iris/transcript-ui/src/selection.rs`) -- all removed once
|
||||
the cause was confirmed, per AGENTS.md's "keep the build clean."
|
||||
Reproduced with `app/iris-scroll.sh` against a real sandbox session
|
||||
(30 sent markdown messages, `EMU_GPU` unset / `-gpu host`,
|
||||
`--features transcript-screen,force-gles`, release build, same
|
||||
recipe as this box's own "-gpu host" pass above).
|
||||
|
||||
*The trace.* Of 24 swipes in one run, 5 produced zero `render()`
|
||||
calls each -- one at the very start of the run, four consecutive
|
||||
later (swipes 22-25) -- exactly the "several consecutive swipes
|
||||
produce nothing, an identical retry then works" shape from the
|
||||
earlier pass's report. Correlating the three log streams by
|
||||
timestamp: every one of those 5 swipes delivered a normal
|
||||
`Down`/`Move`×N/`Up` sequence to `on_touch_event` (touch delivery
|
||||
was never the problem), but `Selection::drag` never once saw
|
||||
`PressStart` for the whole gesture -- only `Pressing`, starting from
|
||||
the very first `Move`. `DragArbiter::update`'s `Idle` arm answers
|
||||
every such frame with `Undecided` and never transitions state (there
|
||||
is no way for pure state to tell "no press is happening" from "a
|
||||
press is happening but I missed its start"), so the arbiter sat in
|
||||
`Idle` from the gesture's first frame to its last, `release()` on
|
||||
`Up` its only state change (`Idle` -> `Idle`, a no-op). The row this
|
||||
landed on registered its `PressStart` correctly on a *different*
|
||||
point in the very next successful swipe at the identical screen
|
||||
coordinate -- confirming the miss is about *where the content
|
||||
happens to be under that pixel when `ACTION_DOWN` fires*, not about
|
||||
timing or a coalesced event.
|
||||
|
||||
*Why `ACTION_DOWN` misses a row's sensor.* Each row's `CursorSense`
|
||||
handler is registered only on its `TextEdit` field
|
||||
(`transcript-ui/src/row.rs`'s `build_text_row`), not on the row's
|
||||
`.pad(10)` margin, the `.gap(4)` between the sender-name header and
|
||||
the field, or the header itself (`Span::empty`/a plain `wtext` with
|
||||
no handler). A real touch's down-point is wherever the finger
|
||||
actually is, with no reason to prefer text over padding, and the
|
||||
list has no sensor of its own to fall back to (`iris::widget::list`
|
||||
registers none) -- pan is reachable *only* through a row's own
|
||||
arbiter. So roughly one in five swipes in this run started on a
|
||||
pixel no sensor covered.
|
||||
|
||||
*The fix.* `iris::sense::DragArbiter` gains `pub fn is_idle(&self)`,
|
||||
documented as the recovery signal: a caller that gets a `Pressing`
|
||||
frame while the arbiter reports `is_idle()` knows the button is
|
||||
genuinely down (that is what `Pressing` means) with no matching
|
||||
`press_start` on record, which can only mean it was missed.
|
||||
`Selection::drag`'s match gains one arm, checked after `PressStart`/
|
||||
`PressEnd` and before the ordinary `_ => update(...)` case: `_ if
|
||||
self.arbiter.is_idle()` starts the press right there instead of
|
||||
where it was missed, using whatever `already_selected` holds at
|
||||
that later frame (the best available answer -- the true value at
|
||||
the actual `ACTION_DOWN` is unrecoverable once missed). This is
|
||||
the caller's fix, not the arbiter's, because only the caller knows
|
||||
what `already_selected` should be; the arbiter's own slop/long-press
|
||||
logic was correct throughout and needed no change.
|
||||
|
||||
*Tests.* Four new, all passing on the fix and the first three failing
|
||||
without it: `sense.rs`'s `drag_arbiter_tests::
|
||||
is_idle_reports_a_press_that_was_never_started`,
|
||||
`::update_on_an_idle_arbiter_stays_undecided_forever_without_recovery`
|
||||
(documents the failure mode itself), `::
|
||||
a_caller_can_recover_a_missed_press_start_via_is_idle` (the pure-state
|
||||
half); and `transcript-ui/src/selection.rs`'s `tests::
|
||||
a_missed_press_start_recovers_on_the_next_pressing_frame`, which
|
||||
drives `Selection::drag` directly with only `Pressing` frames (no
|
||||
`PressStart` ever sent) and asserts the arbiter is no longer idle
|
||||
afterward -- this one fails on the pre-fix code (`is_idle()` stays
|
||||
true forever, matching the real trace).
|
||||
|
||||
**Not completed this pass, and why.** The task asked for
|
||||
`iris-scroll.sh` run three times clean and a re-taken host-GPU
|
||||
`FrameReport` row. Partway through that verification, this
|
||||
checkout's shared emulator (`ai-app-2`, per-checkout per AGENTS.md)
|
||||
turned out to be concurrently in use by another session actively
|
||||
working the P0 phone-benchmark item added to this same file earlier
|
||||
today: `adb shell dumpsys activity processes` showed
|
||||
`com.example.aiapp`/`com.example.aiapp.bench` processes running
|
||||
alongside `dev.iris.android.demo`, window focus was observed to have
|
||||
moved to the Compose app mid-test, and the sandbox server's own log
|
||||
showed a fresh `start` (not `keep`) at 00:55 that wiped this pass's
|
||||
30-message test session and replaced it with the peer's own
|
||||
`bench-check` session -- confirmed by `ui-sandbox.sh api /sessions`
|
||||
returning "no session" for the id this pass had been sending to.
|
||||
Rather than disrupt that session's work (deleting its session,
|
||||
restarting its server, or fighting over emulator focus), this pass
|
||||
stopped chasing a clean aggregate number once the cause was
|
||||
confirmed external. What *is* verified is the fix itself, from
|
||||
direct traces taken before the interference began (above) plus two
|
||||
manual, shorter `ui-trace` swipe sequences (not the full script) that
|
||||
each showed full, healthy per-swipe `render()`/`selection::drag`
|
||||
coverage with the fix in place. The FrameReport row in this box's
|
||||
own table above is therefore **not re-taken this pass** -- a future
|
||||
pass should re-run `iris-scroll.sh` three times and retake it once
|
||||
the emulator is free, per AGENTS.md's "ask before/tell peers" and
|
||||
"coordinate with peer agents" guidance rather than contending for it.
|
||||
Also correctly ruled out, not left ambiguous: a hypothesis raised
|
||||
mid-pass that `iris::widget::list::List::scroll`'s deliberately
|
||||
unclamped anchor (its own module doc, "no overscroll clamping ...
|
||||
leaves a gap rather than rubber-banding back") could itself explain
|
||||
a run of consecutive failed swipes once enough net drift
|
||||
accumulates -- plausible in isolation, but the run where it seemed
|
||||
to reproduce is exactly the run now attributed to the peer
|
||||
session's interference (same timestamps), so it was not
|
||||
independently confirmed and is recorded here as ruled out for now
|
||||
rather than as a second bug.
|
||||
|
||||
## The port, in order (decided 2026-09-05)
|
||||
|
||||
Iris decided iris over Masonry (`DECISIONS.md`). This is the ordered plan
|
||||
@@ -3271,6 +3404,89 @@ device.
|
||||
margin of Compose on p50, p99 and CPU time, no crash, no stutter she
|
||||
can see. Fail stops the port.
|
||||
|
||||
**Compose half: done, 2026-09-05.** `app/androidApp`'s `bench` build
|
||||
type, `app/bench-fixture/` (generator + generated `assets/`),
|
||||
`BenchFixture.kt`/`BenchNetwork.kt` (an in-process fake backend: a
|
||||
`URLStreamHandlerFactory` installed only in `FIXTURE_MODE` answers
|
||||
`https://bench.fixture.invalid:1/...` from an in-memory event log
|
||||
instead of opening a socket, so `TranscriptSource`, `EventStream`,
|
||||
the fold and the paging are the *real* ones, unmodified), and
|
||||
`BenchRun.kt` (the scripted scroll-and-stream, driven against the
|
||||
real `LazyListState`) are all in. "Run benchmark" sits beside "Copy"
|
||||
in the session settings dialog, bench-build only
|
||||
(`SessionSettingsDialog`'s `onRunBenchmark`). `./build-apk.sh bench`
|
||||
works, produces a universal APK (no ABI splits in this project, so
|
||||
arm64-v8a is included alongside the others — confirmed with `aapt2
|
||||
dump badging`), signed with the same release key, own application id
|
||||
`com.example.aiapp.bench`, own label "AI Sessions bench" via a
|
||||
build-type `resValue` overriding `@string/app_name`.
|
||||
|
||||
Checks all clean: `ktfmtFormat`, `compileDebugKotlin`,
|
||||
`compileBenchKotlin`, `lintDebug`, `lintBench` (both "No issues
|
||||
found"), `testDebugUnitTest`. `grep -n "tap [0-9]" app/*.sh` has one
|
||||
hit, pre-existing and unrelated — a comment in `bench-lib.sh`
|
||||
recounting the 2026-09-03 incident that made that grep a rule, not a
|
||||
literal `tap` call.
|
||||
|
||||
**Emulator smoke run, 2026-09-05** (this checkout's AVD,
|
||||
`ui-trace` tap-by-label throughout — `tap 'Session settings'` then
|
||||
`tap 'Run benchmark'`, report read back over `adb logcat`):
|
||||
|
||||
ai-app render report
|
||||
device: sdk_gphone64_x86_64 (Google), Android 16
|
||||
build: release
|
||||
|
||||
transcript:
|
||||
28 events, 26 rows, 58 units loaded
|
||||
viewport 1536px, 2 units visible
|
||||
on screen: the list's own 0px, AssistantMsg 18732px
|
||||
0 tool calls and 0 groups open
|
||||
|
||||
frames:
|
||||
1361 frames over 38.1s at 60Hz (16.7ms budget)
|
||||
late: 1353 (99.4%)
|
||||
total p50 27.8ms p90 37.7ms p99 50.1ms
|
||||
gpu p50 18.9ms p90 28.9ms p99 31.5ms
|
||||
|
||||
where the draw phase went:
|
||||
draw phase 3.12ms per frame, of which:
|
||||
the transcript: 0.33ms (measure 0.18, place 0.14, record 0.00)
|
||||
everything else: 2.79ms (89%)
|
||||
|
||||
bench:
|
||||
scroll: 6 cycles (24 swipes), streamed 400/400 fixture events
|
||||
process CPU time over this run: 23005ms
|
||||
peak RSS: 209348kB
|
||||
battery current: mean 900000µA over 39 samples (min 900000, max 900000)
|
||||
|
||||
Read this as "the harness runs end to end and produces every field
|
||||
P0 asked for," not as a phone number: it is software-rendered
|
||||
emulator rasterisation (this-machine-android's skill — the stock
|
||||
Settings app scrolls worse on the same device), and the battery
|
||||
current is a fixed 900mA on every sample, which is the emulator's
|
||||
mocked charger reporting a constant rather than a real battery —
|
||||
expect that field to read "unavailable" or a real varying number
|
||||
only on Iris's own phone. The ordinary debug build was rebuilt and
|
||||
driven with `./transcript-bench.sh` against `ui-sandbox.sh` alongside
|
||||
this and produced its usual report with no `bench:` section, so nothing
|
||||
changed for it.
|
||||
|
||||
`~/host/bench/compose-bench-arm64.apk` (9.7M) and
|
||||
`~/host/bench/README.md` are written, with a heading left for the
|
||||
iris half. **Known interaction**: the bench build keeps the same
|
||||
`aiapp://enroll` intent filter as the ordinary app (it never uses
|
||||
it), so with both installed, driving enrollment through a raw `am
|
||||
start -d aiapp://...` intent (not the in-app QR scanner, which is
|
||||
the primary path and calls straight into the matched activity) opens
|
||||
Android's "Open with" chooser between the two. Cosmetic — the real
|
||||
enrollment path is unaffected — and left as is rather than pulling
|
||||
the intent-filter out of the bench manifest via source-set merging,
|
||||
which was more diff than the problem was worth.
|
||||
|
||||
**Not done this pass**: the iris half (a separate agent's scope —
|
||||
this session was told not to touch `iris/`), and anything past the
|
||||
emulator — the actual on-phone runs and Iris's pass/fail call.
|
||||
|
||||
- [ ] **P1 — session screen parity.** History paging backward (with the
|
||||
page-boundary healing `client-core` does not have yet, below),
|
||||
`TranscriptSource`-backed cache/server stitching, jump-to-latest,
|
||||
|
||||
Reference in new issue
Block a user