Profiling the Catalog

Profiling BookNest's Catalog Scroll

./gradlew :benchmark:connectedBenchmarkReleaseAndroidTest builds the benchmarkRelease variant (release, minified, but profileable) and runs each test with CompilationMode.None(), which throws away all compiled code first, and with CompilationMode.Partial(BaselineProfileMode.Require), which installs the profile and fails if there is none. Ten cold starts each, on this book's emulator:

BookNest's timeToInitialDisplayMs from StartupTimingMetric (10 iterations each)
Mode Min Median Max
No compilation 1,566 ms 1,858.5 ms 2,034 ms
Baseline profile 1,640 ms 1,770.5 ms 1,900 ms

The profile cut the median by 88 ms, about 5%, well short of the 30% Google reports. The scroll benchmark fared worse: it failed with IllegalArgumentException: At least one result is necessary, 0 found for frameDurationCpuMs, so the trace held no frames from BookNest on this emulator's software renderer. dumpsys gfxinfo did count them, three runs per ART filter after three swipes each:

Catalog scroll frames by compiler filter, benchmarkRelease build (three runs each)
Filter Frames Janky P50 P90
verify 18-47 72-77% 53-69 ms 105-121 ms
speed-profile 50-54 74-80% 57-69 ms 101-113 ms

The real result wins over the brochure: on this machine the frames are bound by software rendering (the GPU percentiles read 4,950 ms), not by interpreted code, so precompiling barely moves them; speed-profile rendered more frames in the same swipes, a hint of less CPU per frame. Cold start moved a little because startup is where interpreted code dominates. On a real phone with a hardware GPU, the same benchmarks are the ones to trust, and the Perfetto 440,880 traces they save beside the results (BookNestBenchmarks_startupNoProfile_iter000_....perfetto-trace) open in ui.perfetto.dev to show what each millisecond was. The recomposition counts of Recomposition Counts already showed the Compose 234 side behaving: rows skipped while typing, and only visible rows composed.