Skip to content

fix: select one discovered HTP session for Hexagon inference - #907

Merged
a-ghorbani merged 4 commits into
mainfrom
feature/TASK-20260909-1820
Sep 10, 2026
Merged

fix: select one discovered HTP session for Hexagon inference#907
a-ghorbani merged 4 commits into
mainfrom
feature/TASK-20260909-1820

Conversation

@pocketpal-dev-team

@pocketpal-dev-team pocketpal-dev-team Bot commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

Hexagon selection currently emits ['HTP*'], which requests every discovered HTP session. On devices exposing several sessions, this substantially increases native compute-buffer allocation. Select the first exact runtime HTP name instead, and resolve existing saved wildcard, stale, or multiple HTP selections before normal model initialization.

Resolution preserves saved preferences and the GPU-layer setting. If discovery fails or no HTP device is available, that load uses CPU with zero GPU layers; a later load can recover Hexagon. Settings and benchmarks share the same device choice, while unavailable benchmark backends still fail their cell rather than measure CPU. CPU/OpenCL/iOS selection and flash policy are unchanged.

Fixes #904. Based on main after #901 merged; the diff contains only the HTP selection fix and its tests.

Validation

  • Full lint and typecheck pass. Full Jest coverage run: 275 suites, 4,418 tests, 310 snapshots pass; two tests skipped. Coverage is 77.08% statements, 72.06% branches, 73.13% functions and 77.14% lines.
  • Native-boundary tests cover legacy preferences, unavailable/rejected discovery, CPU fallback and recovery, exact runtime ordering, zero GPU layers, coherent settings during asynchronous discovery, and benchmark isolation.
  • Complete Android E2E Gradle build passes. No native or dependency files change relative to Upgrade llama.rn to 0.13.0-rc.3 (llama.cpp b10588 -> b10829) #901.
  • Myron and Galaxy S23: baseline/candidate benchmark matrices, CPU/OpenCL controls, and flash-OFF Qwen chat before visiting Settings and after selection/restart. The candidate saves one exact HTP name, retains flash OFF and produces the correct answer on both phones.

Paired Android measurements

Benchmarks use identical per-model configs: pp256/tg64, three repetitions, context2048, six threads, mmap off; Hexagon/CPU flash ON and OpenCL flash OFF. These are individual matrix trials, not statistical performance guarantees. All before/after thermal snapshots report status 0, but temperature and OS memory/cache state differ; Myron candidate runs occurred the following day.

Myron case Baseline Candidate
Qwen3 1.7B Q4_K_M HTP 201.39 pp / 18.64 tg tok/s 213.33 pp / 19.86 tg tok/s
Qwen native buffers 5,118.39 MiB across six HTP sessions 1,764.11 MiB with HTP0
Phi4 mini Q4_K_M Android LOW_MEMORY Passed: 91.30 pp / 13.67 tg
Phi4 mini Q6_K Android LOW_MEMORY Passed: 66.80 pp / 11.01 tg
Gemma4 E2B Q4_K_M Android LOW_MEMORY Passed: 134.22 pp / 13.45 tg
Gemma4 E2B Q6_K Passed: 64.85 pp / 2.49 tg Passed: 114.98 pp / 13.59 tg
Gemma4 E2B Q8_0 Passed: 382.06 pp / 9.10 tg Passed: 801.00 pp / 20.95 tg

All eight Myron candidate cells passed, including CPU/OpenCL controls whose native allocations remain unchanged. Native buffer totals are allocation-log measurements, not process PSS or total system memory. Qwen reported peak memory was 1,681.10 → 1,548.58 MiB.

S23 already uses one HTP session in the baseline: Qwen native buffers remain 1,764.11 MiB; prefill/decode measured 95.30/12.72 → 90.09/14.38 tok/s. Known Phi Q4_0 native abort, Phi Q8_0 allocation failure and Gemma Q6_K/Q8_0 memory kills remain. Phi Q6_K initially passed only on baseline; a reversed-order repeat passed only on candidate, leaving each build with one pass and one confirmed memory kill. This does not establish a repeatable candidate-specific regression or workload reliability.

With flash OFF, Myron Phi Q4_K_M and Gemma Q4_K_M both completed correct Paris answers on candidate; baseline attempts ended in confirmed LOW_MEMORY without a complete answer. S23 Phi Q4_0 still aborts after loading and Gemma Q6_K still dies during load on both builds. Qwen completed correct answers before Settings and after selection/restart on both phones.

Artifact provenance and limits

Paired measurements use preserved PR #901 CI build 6408291 (runtime-equivalent to base 986f402a) and a controlled comparison APK with only its Hermes JavaScript bundle replaced by the freshly compiled candidate bundle. Every other non-signing APK entry, including native libraries/assets, is byte-identical. Installed APK hashes and model hashes were verified. The independent complete Gradle build uses package native prebuilts and passes; it is a separate artifact from the controlled measurement APK. Candidate runtime source 175a7016 is unchanged by test-only commit 47a26875. After rebasing onto merged #901, current head d50f9646 has the identical complete Git tree as the reviewed and tested 47a26875; all four patches are unchanged.

Myron E2E data had to be recreated after an installation problem, so its legacy selection was reseeded; S23 retained its E2E datastore across replacement installs. UI text/captures establish actual generation; full native arguments and allocation evidence come from benchmark traces and complementary native-boundary tests. These results do not claim that all Hexagon model failures are fixed.

Visual captures are complete; uploading them is pending GitHub browser-session authentication.

Generated by PocketPal Dev Team

@pocketpal-dev-team pocketpal-dev-team Bot left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent review of 47a26875 is complete across architecture, QA, security, performance, mobile, data, UX and local invariants. No source-code findings were identified. Independent lint, typecheck, full tests and coverage pass; the paired device and actual-generation evidence was reviewed with its documented artifact and experimental limits.

Approval remains pending the required visual evidence: all eight captures exist and were inspected, but uploading them failed because the GitHub browser session expired (HTTP 302). Refreshing that authentication and posting the captures resolves this prerequisite; no code change or new device run is requested for it.

This draft targets the pinned base for #901. GitHub currently lists no status checks for this temporary base; retarget after #901 merges and run the applicable checks on the resulting branch.

Generated by PocketPal Dev Team

@pocketpal-dev-team
pocketpal-dev-team Bot changed the base branch from pr-901-base to main September 10, 2026 09:31
@a-ghorbani
a-ghorbani marked this pull request as ready for review September 10, 2026 09:32
@a-ghorbani
a-ghorbani force-pushed the feature/TASK-20260909-1820 branch from 47a2687 to d50f964 Compare September 10, 2026 09:37
@a-ghorbani

Copy link
Copy Markdown
Owner

PR #907 benchmark results

Run Configuration
PR / runtime #907 / llama.rn 0.13.0-rc.3
APK E2E build 34461875626 · d50f96466c9440258031cfa46cb6a822f6e762aa (PR head); installed SHA-256 verified on all three phones
Δ reference PR #901 bench (APK 6408291f); all 68 native libraries byte-identical to this APK
Workload 256 prefill tokens / 64 decode tokens / 3 repetitions / 1 sequence / 30 s between cells
Context / batch / microbatch / threads 2048 / 512 / 512 / 6
Flash attention CPU + Hexagon: on; GPU/OpenCL: off
Memory settings mmap=false; no_extra_bufts=false; K/V cache=f16
Hexagon device request poco-myron: #901 ["HTP*"] → 6 HTP sessions; #907 ["HTP0"] → 1 HTP session
samsung-s23: #901 ["HTP*"] → 1 HTP session; #907 ["HTP0"] → 1 HTP session
Coverage 203 planned; 185 passed; 18 not passed (#901: 182 passed); Klee supports CPU only
Device Backend Tier Planned Passed Of passed: rerun Not passed
poco-myron cpu smoke 9 9 0 0
poco-myron cpu focused 20 20 0 0
poco-myron hexagon smoke 9 9 0 0
poco-myron hexagon focused 20 20 0 0
poco-myron gpu smoke 9 9 0 0
poco-myron gpu focused 20 20 0 0
samsung-s23 cpu smoke 9 9 0 0
samsung-s23 cpu focused 20 16 3 4
samsung-s23 hexagon smoke 9 9 0 0
samsung-s23 hexagon focused 20 14 2 6
samsung-s23 gpu smoke 9 9 0 0
samsung-s23 gpu focused 20 14 0 6
poco-x7-klee cpu smoke 9 9 0 0
poco-x7-klee cpu focused 20 18 3 2
Reading the tables Meaning
Prefill / Decode tok/s; higher is faster
Weights / KV / Compute / Total Allocated buffers, MiB; not process RSS
Parentheses Δ versus the stated reference, matched settings; + means increase
No measurement or no matched reference; never zero
Pass / Pass (rerun) Completed on requested backend / selected measurement from a rerun
Not passed Failures, timeouts and interruptions remain in coverage
Note Detail
Fix (Myron Hexagon) Every Hexagon cell requests HTP0 and allocates 1 HTP session (#901: HTP* → 6 sessions). Allocated total median −19.2% smoke / −24.0% focused (range −8.7% to −75.1%). The 5 #901 non-passes (Phi-4-mini Q4_K_M, Q6_K; Gemma-4-E2B Q4_K_M, Q6_K, Q8_0) pass.
Myron coverage 87/87 in the uninterrupted matrix, 0 reruns. #901: 82/87, 33 of them from reruns.
Myron Hexagon throughput Matched cells vs #901: prefill median +6.0% smoke (n=9) / +5.4% focused (n=15); decode +11.3% / +12.0%.
Myron CPU Δ CPU path and native libraries unchanged. Deltas run both ways (−23.8% to +41.2% prefill) because #901 CPU references mix an interrupted original matrix with fresh-process reruns. #907 smoke vs focused on shared cells: median −1.2% prefill, −0.7% decode.
S23 Already 1 HTP session on #901; allocated memory identical in every matched cell. Known Phi-4-mini Q4_0 Hexagon native crash and Q8_0 Hexagon load failure (effective CPU) remain.
S23 status changes Phi-4-mini Q6_K CPU, Gemma-4-E2B Q4_K_M Hexagon and Gemma-4-E2B Q4_0 GPU passed on #901 rerun, LMK on #907 rerun. Same-device A/B afterwards (fresh process each, order 901/907/907/901): Phi Q6_K CPU 1/2 vs 1/2, Gemma Q4_K_M Hexagon 2/2 vs 2/2, Gemma Q4_0 GPU 0/2 vs 0/2. No APK difference.
Klee CPU only. Gemma-4-E2B CPU reruns ran directly after the focused matrix (CPU 41 °C). Cool repeat (31.8 °C) vs #901 prefill/decode: Q4_0 92.2/13.7 (−5.1%/−1.2%), Q4_K_M 69.2/13.5 (+6.8%/+0.5%), Q6_K 51.0/10.3 (+11.8%/+5.0%). Table keeps the first rerun.
Klee non-passes Phi-4-mini Q8_0: app died under memory thrashing in the matrix; rerun timed out while thrashing. Gemma-4-E2B Q8_0: LMK (also LMK on #901).
poco-myron · cpu · 29 cells
Tier Model Quant Status Prefill tok/s (Δ) Decode tok/s (Δ) Weights MiB KV MiB Compute MiB Total MiB (Δ)
smoke qwen3.5-0.8b Q4_0 Pass 545.78 (+0.8%) 79.07 (+2.5%) 678.28 24.00 491.02 1,193.30 (+0.0%)
smoke qwen3.5-0.8b Q4_K_M Pass 380.63 (-0.2%) 68.40 (-1.7%) 719.65 24.00 491.02 1,234.67 (+0.0%)
smoke qwen3.5-0.8b Q8_0 Pass 571.24 (+0.3%) 61.42 (+2.3%) 1,023.09 24.00 491.02 1,538.11 (+0.0%)
smoke gemma-3-1b Q4_0 Pass 533.63 (-4.9%) 70.94 (+1.2%) 988.24 22.00 518.75 1,528.99 (+0.0%)
smoke gemma-3-1b Q4_K_M Pass 146.97 (+2.1%) 58.53 (+2.2%) 1,068.49 22.00 518.75 1,609.24 (+0.0%)
smoke gemma-3-1b Q8_0 Pass 588.77 (-3.3%) 53.32 (-1.4%) 1,319.54 22.00 518.75 1,860.29 (+0.0%)
smoke qwen3-1.7b Q4_0 Pass 306.86 (+0.2%) 44.31 (+0.2%) 1,169.07 224.00 304.75 1,697.82 (+0.0%)
smoke qwen3-1.7b Q4_K_M Pass 203.27 (-0.4%) 40.95 (+0.5%) 1,217.35 224.00 304.75 1,746.10 (+0.0%)
smoke qwen3-1.7b Q8_0 Pass 318.90 (-0.7%) 30.06 (-2.4%) 2,059.07 224.00 304.75 2,587.82 (+0.0%)
focused qwen3.5-0.8b Q4_0 Pass 544.30 (+3.0%) 78.96 (+0.6%) 678.28 24.00 491.02 1,193.30 (+0.0%)
focused qwen3.5-0.8b Q4_K_M Pass 379.66 (+8.2%) 69.84 (+1.7%) 719.65 24.00 491.02 1,234.67 (+0.0%)
focused qwen3.5-0.8b Q6_K Pass 357.59 (+23.8%) 61.48 (+8.8%) 824.50 24.00 491.02 1,339.52 (+0.0%)
focused qwen3.5-0.8b Q8_0 Pass 572.08 (+41.2%) 61.02 (+22.3%) 1,023.09 24.00 491.02 1,538.11 (+0.0%)
focused lfm2.5-1.2b-instruct Q4_0 Pass 488.22 (+33.8%) 73.63 (+4.0%) 766.25 24.00 142.01 932.26 (+0.0%)
focused lfm2.5-1.2b-instruct Q4_K_M Pass 310.25 (+11.7%) 69.55 (+1.8%) 799.77 24.00 142.01 965.78 (+0.0%)
focused lfm2.5-1.2b-instruct Q6_K Pass 215.87 (-1.0%) 50.65 (+0.4%) 1,020.97 24.00 142.01 1,186.98 (+0.0%)
focused lfm2.5-1.2b-instruct Q8_0 Pass 496.42 (-0.5%) 49.65 (+1.4%) 1,322.25 24.00 142.01 1,488.26 (+0.0%)
focused qwen3-1.7b Q4_0 Pass 300.45 (-1.8%) 42.91 (-2.8%) 1,169.07 224.00 304.75 1,697.82 (+0.0%)
focused qwen3-1.7b Q4_K_M Pass 196.38 (-2.2%) 40.66 (-0.1%) 1,217.35 224.00 304.75 1,746.10 (+0.0%)
focused qwen3-1.7b Q6_K Pass 138.39 (-4.9%) 30.58 (-1.1%) 1,589.83 224.00 304.75 2,118.58 (+0.0%)
focused qwen3-1.7b Q8_0 Pass 267.85 (-14.5%) 29.80 (-2.8%) 2,059.07 224.00 304.75 2,587.82 (+0.0%)
focused phi-4-mini Q4_0 Pass 124.71 (-5.7%) 21.12 (-1.0%) 2,696.38 256.00 416.76 3,369.14 (+0.0%)
focused phi-4-mini Q4_K_M Pass 75.58 (-13.2%) 18.35 (-2.4%) 2,849.38 256.00 416.76 3,522.14 (+0.0%)
focused phi-4-mini Q6_K Pass 55.64 (-13.0%) 13.97 (+0.7%) 3,482.38 256.00 416.76 4,155.14 (+0.0%)
focused phi-4-mini Q8_0 Pass 108.59 (-17.2%) 14.05 (-0.2%) 4,510.28 256.00 416.76 5,183.04 (+0.0%)
focused gemma-4-e2b Q4_0 Pass 130.93 (-23.8%) 26.19 (-11.0%) 3,522.14 12.00 579.52 4,113.66 (+0.0%)
focused gemma-4-e2b Q4_K_M Pass 93.93 (-20.8%) 24.25 (-10.9%) 3,602.19 12.00 579.52 4,193.71 (+0.0%)
focused gemma-4-e2b Q6_K Pass 83.67 (+14.0%) 19.90 (+4.0%) 4,017.80 12.00 579.52 4,609.32 (+0.0%)
focused gemma-4-e2b Q8_0 Pass 126.57 (+21.4%) 19.63 (+17.6%) 5,130.29 12.00 579.52 5,721.81 (+0.0%)
poco-myron · hexagon · 29 cells
Tier Model Quant Status Prefill tok/s (Δ) Decode tok/s (Δ) Weights MiB KV MiB Compute MiB Total MiB (Δ)
smoke qwen3.5-0.8b Q4_0 Pass 1,250.44 (+6.0%) 39.53 (+27.5%) 678.30 24.00 531.16 1,233.46 (-20.0%)
smoke qwen3.5-0.8b Q4_K_M Pass 364.37 (+5.2%) 26.17 (+11.3%) 719.66 24.00 529.14 1,272.80 (-67.6%)
smoke qwen3.5-0.8b Q8_0 Pass 1,259.88 (+5.7%) 46.90 (+24.5%) 1,023.10 24.00 495.02 1,542.12 (-16.4%)
smoke gemma-3-1b Q4_0 Pass 3,886.46 (+14.2%) 68.09 (+28.7%) 988.24 22.00 521.77 1,532.01 (-19.2%)
smoke gemma-3-1b Q4_K_M Pass 149.57 (+1.4%) 23.54 (+9.8%) 1,068.49 22.00 562.27 1,652.76 (-49.0%)
smoke gemma-3-1b Q8_0 Pass 3,483.88 (+4.0%) 50.11 (+18.3%) 1,319.54 22.00 521.77 1,863.31 (-16.4%)
smoke qwen3-1.7b Q4_0 Pass 2,476.72 (+7.9%) 30.63 (+8.3%) 1,169.07 224.00 354.76 1,747.83 (-17.9%)
smoke qwen3-1.7b Q4_K_M Pass 211.25 (+6.7%) 20.09 (+7.7%) 1,217.35 224.00 322.76 1,764.11 (-65.5%)
smoke qwen3-1.7b Q8_0 Pass 2,514.14 (+10.2%) 30.08 (+10.6%) 2,059.07 224.00 310.76 2,593.83 (-12.5%)
focused qwen3.5-0.8b Q4_0 Pass 1,141.01 (+7.8%) 37.88 (+22.4%) 678.30 24.00 531.16 1,233.46 (-20.0%)
focused qwen3.5-0.8b Q4_K_M Pass 359.71 (+13.2%) 25.61 (+11.7%) 719.66 24.00 529.14 1,272.80 (-67.6%)
focused qwen3.5-0.8b Q6_K Pass 375.14 (+13.6%) 26.91 (+8.6%) 824.50 24.00 529.14 1,377.64 (-54.8%)
focused qwen3.5-0.8b Q8_0 Pass 1,166.87 (+10.6%) 45.61 (+24.2%) 1,023.10 24.00 495.02 1,542.12 (-16.4%)
focused lfm2.5-1.2b-instruct Q4_0 Pass 3,437.39 (+13.1%) 49.10 (+40.5%) 766.25 24.00 194.03 984.28 (-30.8%)
focused lfm2.5-1.2b-instruct Q4_K_M Pass 307.23 (+5.0%) 34.49 (+14.5%) 799.77 24.00 183.96 1,007.73 (-75.1%)
focused lfm2.5-1.2b-instruct Q6_K Pass 207.11 (-1.4%) 29.21 (+10.7%) 1,020.97 24.00 183.96 1,228.93 (-71.2%)
focused lfm2.5-1.2b-instruct Q8_0 Pass 3,836.68 (+17.0%) 45.69 (+37.5%) 1,322.25 24.00 148.02 1,494.27 (-22.0%)
focused qwen3-1.7b Q4_0 Pass 2,442.63 (+6.0%) 31.06 (+12.1%) 1,169.07 224.00 354.76 1,747.83 (-17.9%)
focused qwen3-1.7b Q4_K_M Pass 211.41 (+3.7%) 19.66 (+6.4%) 1,217.35 224.00 322.76 1,764.11 (-65.5%)
focused qwen3-1.7b Q6_K Pass 139.09 (-5.7%) 17.04 (+5.8%) 1,589.83 224.00 322.76 2,136.59 (-61.1%)
focused qwen3-1.7b Q8_0 Pass 2,336.18 (+5.4%) 29.51 (+12.0%) 2,059.07 224.00 310.76 2,593.83 (-12.5%)
focused phi-4-mini Q4_0 Pass 1,302.98 (-1.4%) 18.29 (+7.7%) 2,696.38 256.00 470.76 3,423.14 (-13.1%)
focused phi-4-mini Q4_K_M Pass 71.65 (—) 13.25 (—) 2,849.38 256.00 450.71 3,556.09 (—)
focused phi-4-mini Q6_K Pass 51.10 (—) 9.81 (—) 3,482.38 256.00 450.71 4,189.09 (—)
focused phi-4-mini Q8_0 Pass 1,187.14 (-2.6%) 13.58 (+0.9%) 4,510.28 256.00 424.77 5,191.05 (-8.7%)
focused gemma-4-e2b Q4_0 Pass 493.29 (-12.9%) 23.63 (+67.1%) 3,522.14 12.00 636.01 4,170.15 (-24.0%)
focused gemma-4-e2b Q4_K_M Pass 106.95 (—) 11.39 (—) 3,602.19 12.00 582.51 4,196.70 (—)
focused gemma-4-e2b Q6_K Pass 89.70 (—) 13.30 (—) 4,017.81 12.00 577.51 4,607.32 (—)
focused gemma-4-e2b Q8_0 Pass 626.88 (—) 20.46 (—) 5,130.29 12.00 561.03 5,703.32 (—)
poco-myron · gpu · 29 cells
Tier Model Quant Status Prefill tok/s (Δ) Decode tok/s (Δ) Weights MiB KV MiB Compute MiB Total MiB (Δ)
smoke qwen3.5-0.8b Q4_0 Pass 1,014.42 (+0.6%) 65.22 (-0.9%) 678.36 24.00 499.02 1,201.38 (+0.0%)
smoke qwen3.5-0.8b Q4_K_M Pass 994.76 (-0.5%) 63.16 (+1.0%) 719.73 24.00 499.02 1,242.75 (+0.0%)
smoke qwen3.5-0.8b Q8_0 Pass 922.71 (+0.7%) 50.88 (+0.9%) 1,023.17 24.00 499.02 1,546.19 (+0.0%)
smoke gemma-3-1b Q4_0 Pass 1,065.14 (+0.5%) 54.27 (+1.1%) 988.33 22.00 526.76 1,537.09 (+0.0%)
smoke gemma-3-1b Q4_K_M Pass 158.46 (+1.1%) 23.62 (+0.8%) 1,068.52 22.00 551.51 1,642.03 (+0.0%)
smoke gemma-3-1b Q8_0 Pass 936.48 (+0.3%) 43.57 (+0.6%) 1,319.63 22.00 526.76 1,868.39 (+0.0%)
smoke qwen3-1.7b Q4_0 Pass 602.42 (-0.1%) 41.37 (+0.8%) 1,169.17 224.00 316.76 1,709.93 (+0.0%)
smoke qwen3-1.7b Q4_K_M Pass 597.55 (+0.5%) 39.05 (+0.2%) 1,217.45 224.00 316.76 1,758.21 (+0.0%)
smoke qwen3-1.7b Q8_0 Pass 508.69 (+0.0%) 28.03 (+0.2%) 2,059.17 224.00 316.76 2,599.93 (+0.0%)
focused qwen3.5-0.8b Q4_0 Pass 947.57 (-6.2%) 62.94 (-4.0%) 678.36 24.00 499.02 1,201.38 (+0.0%)
focused qwen3.5-0.8b Q4_K_M Pass 936.12 (-6.0%) 60.88 (-2.4%) 719.73 24.00 499.02 1,242.75 (+0.0%)
focused qwen3.5-0.8b Q6_K Pass 921.25 (-6.7%) 55.67 (-0.4%) 824.58 24.00 499.02 1,347.60 (+0.0%)
focused qwen3.5-0.8b Q8_0 Pass 868.68 (-5.5%) 49.89 (-1.9%) 1,023.17 24.00 499.02 1,546.19 (+0.0%)
focused lfm2.5-1.2b-instruct Q4_0 Pass 859.64 (-6.7%) 65.17 (-1.9%) 766.29 24.00 164.02 954.31 (+0.0%)
focused lfm2.5-1.2b-instruct Q4_K_M Pass 849.14 (-6.4%) 62.26 (+0.2%) 799.81 24.00 164.02 987.83 (+0.0%)
focused lfm2.5-1.2b-instruct Q6_K Pass 931.84 (+0.7%) 51.69 (+0.9%) 1,021.01 24.00 164.02 1,209.03 (+0.0%)
focused lfm2.5-1.2b-instruct Q8_0 Pass 767.40 (-0.4%) 43.74 (+1.3%) 1,322.29 24.00 164.02 1,510.31 (+0.0%)
focused qwen3-1.7b Q4_0 Pass 604.20 (+0.4%) 41.43 (+1.4%) 1,169.17 224.00 316.76 1,709.93 (+0.0%)
focused qwen3-1.7b Q4_K_M Pass 598.51 (+0.6%) 39.35 (+1.1%) 1,217.45 224.00 316.76 1,758.21 (+0.0%)
focused qwen3-1.7b Q6_K Pass 604.31 (-0.2%) 32.85 (+0.6%) 1,589.93 224.00 316.76 2,130.69 (+0.0%)
focused qwen3-1.7b Q8_0 Pass 507.19 (+0.6%) 28.29 (+1.0%) 2,059.17 224.00 316.76 2,599.93 (+0.0%)
focused phi-4-mini Q4_0 Pass 289.11 (-1.4%) 22.59 (+0.5%) 2,696.44 256.00 422.76 3,375.20 (+0.0%)
focused phi-4-mini Q4_K_M Pass 253.51 (-0.9%) 20.62 (+0.6%) 2,849.44 256.00 422.76 3,528.20 (+0.0%)
focused phi-4-mini Q6_K Pass 285.78 (+0.9%) 17.49 (+0.4%) 3,482.44 256.00 422.76 4,161.20 (+0.0%)
focused phi-4-mini Q8_0 Pass 240.27 (-1.1%) 13.76 (+0.2%) 4,510.34 256.00 422.76 5,189.10 (+0.0%)
focused gemma-4-e2b Q4_0 Pass 397.51 (+0.8%) 27.21 (-1.1%) 3,522.29 18.00 582.03 4,122.32 (+0.0%)
focused gemma-4-e2b Q4_K_M Pass 325.17 (+1.6%) 17.64 (-0.4%) 3,602.27 18.00 597.53 4,217.80 (+0.0%)
focused gemma-4-e2b Q6_K Pass 326.13 (-0.9%) 15.83 (-0.2%) 4,017.92 18.00 596.53 4,632.45 (+0.0%)
focused gemma-4-e2b Q8_0 Pass 280.80 (-17.7%) 17.46 (-14.3%) 5,130.45 18.00 582.03 5,730.48 (+0.0%)
samsung-s23 · cpu · 29 cells
Tier Model Quant Status Prefill tok/s (Δ) Decode tok/s (Δ) Weights MiB KV MiB Compute MiB Total MiB (Δ)
smoke qwen3.5-0.8b Q4_0 Pass 219.11 (+0.3%) 42.12 (+0.7%) 678.28 24.00 491.02 1,193.30 (+0.0%)
smoke qwen3.5-0.8b Q4_K_M Pass 160.15 (+1.0%) 37.78 (+4.5%) 719.65 24.00 491.02 1,234.67 (+0.0%)
smoke qwen3.5-0.8b Q8_0 Pass 219.65 (+0.1%) 37.69 (+0.3%) 1,023.09 24.00 491.02 1,538.11 (+0.0%)
smoke gemma-3-1b Q4_0 Pass 211.04 (+3.3%) 41.48 (+1.2%) 988.24 22.00 518.75 1,528.99 (+0.0%)
smoke gemma-3-1b Q4_K_M Pass 68.95 (+0.7%) 30.79 (+1.8%) 1,068.49 22.00 518.75 1,609.24 (+0.0%)
smoke gemma-3-1b Q8_0 Pass 203.29 (+3.7%) 33.17 (+5.6%) 1,319.54 22.00 518.75 1,860.29 (+0.0%)
smoke qwen3-1.7b Q4_0 Pass 124.34 (+2.8%) 20.69 (+0.9%) 1,169.07 224.00 304.75 1,697.82 (+0.0%)
smoke qwen3-1.7b Q4_K_M Pass 84.05 (+1.5%) 16.31 (-12.7%) 1,217.35 224.00 304.75 1,746.10 (+0.0%)
smoke qwen3-1.7b Q8_0 Pass 114.06 (+8.8%) 15.67 (+6.1%) 2,059.07 224.00 304.75 2,587.82 (+0.0%)
focused qwen3.5-0.8b Q4_0 Pass 208.30 (+8.7%) 39.42 (+5.7%) 678.28 24.00 491.02 1,193.30 (+0.0%)
focused qwen3.5-0.8b Q4_K_M Pass 153.09 (+1.9%) 35.67 (+7.1%) 719.65 24.00 491.02 1,234.67 (+0.0%)
focused qwen3.5-0.8b Q6_K Pass 140.86 (+0.8%) 30.15 (+0.8%) 824.50 24.00 491.02 1,339.52 (+0.0%)
focused qwen3.5-0.8b Q8_0 Pass 203.96 (+6.6%) 36.63 (+10.4%) 1,023.09 24.00 491.02 1,538.11 (+0.0%)
focused lfm2.5-1.2b-instruct Q4_0 Pass 198.91 (+4.4%) 42.30 (+4.3%) 766.25 24.00 142.01 932.26 (+0.0%)
focused lfm2.5-1.2b-instruct Q4_K_M Pass 128.46 (+3.1%) 37.48 (+0.3%) 799.77 24.00 142.01 965.78 (+0.0%)
focused lfm2.5-1.2b-instruct Q6_K Pass 91.14 (+2.6%) 22.04 (+3.3%) 1,020.97 24.00 142.01 1,186.98 (+0.0%)
focused lfm2.5-1.2b-instruct Q8_0 Pass 175.23 (+1.1%) 29.35 (+7.8%) 1,322.25 24.00 142.01 1,488.26 (+0.0%)
focused qwen3-1.7b Q4_0 Pass 118.22 (+13.7%) 20.04 (+9.4%) 1,169.07 224.00 304.75 1,697.82 (+0.0%)
focused qwen3-1.7b Q4_K_M Pass 78.27 (+6.9%) 18.42 (+10.3%) 1,217.35 224.00 304.75 1,746.10 (+0.0%)
focused qwen3-1.7b Q6_K Pass 57.28 (+0.8%) 12.66 (+15.6%) 1,589.83 224.00 304.75 2,118.58 (+0.0%)
focused qwen3-1.7b Q8_0 Pass 103.95 (+16.2%) 14.31 (+1.8%) 2,059.07 224.00 304.75 2,587.82 (+0.0%)
focused phi-4-mini Q4_0 Pass 47.57 (-6.1%) 8.06 (-11.0%) 2,696.38 256.00 416.76 3,369.14 (+0.0%)
focused phi-4-mini Q4_K_M Pass (rerun) 35.13 (-7.4%) 7.94 (-10.7%) 2,849.38 256.00 416.76 3,522.14 (+0.0%)
focused phi-4-mini Q6_K LMK
focused phi-4-mini Q8_0 LMK
focused gemma-4-e2b Q4_0 Pass (rerun) 64.54 (-0.0%) 14.67 (+2.0%) 3,522.14 12.00 579.52 4,113.66 (+0.0%)
focused gemma-4-e2b Q4_K_M Pass (rerun) 33.72 (—) 9.73 (—) 3,602.19 12.00 579.52 4,193.71 (—)
focused gemma-4-e2b Q6_K LMK
focused gemma-4-e2b Q8_0 LMK
samsung-s23 · hexagon · 29 cells
Tier Model Quant Status Prefill tok/s (Δ) Decode tok/s (Δ) Weights MiB KV MiB Compute MiB Total MiB (Δ)
smoke qwen3.5-0.8b Q4_0 Pass 482.78 (-0.1%) 22.14 (-0.4%) 678.30 24.00 531.16 1,233.46 (+0.0%)
smoke qwen3.5-0.8b Q4_K_M Pass 158.31 (-0.6%) 15.67 (-1.0%) 719.66 24.00 529.14 1,272.80 (+0.0%)
smoke qwen3.5-0.8b Q8_0 Pass 491.37 (-1.9%) 27.95 (-2.1%) 1,023.10 24.00 495.02 1,542.12 (+0.0%)
smoke gemma-3-1b Q4_0 Pass 1,362.75 (+0.9%) 46.91 (+1.1%) 988.24 22.00 521.77 1,532.01 (+0.0%)
smoke gemma-3-1b Q4_K_M Pass 69.52 (+2.4%) 16.24 (-17.4%) 1,068.49 22.00 562.27 1,652.76 (+0.0%)
smoke gemma-3-1b Q8_0 Pass 1,351.19 (+0.7%) 31.59 (+0.1%) 1,319.54 22.00 521.77 1,863.31 (+0.0%)
smoke qwen3-1.7b Q4_0 Pass 1,038.01 (+1.0%) 25.81 (+1.0%) 1,169.07 224.00 354.76 1,747.83 (+0.0%)
smoke qwen3-1.7b Q4_K_M Pass 87.55 (+3.2%) 14.10 (-0.0%) 1,217.35 224.00 322.76 1,764.11 (+0.0%)
smoke qwen3-1.7b Q8_0 Pass 1,032.45 (-2.1%) 19.29 (-9.3%) 2,059.07 224.00 310.76 2,593.83 (+0.0%)
focused qwen3.5-0.8b Q4_0 Pass 477.88 (-0.7%) 21.15 (-4.8%) 678.30 24.00 531.16 1,233.46 (+0.0%)
focused qwen3.5-0.8b Q4_K_M Pass 150.91 (-0.5%) 16.27 (+6.1%) 719.66 24.00 529.14 1,272.80 (+0.0%)
focused qwen3.5-0.8b Q6_K Pass 168.78 (+1.2%) 19.57 (+14.8%) 824.50 24.00 529.14 1,377.64 (+0.0%)
focused qwen3.5-0.8b Q8_0 Pass 496.35 (+0.6%) 29.79 (+5.9%) 1,023.10 24.00 495.02 1,542.12 (+0.0%)
focused lfm2.5-1.2b-instruct Q4_0 Pass 1,744.13 (+12.4%) 39.47 (+48.8%) 766.25 24.00 194.03 984.28 (+0.0%)
focused lfm2.5-1.2b-instruct Q4_K_M Pass 128.56 (+3.3%) 26.02 (+15.2%) 799.77 24.00 183.96 1,007.73 (+0.0%)
focused lfm2.5-1.2b-instruct Q6_K Pass 91.17 (+5.5%) 18.67 (+5.8%) 1,020.97 24.00 183.96 1,228.93 (+0.0%)
focused lfm2.5-1.2b-instruct Q8_0 Pass 1,813.94 (-2.6%) 28.43 (-15.2%) 1,322.25 24.00 148.02 1,494.27 (+0.0%)
focused qwen3-1.7b Q4_0 Pass 995.45 (-3.0%) 24.60 (-0.5%) 1,169.07 224.00 354.76 1,747.83 (+0.0%)
focused qwen3-1.7b Q4_K_M Pass 82.66 (+5.6%) 14.07 (+1.9%) 1,217.35 224.00 322.76 1,764.11 (+0.0%)
focused qwen3-1.7b Q6_K Pass 62.58 (+8.8%) 10.66 (+9.5%) 1,589.83 224.00 322.76 2,136.59 (+0.0%)
focused qwen3-1.7b Q8_0 Pass 1,023.43 (+0.9%) 19.03 (-1.1%) 2,059.07 224.00 310.76 2,593.83 (+0.0%)
focused phi-4-mini Q4_0 Crash
focused phi-4-mini Q4_K_M Pass (rerun) 32.26 (-15.2%) 7.44 (-7.6%) 2,849.38 256.00 450.71 3,556.09 (+0.0%)
focused phi-4-mini Q6_K LMK
focused phi-4-mini Q8_0 Load error (effective cpu)
focused gemma-4-e2b Q4_0 Pass (rerun) 264.26 (+7.7%) 15.52 (+6.7%) 3,522.14 12.00 636.01 4,170.15 (+0.0%)
focused gemma-4-e2b Q4_K_M LMK
focused gemma-4-e2b Q6_K LMK
focused gemma-4-e2b Q8_0 LMK
samsung-s23 · gpu · 29 cells
Tier Model Quant Status Prefill tok/s (Δ) Decode tok/s (Δ) Weights MiB KV MiB Compute MiB Total MiB (Δ)
smoke qwen3.5-0.8b Q4_0 Pass 243.00 (-1.1%) 14.11 (+1.4%) 678.37 24.00 543.01 1,245.38 (+0.0%)
smoke qwen3.5-0.8b Q4_K_M Pass 221.18 (-2.1%) 12.99 (-2.2%) 719.74 24.00 543.01 1,286.75 (+0.0%)
smoke qwen3.5-0.8b Q8_0 Pass 252.17 (-0.3%) 16.11 (+1.1%) 1,023.17 24.00 499.02 1,546.19 (+0.0%)
smoke gemma-3-1b Q4_0 Pass 114.52 (-1.5%) 26.59 (-7.4%) 988.33 22.00 526.76 1,537.09 (+0.0%)
smoke gemma-3-1b Q4_K_M Pass 60.38 (+2.4%) 11.69 (-0.7%) 1,068.52 22.00 551.51 1,642.03 (+0.0%)
smoke gemma-3-1b Q8_0 Pass 130.72 (-4.7%) 24.25 (-9.9%) 1,319.63 22.00 526.76 1,868.39 (+0.0%)
smoke qwen3-1.7b Q4_0 Pass 244.65 (-3.1%) 15.87 (-0.1%) 1,169.17 224.00 392.76 1,785.93 (+0.0%)
smoke qwen3-1.7b Q4_K_M Pass 227.04 (-5.2%) 13.94 (-3.2%) 1,217.45 224.00 392.76 1,834.21 (+0.0%)
smoke qwen3-1.7b Q8_0 Pass 261.02 (-3.5%) 16.37 (+0.8%) 2,059.17 224.00 316.76 2,599.93 (+0.0%)
focused qwen3.5-0.8b Q4_0 Pass 243.36 (+0.1%) 13.96 (+0.3%) 678.37 24.00 543.01 1,245.38 (+0.0%)
focused qwen3.5-0.8b Q4_K_M Pass 221.89 (+0.1%) 13.57 (-0.2%) 719.74 24.00 543.01 1,286.75 (+0.0%)
focused qwen3.5-0.8b Q6_K Pass 204.83 (+2.7%) 12.93 (+6.0%) 824.58 24.00 543.01 1,391.59 (+0.0%)
focused qwen3.5-0.8b Q8_0 Pass 249.87 (+0.1%) 15.96 (+3.5%) 1,023.17 24.00 499.02 1,546.19 (+0.0%)
focused lfm2.5-1.2b-instruct Q4_0 Pass 416.52 (+5.3%) 23.84 (+4.6%) 766.29 24.00 286.01 1,076.30 (+0.0%)
focused lfm2.5-1.2b-instruct Q4_K_M Pass 399.02 (+0.8%) 21.97 (+1.3%) 799.81 24.00 286.01 1,109.82 (+0.0%)
focused lfm2.5-1.2b-instruct Q6_K Pass 204.30 (+3.3%) 19.85 (+2.4%) 1,021.01 24.00 286.01 1,331.02 (+0.0%)
focused lfm2.5-1.2b-instruct Q8_0 Pass 459.53 (+5.4%) 29.65 (+2.0%) 1,322.29 24.00 164.02 1,510.31 (+0.0%)
focused qwen3-1.7b Q4_0 Pass 269.43 (+6.7%) 16.10 (+1.6%) 1,169.17 224.00 392.76 1,785.93 (+0.0%)
focused qwen3-1.7b Q4_K_M Pass 240.61 (-0.2%) 14.86 (+3.9%) 1,217.45 224.00 392.76 1,834.21 (+0.0%)
focused qwen3-1.7b Q6_K Pass 131.51 (+2.7%) 12.92 (+3.0%) 1,589.93 224.00 392.76 2,206.69 (+0.0%)
focused qwen3-1.7b Q8_0 Pass 266.98 (-2.4%) 16.73 (-0.8%) 2,059.17 224.00 316.76 2,599.93 (+0.0%)
focused phi-4-mini Q4_0 Pass 120.02 (+5.1%) 9.73 (+5.8%) 2,696.44 256.00 518.76 3,471.20 (+0.0%)
focused phi-4-mini Q4_K_M Pass 96.49 (+16.7%) 8.33 (+8.3%) 2,849.44 256.00 518.76 3,624.20 (+0.0%)
focused phi-4-mini Q6_K LMK
focused phi-4-mini Q8_0 LMK
focused gemma-4-e2b Q4_0 LMK
focused gemma-4-e2b Q4_K_M LMK
focused gemma-4-e2b Q6_K LMK
focused gemma-4-e2b Q8_0 LMK
poco-x7-klee · cpu · 29 cells
Tier Model Quant Status Prefill tok/s (Δ) Decode tok/s (Δ) Weights MiB KV MiB Compute MiB Total MiB (Δ)
smoke qwen3.5-0.8b Q4_0 Pass 281.18 (+2.1%) 40.95 (+1.1%) 678.28 24.00 491.02 1,193.30 (+0.0%)
smoke qwen3.5-0.8b Q4_K_M Pass 205.84 (+0.6%) 36.69 (-0.7%) 719.65 24.00 491.02 1,234.67 (+0.0%)
smoke qwen3.5-0.8b Q8_0 Pass 269.62 (-2.1%) 36.79 (+0.3%) 1,023.09 24.00 491.02 1,538.11 (+0.0%)
smoke gemma-3-1b Q4_0 Pass 300.86 (+0.0%) 40.89 (-0.4%) 988.24 22.00 518.75 1,528.99 (+0.0%)
smoke gemma-3-1b Q4_K_M Pass 92.94 (-0.3%) 32.31 (-4.1%) 1,068.49 22.00 518.75 1,609.24 (+0.0%)
smoke gemma-3-1b Q8_0 Pass 307.66 (+2.3%) 30.91 (-1.5%) 1,319.54 22.00 518.75 1,860.29 (+0.0%)
smoke qwen3-1.7b Q4_0 Pass 175.86 (-1.9%) 22.78 (-2.6%) 1,169.07 224.00 304.75 1,697.82 (+0.0%)
smoke qwen3-1.7b Q4_K_M Pass 118.29 (+1.5%) 21.39 (+1.6%) 1,217.35 224.00 304.75 1,746.10 (+0.0%)
smoke qwen3-1.7b Q8_0 Pass 156.79 (+3.9%) 17.71 (+1.5%) 2,059.07 224.00 304.75 2,587.82 (+0.0%)
focused qwen3.5-0.8b Q4_0 Pass 279.42 (+2.5%) 40.56 (+2.7%) 678.28 24.00 491.02 1,193.30 (+0.0%)
focused qwen3.5-0.8b Q4_K_M Pass 195.90 (+4.4%) 36.37 (+0.7%) 719.65 24.00 491.02 1,234.67 (+0.0%)
focused qwen3.5-0.8b Q6_K Pass 185.18 (+5.0%) 32.15 (-0.2%) 824.50 24.00 491.02 1,339.52 (+0.0%)
focused qwen3.5-0.8b Q8_0 Pass 269.05 (+6.1%) 36.38 (+1.3%) 1,023.09 24.00 491.02 1,538.11 (+0.0%)
focused lfm2.5-1.2b-instruct Q4_0 Pass 280.60 (+10.8%) 46.52 (+3.0%) 766.25 24.00 142.01 932.26 (+0.0%)
focused lfm2.5-1.2b-instruct Q4_K_M Pass 173.19 (+7.5%) 41.42 (+0.9%) 799.77 24.00 142.01 965.78 (+0.0%)
focused lfm2.5-1.2b-instruct Q6_K Pass 106.00 (+3.5%) 22.96 (+1.2%) 1,020.97 24.00 142.01 1,186.98 (+0.0%)
focused lfm2.5-1.2b-instruct Q8_0 Pass 205.70 (+6.2%) 30.15 (-2.3%) 1,322.25 24.00 142.01 1,488.26 (+0.0%)
focused qwen3-1.7b Q4_0 Pass 158.54 (+7.5%) 22.55 (+0.9%) 1,169.07 224.00 304.75 1,697.82 (+0.0%)
focused qwen3-1.7b Q4_K_M Pass 106.55 (+5.6%) 21.16 (+1.4%) 1,217.35 224.00 304.75 1,746.10 (+0.0%)
focused qwen3-1.7b Q6_K Pass 72.52 (+2.3%) 13.41 (-0.2%) 1,589.83 224.00 304.75 2,118.58 (+0.0%)
focused qwen3-1.7b Q8_0 Pass 135.61 (+1.6%) 17.67 (+0.6%) 2,059.07 224.00 304.75 2,587.82 (+0.0%)
focused phi-4-mini Q4_0 Pass 68.19 (+2.7%) 11.27 (+1.1%) 2,696.38 256.00 416.76 3,369.14 (+0.0%)
focused phi-4-mini Q4_K_M Pass 41.56 (+3.1%) 9.57 (+1.5%) 2,849.38 256.00 416.76 3,522.14 (+0.0%)
focused phi-4-mini Q6_K Pass 29.63 (-1.3%) 6.45 (-4.5%) 3,482.38 256.00 416.76 4,155.14 (+0.0%)
focused phi-4-mini Q8_0 Timeout
focused gemma-4-e2b Q4_0 Pass (rerun) 66.60 (-31.4%) 11.40 (-17.9%) 3,522.14 12.00 579.52 4,113.66 (+0.0%)
focused gemma-4-e2b Q4_K_M Pass (rerun) 54.01 (-16.6%) 11.32 (-15.8%) 3,602.19 12.00 579.52 4,193.71 (+0.0%)
focused gemma-4-e2b Q6_K Pass (rerun) 41.97 (-7.9%) 9.37 (-4.1%) 4,017.80 12.00 579.52 4,609.32 (+0.0%)
focused gemma-4-e2b Q8_0 LMK

PR #907 feature e2e

Run Configuration
APK Same E2E build as above; Pixel 9 and S23 with full reset per spec, Myron without (HyperOS install dialog)
Specs 17 Android feature specs (release set; purchase-flow excluded); test code at d50f9646
Remote endpoint llama.cpp server on the bench host (Qwen3-1.7B text, Gemma-4-E2B vision)
Spec Pixel 9 (no Hexagon) poco-myron samsung-s23
quick-smoke Pass (1/1) Pass (1/1) Pass (1/1)
thinking Pass (3/3) Pass (3/3) Pass (3/3)
thinking-pal-override Pass (3/3) Pass (3/3) Fail (0/3)
language Pass (1/1) Pass (1/1) Pass (1/1)
draft-autosave Pass (4/4) Pass (4/4) Pass (4/4)
context-banner Pass (5/5) Pass (5/5) Pass (5/5)
hub-run Pass (2/2) Pass (2/2) Pass (2/2)
pal-greeting Pass (1/1) Pass (1/1) Pass (1/1)
talent-tool-use Pass (1/1) Pass (1/1) Pass (1/1)
graded-effort-override Pass (1/1) Pass (1/1) Pass (1/1)
download-cancel Pass (1/1) Pass (1/1) Pass (1/1)
speculative Pass (3/3) Pass (3/3) Pass (3/3)
speculative-paired Pass (1/1) Fail (0/1) Fail (0/1)
speculative-visual Pass (6/6) Pass (6/6) Pass (7/7)
remote-server Pass (3/3) Pass (3/3) Pass (3/3)
remote-reasoning Pass (5/5) Fail (4/5) Fail (4/5)
remote-vision Fail (0/2) Fail (0/2) Fail (0/2)

Failures were rerun alone on the #907 APK and on the #901 APK (6408291f), same phone and test code. Counts are tests passed / tests in the spec.

Device Spec #907 main run #907 solo rerun #901 APK Error
Pixel 9 remote-vision 0/2 2/2 2/2 Main run overlapped Myron's remote-vision on the same server
poco-myron remote-vision 0/2 2/2 2/2 Same overlap
poco-myron remote-reasoning 4/5 5/5 5/5 Reasoning bubble not rendered (thinking on)
poco-myron speculative-paired 0/1 0/1 MTP draft model card not shown within 900 s (both APKs)
samsung-s23 thinking-pal-override 0/3 0/3 Toggle assertions (both APKs; also failed on S23 in v1.16.0)
samsung-s23 speculative-paired 0/1 0/1 chat-input not displayed after 15 s (both APKs)
samsung-s23 remote-reasoning 4/5 0/5, 5/5, 0/5 5/5, 0/5 0/5 runs fail at the server-type probe, on both APKs
samsung-s23 remote-vision 0/2 0/2 0/2 Remote model gemma-4-e2b not found in chat picker (both APKs)
Note Detail
Result No failure is specific to #907. Every #907 failure either passes when rerun alone or fails the same way on #901.
Harness Myron's first feature attempt ran against the lock screen after the bench released stay-awake; logs kept, full set rerun with the device held awake.

@a-ghorbani

Copy link
Copy Markdown
Owner

PR #907 compared with release v1.17.2

Run Configuration
This PR #907 bench above: E2E build 34461875626, d50f9646, llama.rn 0.13.0-rc.3
Reference release v1.17.2 bench, 2a1467e4, llama.rn 0.13.0-rc.1
Coverage Myron only; v1.17.2 was not benched on S23 or Klee
Workload / settings Same standard configs: pp256 / tg64 / 3 reps / 30 s settle; CPU + Hexagon flash on, GPU flash off; mmap off. Initialization settings identical in every matched cell
Scope The delta includes the llama.rn rc.1 → rc.3 bump from #901 plus the single-HTP fix from this PR
Myron v1.17.2 #907
Cells passed 87/87 87/87
Hexagon devices HTP* → 6 or 12 HTP sessions HTP0 → 1 HTP session
Backend Tier Matched cells Prefill Δ median [min, max] Decode Δ median [min, max] Allocated total Δ median [min, max]
Hexagon smoke 9 +10.6% [−1.8%, +21.7%] +43.3% [+3.1%, +84.2%] −13.5% [−16.4%, −0.5%]
Hexagon focused 20 +16.5% [+2.9%, +26.5%] +20.6% [+3.5%, +101.8%] −8.8% [−22.4%, −2.7%]
CPU smoke 9 +1.5% [−2.7%, +19.9%] +12.9% [+8.2%, +25.9%] 0.0%
CPU focused 20 +28.8% [+3.7%, +42.5%] +18.2% [+9.9%, +27.5%] 0.0%
GPU smoke 9 0.0% [−83.4%, +1.9%] +3.2% [−51.1%, +14.3%] 0.0% [0.0%, +1.5%]
GPU focused 20 +3.5% [−10.3%, +8.4%] +5.6% [−25.7%, +17.7%] 0.0% [0.0%, +0.4%]
Note Detail
Hexagon Gains come from both changes. Example, Qwen3.5-0.8B Q8_0 smoke decode: 25.5 (v1.17.2) → 37.7 (#901) → 46.9 tok/s (#907).
CPU focused Do not read the +28.8% as a gain. v1.17.2's own focused-vs-smoke prefill on 6 shared CPU cells is −25.3% [−30.2%, −10.9%], so its focused CPU reference is confounded. #907 on the same cells: −1.2%. Use the smoke row for CPU.
GPU Three cells regress with the rc.3 engine and are unchanged between #901 and #907. See below.

Myron GPU Q4_K_M regression

Build llama.rn Gemma-3-1B Q4_K_M Qwen3.5-0.8B Q4_K_M Qwen3-1.7B Q4_K_M Gemma-3-1B Q4_0
PR #713 baseline 132 / 22.1 646 / 35.4 371 / 23.3 1,025 / 50.4
v1.15.2 0.12.4 145 / 22.7 645 / 31.8 408 / 24.6 1,020 / 47.0
v1.16.0 0.12.4 130 / 22.3 616 / 32.6 417 / 24.1 958 / 49.4
v1.17.0 0.13.0-rc.0 908 / 40.8 941 / 56.3 562 / 29.4 991 / 43.6
v1.17.2 0.13.0-rc.1 953 / 48.3 996 / 59.1 594 / 36.1 1,045 / 52.6
#901 0.13.0-rc.3 157 / 23.4 999 / 62.6 595 / 39.0 1,060 / 53.7
#907 0.13.0-rc.3 158 / 23.6 995 / 63.2 598 / 39.0 1,065 / 54.3

Smoke tier, prefill / decode tok/s. v1.16.0 uses its cool smoke rerun (its focused tier is contaminated).

Gemma-3-1B Q4_K_M, GPU Weights on CPU Weights on OpenCL Compute on CPU
v1.16.0 (0.12.4) 605.13 MiB 463.36 MiB 37.26 MiB
v1.17.2 (rc.1) 306.00 MiB 762.58 MiB 12.51 MiB
#907 (rc.3) 605.13 MiB 463.39 MiB 37.26 MiB
Finding Evidence
Cell history 130–157 tok/s prefill through v1.16; about 6× faster on 0.13.0-rc.0/rc.1; back to the old level on rc.3. Q4_K_M on the Qwen models keeps the rc.0 gain; Gemma-3-1B Q4_0 and Q8_0 are unaffected.
Model file Gemma-3-1B Q4_K_M stores 117 tensors as Q5_0, 299.1 MiB in total; its hidden size is 1152, likely not divisible into the 256-element K-quant blocks. The Qwen Q4_K_M files contain no Q5_0. The 306 MiB kept on CPU in every build is the Q8_0 token embedding, also on CPU for Q4_0.
Placement On rc.3, CPU weights grow by 299.13 MiB (605.13 − 306.00), matching the Q5_0 total.
Code change llama.cpp 95ef7fc16 (#26477, "opencl: quant lm_head / decode GEMV and medium-batch GEMM optimizations"), between b10588 (rc.1) and b10829 (rc.3), makes OpenCL supports_op decline MUL_MAT when src1->ne[1] >= 512 for quant types without an Adreno GEMM kernel. Q5_0 has none. Present in the rc.3 package source, absent in rc.1.
Mechanism llama.cpp tests weight placement (weight_buft_supported, llama-model-loader.cpp) with a 512-row MUL_MAT. The OpenCL device declines Q5_0, so those weights load on CPU and every layer crosses CPU/GPU, as on 0.12.4.
Confidence Established from source and matching allocation sizes. Not yet confirmed with a patched build.
Gemma-4-E2B GPU decode vs v1.17.2 (focused): Q4_K_M −25.7%, Q6_K −24.4%, Q4_0 +9.3%. On rc.3, 29.24 / 25.84 MiB of weights move from OpenCL to a new CPU_REPACK buffer in those two files only. Tensors not traced.
This PR No effect. #901 and #907 match for all three cells.

Feature e2e compared with v1.17.2

Device v1.17.2 #907 main run #907 after solo reruns
Pixel 9 13/17 16/17 17/17
poco-myron 12/17 14/17 16/17
samsung-s23 — (not run) 13/17 13/17; all 4 fail the same way on the #901 APK
Spec v1.17.2 Pixel 9 v1.17.2 Myron #907 Pixel 9 #907 Myron
pal-greeting Pass Fail Pass Pass
speculative-paired Fail Fail Pass Fail (same error on #901)
remote-server Fail Fail Pass Pass
remote-reasoning Fail Fail Pass Pass on solo rerun
remote-vision Fail Fail Pass on solo rerun Pass on solo rerun
Other 12 specs Pass Pass Pass Pass

@a-ghorbani a-ghorbani left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@a-ghorbani
a-ghorbani merged commit d7df7b7 into main Sep 10, 2026
5 checks passed
@a-ghorbani
a-ghorbani deleted the feature/TASK-20260909-1820 branch September 10, 2026 18:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Select one discovered HTP session for PocketPal Hexagon inference

1 participant