Skip to content

Instantly share code, notes, and snippets.

@samhenrigold
Last active September 27, 2026 19:19
Show Gist options
  • Select an option

  • Save samhenrigold/e958109a1774c8a466c0ae33e1fb6816 to your computer and use it in GitHub Desktop.

Select an option

Save samhenrigold/e958109a1774c8a466c0ae33e1fb6816 to your computer and use it in GitHub Desktop.
Light Touch iOS port — feasibility study and prior experiments (see https://github.com/samhenrigold/LightTouchMac/issues/9)

Running Light Touch on an iPhone — feasibility study and prior experiments

Status: viable but blocked on JIT. This is a record of what was measured and tried so nobody has to re-derive it.

The appeal is obvious — an iPod touch 2G emulated on a phone, driven by touch, is the device in something close to its original form. The question was never whether it would be nice, but whether it would be fast enough and whether iOS would allow it.

The performance question, measured rather than argued

QEMU's TCG needs to write executable memory, which iOS forbids for ordinary apps. So the number that decides everything is: how slow is QEMU without JIT?

Rather than quote someone else's benchmark, an interpreter-only QEMU was built from the Light Touch tree (--enable-tcg-interpreter) and the same boot measured through it. Host was an M4 Max under load 7–15 from other work, so these are conservative. All three runs reach a pixel-identical 3.1.3 lock screen.

JIT run 1 JIT run 2 interpreter
boot, wall clock 39.0 s 24.7 s 115.9 s
boot, CPU-seconds 30.89 19.37 101.58
idle, % of one core 4.98 4.43 24.8
driven, % of one core 10.9 6.49 15.0

Measured penalty for losing JIT: 3.3–5.2×. Boot CPU-seconds is the defensible comparison; the interpreter's idle figure is contaminated by post-boot settling work spilling into the measurement window. This is plain TCI — UTM SE ships TCTI, which is faster — so ~4× is an upper bound. It agrees with UTM's own published 3.5–6.7× range.

Browser / WASM: tried, 24× too slow

A previous study found a browser/WASM port ~24× too slow and abandoned it. That number does not carry over to a native iOS build: it stacked a browser, WASM, and no native code generation. Native ARM64 interpretation is 4×.

Scaling to a phone

An A19 Pro is ~0.90 of an M4 Max core single-threaded, and sustained thermal derate is ~0.75, so call it 0.68 M4-Max-core-equivalents.

  • With JIT: boot 29–45 s, interactive 10–16% of one core — 6–10× headroom.
  • Without JIT: boot ~2m50s, interactive 22% of one core.

Both are usable. And Light Touch already has working save/restore (migrate file: plus -incoming file:, 86 MB in ~1.1 s), so an app could ship a post-boot snapshot and skip the boot entirely.

JIT: available, but on a leash

  • StikDebug works on iOS 17.4 through 26.x. get-task-allow plus ptrace(PT_TRACE_ME) sets CS_DEBUGGED, which lifts the dynamic-codesigning gate on PROT_EXEC. Needs a device-specific pairing file made once on a computer, and must be re-armed on every launch.
  • TrollStore is dead — the CoreTrust bug was patched in iOS 17.0.1.
  • The EU browser-engine entitlement does not apply; it is for browser engines, with a paid account and Apple approval.

This is the real blocker — not speed, and not whether JIT exists. It is that JIT depends on third-party tooling, re-armed every launch, on top of a 7-day free certificate ($99/year buys a year), and that arrangement tends to break with each iOS release. It is survivable only because the fallback is 4× rather than 40×.

Graphics: iOS is a better host than macOS

In Xcode's device support for an iPhone 17, OpenGLES.framework still exports 457 GL entry points, including the full fixed-function set — glMatrixMode, glOrthof, glTexEnvf, glEnableClientState, glLightfv, and so on — and the iOS 27 SDK ships matching ES1 headers with an arm64e .tbd.

The guest API is GLES 1.1. So on iOS there is no Metal translation layer and no fixed-function emulation needed — a closer match than the Mac's OpenGL 2.1 compatibility profile, which is what the GLES host bridge targets today.

That host bridge is the only macOS-specific file. Its CGL portion is ~40 lines (CGLChoosePixelFormat/CGLCreateContext/CGLSetCurrentContext becoming EAGLContext), plus *EXT → *OES FBO suffixes.

Sensors

State Effort
Touch Already takes normalised coordinates and calls multitouch handlers Near-free; a UIKit handler calls those three functions
Multitouch One FingerData, scalar touch_x/touch_y. Pinch and rotate do not work. Contained but real — the wire header already has numFingers/fingerDataLen
Accelerometer Modelled LIS302DL. X is inverted versus UIKit; Y and Z are not. Near-free: x = -a.x*0x40, y = a.y*0x40, z = a.z*0x40
Audio out Host path proven, but the guest never starts the transfer Blocked on an open bug
Microphone Unbuilt — no AUD_open_in, SWVoiceIn or AUD_read anywhere A whole new capture path

One accelerometer gotcha: the orientation handler also rotates the host window. On a phone that should be suppressed — the physical device is already rotating.

The rebase

This fork is QEMU 8.2.0; UTM's iOS host QEMU is 10.0.2, so a port means rebasing onto utmapp/qemu. That is less alarming than it sounds. git diff --shortstat v8.2.0 HEAD is 170 files, 29,657 insertions and 214 deletions — almost purely additive. Outside our own files, the only upstream code touched is hw/dma/pl080.c, hw/char/exynos4210_uart.c, one line of ui/sdl2.c and one of include/qemu/osdep.h.

Nothing in tcg/, accel/ or target/arm/ — so both the rebase and swapping the TCG backend are low-risk.

Routes tried

Approach Result
Browser / WASM ~24× too slow, abandoned
TCI interpreter (native) ~4× penalty, usable but not great for boot
JIT via debugger attach (StikDebug) Works, but requires re-arming every launch
TrollStore (persistent JIT) Dead on iOS 17.0.1+

Smallest experiment that would settle it

The expensive half is already done: the interpreter multiplier for this exact workload. What remains is device-side — rebase the machine onto utmapp/qemu 10.0.2, build a minimal iOS app with the NAND/NOR/bootrom in the bundle, blit the DisplaySurface into a UIView, and time a boot with and without JIT.

Roughly 1–2 weeks for someone who has shipped an iOS build before. The risk sits in the rebase, not the app.


Addendum (2026-09-27): no JIT needed, because the guest has none either

Three facts measured against the shipping 3.1.3 (7E18) image:

  1. iPhone OS 3.1.3 on ARMv6 never generates code at runtime. JavaScriptCore lives in the dyld shared cache with zero ExecutableAllocator/JITStub strings (JSC's JIT never supported ARMv6); the kernelcache has no dynamic-codesigning or cs_enforcement strings; MobileSafari carries no JIT entitlement (Nitro came in 4.3). Every instruction the guest will ever execute is on the NAND image before the app is signed.
  2. The whole executable world is ~93 MiB. Shared cache r-x mapping 67.7 MiB (273 images), 143 standalone Mach-Os 23.7 MiB, kernel __TEXT 1.9 MiB.
  3. A 10-minute session at the lock screen touched 570,681 live TBs, avg 12 bytes guest / 354 bytes host (27.6× expansion), zero cross-page TBs, 42% direct-jump chained. That is ~6.8 MB of distinct guest code — 7% of the shipped text — and ~200 MB of native code if TCG's verbose output is kept.

Bake the JIT

Tier 1 — AOT-TCG. Run the guest on the Mac under the existing JIT across a corpus (boot, every stock app, the Legacy Store catalog). Record every TB as (4K page content hash, offset, mode flags). Re-emit those TBs with the existing aarch64 TCG backend into a relocatable object and link it into the app's __TEXT. TCG's output is almost position-independent already: helper calls are PC-relative BL or a MOVZ/MOVK literal (tcg_out_call_int), TB chaining is a 26-bit B patched by tb_target_set_jmp_target, and the epilogue is one absolute tb_ret_addr. Three relocation kinds. At runtime tb_gen_code() becomes "hash the page on TLB fill → table lookup → native entry".

Result: today's Mac JIT speed on the phone with no writable+executable page anywhere. No pairing, no debugger, no entitlement, works under Lockdown Mode, and it is App-Store-shaped because it is just code in the signed binary. The ~2M TB invalidations observed are page recycling (a code page freed and reused for data), which content hashing handles for free. Ship the post-boot snapshot too and boot is ~1 s.

Tier 2 — the miss path (user-installed IPAs' own __TEXT). Not TCI. A direct-threaded template interpreter: precompiled native stubs for every ARMv6/Thumb instruction form, operands pre-decoded into a side table, each stub tail-branching to the next, plus superinstructions mined from the corpus for common 2–3-grams. Roughly 2–3× faster than TCI. Frameworks (the 67 MiB) are still Tier 1, so only the app's own game loop pays.

Tier 2.5 — the learning loop. Devices record uncovered (page-hash, offset) pairs; each app update ships AOT packs for the top-N Legacy Store titles. Rosetta's playbook; coverage converges.

Tier 3 — the long version. TCG's 12-byte TBs and 27× expansion come from translating blindly. A real static recompiler (LLVM, whole-function, Mach-O symbols + objc metadata + corpus as seeds) over the 93 MiB would give ~3–5× expansion and better-than-JIT code. Months, not weeks.

The one real-JIT route that needs no computer

WKWebView. The WebContent process is Apple's own and keeps its JIT under iOS 26's Trusted Execution Monitor, and App Review guideline 2.5.2 explicitly permits code run by WebKit. Kohei Tokunaga's wasm32 TCG backend compiles hot TBs to Wasm modules and runs cold ones on a forked TCI. It is not in upstream QEMU master (tcg/ has no wasm32) but lives in ktock/qemu-wasm (patch series: tcg: Add WebAssembly backend). The earlier 24× WASM number was TCI-inside-Wasm; this is JIT-inside-Wasm. Expect 30–60% of native TCG, with WebGL doing GLES 1.1 fixed-function in shaders. Quickest to prototype; Tier 1 beats it on speed and doesn't depend on WebKit.

What's dead

  • StikDebug-style JIT: iOS 26's TXM requires the debugger to touch every page of the code region on every launch, and it broke again in 26.4/26.6 (iCube #11, PiunikaWeb).
  • Metal's runtime shader compiler is a legal JIT, but a single GPU lane running branchy scalar code with ~100-cycle memory latency doesn't reach a 533 MHz ARM11 equivalent.
  • A10 and earlier silicon can execute AArch32, but iOS won't run a 32-bit process. Hypervisor.framework doesn't exist on iOS.

The 3-day experiment that settles it, on a Mac

Dump the TB corpus from one session to a .o, relink qemu-system-arm with the JIT buffer mprotected read-only so any codegen attempt crashes, and boot from the frozen blob. If it reaches the lock screen in ~25 s, the concept is proven before an iPhone is involved. The rebase onto utmapp/qemu and the EAGL/UIKit shell are the boring part after that.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment