Lab · Field review

PocketPal AI on iPhone

Most phone LLM apps hide the dials. PocketPal exposes nearly all of them, which is why it is the first app we are putting through a controlled comparison against our Mac baselines. This page is the setup guide we wished existed: what each setting does, what to change for a repeatable run, and what to leave alone.

Ongoing field review · started 2026-10-02 · lab notebook, dated sections; hands-on results land here as they are measured

Short version 2026-10-02

What PocketPal is 2026-10-02

PocketPal AI is an open-source (MIT) React Native app by Asghar Ghorbani that runs GGUF models through llama.rn, a binding to llama.cpp with Metal acceleration on Apple devices. Version 1.18.0 shipped on 2026-09-28. Models come from three places: an in-app Hugging Face search (gated repositories included once you have accepted the license), a curated "available to download" list, and sideloading any GGUF through the Files app. It can also act as a client for a remote OpenAI-compatible or llama.cpp server; nothing on this page uses that mode.

The parts that matter to a tester: a per-message footer with tokens per second, milliseconds per token and time to first token; a benchmark screen; JSON and Markdown export; a seed field; and a context-size field. The llama.cpp build advances with every app release (roughly b10829 as of 1.17.3), so the app version belongs in every run record.

The test rig 2026-10-02

DeviceMemoryChipWhy it is here
iPhone 18 Pro Max 12 GB A20 Pro: 2 nm, 7-core GPU with Neural Accelerators, two 16-core Neural Engines, about 50% more memory bandwidth than A19 Pro per Apple The phone. Same RAM as the 17 Pro, so published 17 Pro model-fit guidance applies; decode is bandwidth-bound, so expect roughly 1.5x the 17 Pro generation speeds if the runtime scales. Unmeasured by anyone so far.
iPad Pro M5 (1 TB) 16 GB M5 with the 10-core CPU and 16 GB (the 1 and 2 TB tiers; the 256 and 512 GB tiers get 9 cores and 12 GB). Neural Accelerator in every GPU core, as on the Mac bench. The same M5 GPU design as our bench host, in a tablet power envelope. Lets us separate "small silicon" from "small memory budget", and iPadOS gives an app a larger memory ceiling than iOS.
MacBook Pro M5 Max 128 GB Bench host Holds the controlled baseline for every model we put on the phone, scored on the same suites with the same judge.

One published data point frames the iPad: a February 2026 MLX benchmark of six sub-1 GB models found the iPhone 17 Pro and iPad Pro M5 close at short context, with the iPhone degrading faster as context grows. We will see whether llama.cpp shows the same shape.

First contact: the memory warning 2026-10-02

Loading Unsloth's Gemma 4 E4B GGUF on the iPhone produced an orange "Memory tight" badge and the line "this model needs 7 GB, but only 6.4 GB is estimated as usable." Three numbers are stacked inside that 7 GB:

ComponentSizeWhere it comes from
Language model weights4.98 GBgemma-4-E4B-it-Q4_K_M.gguf (Unsloth). This is what the model card calls the model.
Vision projector0.99 GBmmproj-F16.gguf. Downloaded and loaded only because the Vision toggle is on. PocketPal shows the pair as one 5.97 GB entry.
Context and compute buffers~1 GBKV cache for the context size you set (2048 by default), plus llama.cpp's working buffers. Grows with context.

Why 12 GB of RAM becomes 6.4 GB usable

iOS caps how much memory a single app may hold, well below installed RAM, and kills the app when it crosses the line. Apple does not publish the ceiling per device; apps can request an increased-memory entitlement, and PocketPal's 6.4 GB figure is its own estimate of where the line sits on this phone. So the warning is a prediction about an iOS limit, not a statement that the phone is out of memory. iPadOS is more generous, which is one reason the iPad Pro is in the rig.

Update, later the same day: it loaded anyway

With the warning still showing, the Q4_K_M build loaded and answered on the iPhone 18 Pro Max. Turning Vision off did not change the "needs 7 GB" line at all, which settles what the badge is: an estimate computed from the model files when they are added, not a live measurement of the configuration you are about to load. So it cannot tell you where the iOS kill line is; only loading progressively larger files will. The Q5_K_M (5.48 GB) and Q6_K (7.07 GB) files of the same model are the next probes, and the memory-pressure behaviour is now the first open question below.

What still helps when a model really is too big, in the order we would try it:

For our comparison we keep Q4_K_M, because that is the quant our Mac baseline for this model was scored on, and we turn Vision off for the text cells.

The settings, and what they do 2026-10-02

PocketPal keeps its settings in three places, and the split matters: two of them require a model reload to take effect, one applies to the next message.

1. Model Initialization Settings (global; reload to apply)

SettingDefaultWhat it doesBench profile
Device: Auto / Metal / CPUAutoWhether llama.cpp runs on the GPU through Metal or on the CPU cores. Auto picks Metal on any modern iPhone.Metal, set explicitly, so the record says which path ran. CPU is a diagnostic, not a daily mode.
Layers on GPU99How many transformer layers are offloaded to the GPU. 99 means all of them.99. Partial offload is a desktop trick for models that do not fit VRAM; on unified memory it only slows things down.
Context Size2048The token window the model can see, prompt plus reply. The KV cache is sized from this at load time.8192. Enough for a pasted document and a long answer; matches the fixed 8K ceiling other phone apps use, so results line up. Watch the memory estimate when you raise it.
Advanced: Batch Size512How many prompt tokens llama.cpp processes per step (n_batch). Caps prefill chunking; also the ceiling for the benchmark's prompt-processing size.512. Matches the pp512 shape of our Mac llama-bench runs.
Advanced: Physical Batch Size512The micro-batch actually dispatched to the GPU (n_ubatch). Lower values trade prefill speed for less peak memory.512. Drop to 256 only if a large model is memory-limited during prefill.
Advanced: CPU Threads4 of 6Threads for the CPU side of the work. With all layers on Metal, mostly tokenization and sampling.4. More threads on a phone means more heat for little gain.
Advanced: Image Max Tokens512Token budget for an attached image. Lower is faster and sees less detail.512, and only matters with Vision on.
Advanced: Flash AttentionAuto (on for Metal)Fused, memory-efficient attention kernel.Auto.
Advanced: Key / Value Cache TypeF16 / F16Precision of the attention cache. q8_0 halves its size with negligible quality cost; q4_0 saves more and starts to cost quality.F16 for the first pass so nothing but the model is quantized; q8_0 if memory forces it, noted in the run.
Advanced: Speculative Decodingtoggle (experimental, added Aug 2026)Runs a small draft model ahead of the main one and verifies its guesses in bulk. Can raise decode speed on bandwidth-bound hardware; output is unchanged in theory, the speedup is model-pair dependent.Off for the baseline. A separate speed experiment: Unsloth ships 98 MB MTP draft files alongside the Gemma 4 E4B GGUFs that are built for exactly this.
Use Memory LockoffPins model pages in RAM so iOS cannot compress or evict them.Off. See the memory section.
Memory MappingonMaps the GGUF file instead of reading it into a buffer. Faster loads, lower peak.On.

2. Model Settings (per model; reload to apply)

SettingDefault for Gemma 4 E4BWhat it doesBench profile
BOS / EOSoff / offWhether the app manually inserts beginning- and end-of-sequence tokens. With Jinja templating on, the chat template handles these; forcing them doubles up.Leave off while Jinja is on. Turn on only for a model whose template is missing or broken.
Add Generation PromptonAppends the assistant-turn opener so the model knows it is its turn to speak.On.
System PromptemptyStanding instructions sent before every chat.Empty for the baseline cells; the test prompts carry their own instructions. A fixed system prompt is a separate cell.
Templatefrom the GGUFThe chat format. Read from the file's metadata; editable.Do not edit. A template edit is a different instrument.
Stop words<turn|>Strings that end generation. Gemma 4's end-of-turn marker is pre-filled.Leave as detected. A missing stop word shows up as the model talking to itself; a wrong one truncates answers.
Vision toggle and projection modelon, mmproj-F16Loads the image encoder alongside the language model.Off for text cells (saves ~1 GB). On, recorded, for the vision cell.
Reasoning modelon (shows the thinking control)Exposes the thinking on/off control in chat for this model, even when the app did not auto-detect it.Leave the toggle on so the control is visible. Our Mac baseline for Gemma 4 E4B was scored with thinking on (the model's default), so the comparable cell runs with thinking on. A second cell with thinking off is the phone-realistic one: thinking tokens dominate latency on a phone, and the two together show what the switch costs.
Supports graded effortoffFor models that take a low/medium/high effort hint.Off for Gemma 4; it has a binary thinking switch.

3. Chat Generation Settings (preset; applies to the next message)

SettingDefaultWhat it doesBench profileDaily use
N PredictUnlimitedMaximum reply length in tokens.Custom, 2048. A runaway reply then ends in a measured place instead of a thermal throttle.Unlimited is fine.
Include thinking in contextonWhether earlier thinking blocks are re-sent on later turns.Off. Our cells are single-turn; for multi-turn tests this doubles context use for no measured gain.Off, unless you rely on the model remembering its own reasoning.
Temperature0.7Randomness. 0 picks the most likely token every time.0. Removes one source of variance; it is how the Mac baseline was scored.0.7 for chat, 0.2 to 0.4 for anything structured.
Top K40Only the K most likely tokens are eligible.Irrelevant at temperature 0; leave at 40.40.
Top P0.95Nucleus sampling: smallest set of tokens whose probability sums to P.Irrelevant at temperature 0.0.95.
Min P0.05Drops tokens below 5% of the top token's probability. The modern replacement for fiddling with Top P.Irrelevant at temperature 0.0.05. If you raise temperature, this is the guardrail that keeps it coherent.
XTC threshold / probability0.1 / 0"Exclude top choices" sampler for creative writing. Probability 0 means off.Off.Off unless you write fiction.
Typical P1 (off)Locally typical sampling.Off.Off.
Penalty last N / repeat / frequency / presence64 / 1 / 0 / 0Repetition controls. Repeat 1.0 means no penalty.Defaults. A repeat penalty changes what the model would have said; we want to see the loop if there is one.Defaults. Raise repeat to 1.1 only if a model loops in normal use, and know it costs accuracy on code and lists.
MirostatOffAdaptive sampler that targets a perplexity instead of fixed Top K / Top P.Off.Off.
Seed-1 (random)Random number seed. A fixed value makes a run repeatable at a given temperature and build.Fixed (we use 42) and recorded.-1.
JinjaonUses the GGUF's own Jinja chat template rather than a legacy format guess. Required for current models' tool and thinking tags.On.On.

Save the bench profile as a preset once and reuse it across models, then record the preset name in the run. The export carries the values anyway, which is the point of exporting.

The built-in benchmark 2026-10-02

Easy to miss in the side menu: a Benchmark screen that reports the device (Apple iPhone, iOS 27.0.1 on ours), then runs the loaded model through a fixed prompt-processing and text-generation loop and reports tokens per second for each. It is llama-bench in a phone UI, and its advanced settings expose the same three knobs:

KnobDefault profileFast profileWhat it measures
PP, prompt processing tokens512128Prefill: how fast the model reads a prompt. Compute-bound; this is where GPU matrix accelerators show up. Capped by the Physical Batch Size above.
TG, text generation tokens12832Decode: how fast it writes. Memory-bandwidth-bound; this is the number people quote as "tokens per second".
NR, repetitions33Runs averaged. Three is enough to see a thermal step between repetitions on a phone.

The Default profile is pp512 / tg128, the exact shape our Mac bench records with llama-bench and the shape most vendors quote, so a Default run on the phone sits next to the Mac number for the same GGUF with no conversion. Fast is for a quick sanity check after a settings change. How we use it: Default profile, run once from a cool phone, then again straight after the ten-minute sustained generation, and record both. The screen also offers to share results to a community device leaderboard; we keep our runs in the export.

What we are measuring 2026-10-02

Models on the list 2026-10-02

PocketPal's own "available to download" shelf is current: Gemma 4 E2B (4.09 GB with projector), LFM2.5 8B-A1B (5.16 GB), Qwen3.5 4B (3.41 GB), Spark-X2.5 4B (2.6 GB), plus Gemma 3 4B and 1B. Through the Hugging Face search we add Gemma 4 E4B (the Mac baseline pairing), LFM2.5-2.6B (tied first on the only independent iPhone bake-off we know of, Artificial Analysis, August 2026) and, if the bundled llama.cpp build loads it, Bonsai 2 27B, the ternary build that puts 27B-class knowledge in 5.9 GB. Results land in this section as dated updates.

Open questions 2026-10-02

← Back to the Lab