Lab · Field review
PocketPal AI on iPhone
Most phone LLM apps hide the dials. PocketPal exposes nearly all of them, which is why it is the first app we are putting through a controlled comparison against our Mac baselines. This page is the setup guide we wished existed: what each setting does, what to change for a repeatable run, and what to leave alone.
Ongoing field review · started 2026-10-02 · lab notebook, dated sections; hands-on results land here as they are measured
Short version 2026-10-02
- PocketPal is the only mainstream iPhone app we found that behaves like an instrument. It loads any GGUF by repository and filename, exposes seed, context length, sampling and KV-cache settings, prints tokens per second and time to first token under every reply, and exports the chat as JSON with the generation settings attached. Free, MIT-licensed, llama.cpp under the hood.
- Its "memory tight" warning is advisory, and static. A 12 GB iPhone does not hand 12 GB to one app, and PocketPal flags any model it estimates above roughly 6.4 GB. Gemma 4 E4B at Q4_K_M was flagged at about 7 GB (4.98 GB weights plus the 990 MB vision projector plus a context allowance), then loaded and ran on the iPhone 18 Pro Max anyway. Turning Vision off did not change the on-screen numbers: the estimate is computed from the files, not from the live configuration. Treat it as a hint, not a gate.
- Four settings decide whether a run is comparable: the exact GGUF file, context size, temperature plus seed, and the thinking toggle. Everything else is taste. Our bench profile: Q4_K_M, 8192 context, temperature 0 with a fixed seed, thinking as the Mac baseline had it.
- The built-in Benchmark screen is a phone-sized llama-bench. Its Default profile (512 prompt tokens, 128 generated, 3 repetitions) is the same pp512 / tg128 shape we run on the Mac bench, so its prefill and decode numbers line up with ours and with vendor claims without conversion.
- Devices: iPhone 18 Pro Max (12 GB, A20 Pro, iOS 27.0.1) now, iPad Pro M5 (1 TB tier, 16 GB) next. No published LLM benchmark exists yet for the A20 Pro, so the first speed readouts here are new data.
What PocketPal is 2026-10-02
PocketPal AI is an open-source
(MIT) React Native app by Asghar Ghorbani that runs GGUF models through
llama.rn, a binding to llama.cpp with Metal acceleration on Apple devices.
Version 1.18.0 shipped on 2026-09-28. Models come from three places: an in-app Hugging
Face search (gated repositories included once you have accepted the license), a curated
"available to download" list, and sideloading any GGUF through the Files app. It can also
act as a client for a remote OpenAI-compatible or llama.cpp server; nothing on this page
uses that mode.
The parts that matter to a tester: a per-message footer with tokens per second, milliseconds per token and time to first token; a benchmark screen; JSON and Markdown export; a seed field; and a context-size field. The llama.cpp build advances with every app release (roughly b10829 as of 1.17.3), so the app version belongs in every run record.
The test rig 2026-10-02
| Device | Memory | Chip | Why it is here |
|---|---|---|---|
| iPhone 18 Pro Max | 12 GB | A20 Pro: 2 nm, 7-core GPU with Neural Accelerators, two 16-core Neural Engines, about 50% more memory bandwidth than A19 Pro per Apple | The phone. Same RAM as the 17 Pro, so published 17 Pro model-fit guidance applies; decode is bandwidth-bound, so expect roughly 1.5x the 17 Pro generation speeds if the runtime scales. Unmeasured by anyone so far. |
| iPad Pro M5 (1 TB) | 16 GB | M5 with the 10-core CPU and 16 GB (the 1 and 2 TB tiers; the 256 and 512 GB tiers get 9 cores and 12 GB). Neural Accelerator in every GPU core, as on the Mac bench. | The same M5 GPU design as our bench host, in a tablet power envelope. Lets us separate "small silicon" from "small memory budget", and iPadOS gives an app a larger memory ceiling than iOS. |
| MacBook Pro M5 Max | 128 GB | Bench host | Holds the controlled baseline for every model we put on the phone, scored on the same suites with the same judge. |
One published data point frames the iPad: a February 2026 MLX benchmark of six sub-1 GB models found the iPhone 17 Pro and iPad Pro M5 close at short context, with the iPhone degrading faster as context grows. We will see whether llama.cpp shows the same shape.
First contact: the memory warning 2026-10-02
Loading Unsloth's Gemma 4 E4B GGUF on the iPhone produced an orange "Memory tight" badge and the line "this model needs 7 GB, but only 6.4 GB is estimated as usable." Three numbers are stacked inside that 7 GB:
| Component | Size | Where it comes from |
|---|---|---|
| Language model weights | 4.98 GB | gemma-4-E4B-it-Q4_K_M.gguf (Unsloth). This is what the model card calls the model. |
| Vision projector | 0.99 GB | mmproj-F16.gguf. Downloaded and loaded only because the Vision toggle is on. PocketPal shows the pair as one 5.97 GB entry. |
| Context and compute buffers | ~1 GB | KV cache for the context size you set (2048 by default), plus llama.cpp's working buffers. Grows with context. |
Why 12 GB of RAM becomes 6.4 GB usable
iOS caps how much memory a single app may hold, well below installed RAM, and kills the app when it crosses the line. Apple does not publish the ceiling per device; apps can request an increased-memory entitlement, and PocketPal's 6.4 GB figure is its own estimate of where the line sits on this phone. So the warning is a prediction about an iOS limit, not a statement that the phone is out of memory. iPadOS is more generous, which is one reason the iPad Pro is in the rig.
Update, later the same day: it loaded anyway
With the warning still showing, the Q4_K_M build loaded and answered on the iPhone 18 Pro Max. Turning Vision off did not change the "needs 7 GB" line at all, which settles what the badge is: an estimate computed from the model files when they are added, not a live measurement of the configuration you are about to load. So it cannot tell you where the iOS kill line is; only loading progressively larger files will. The Q5_K_M (5.48 GB) and Q6_K (7.07 GB) files of the same model are the next probes, and the memory-pressure behaviour is now the first open question below.
What still helps when a model really is too big, in the order we would try it:
- Turn Vision off in the model's card (the eye toggle). The 990 MB projector is not loaded for text work, even though the on-screen estimate does not move. Turn it back on only for a vision test.
- Pick a smaller file of the same model if you need the headroom for a longer context. In the same Unsloth repository:
IQ4_XSat 4.72 GB,Q4_K_Sat 4.84 GB,Q3_K_Mat 4.06 GB. Google's own QAT build (gemma-4-E4B_q4_0-it.gguf, 5.15 GB) is slightly larger than Unsloth's Q4_K_M, not smaller. - Quantize the KV cache (Advanced Settings, cache type K and V to
q8_0). Small gain at 2048 context, larger once you push context to 8K or beyond. - Leave Memory Lock off. Locking pages in RAM makes the ceiling harder, not softer. Keep Memory Mapping on; it lets iOS page weights in from flash instead of copying them.
For our comparison we keep Q4_K_M, because that is the quant our Mac baseline for this model was scored on, and we turn Vision off for the text cells.
The settings, and what they do 2026-10-02
PocketPal keeps its settings in three places, and the split matters: two of them require a model reload to take effect, one applies to the next message.
1. Model Initialization Settings (global; reload to apply)
| Setting | Default | What it does | Bench profile |
|---|---|---|---|
| Device: Auto / Metal / CPU | Auto | Whether llama.cpp runs on the GPU through Metal or on the CPU cores. Auto picks Metal on any modern iPhone. | Metal, set explicitly, so the record says which path ran. CPU is a diagnostic, not a daily mode. |
| Layers on GPU | 99 | How many transformer layers are offloaded to the GPU. 99 means all of them. | 99. Partial offload is a desktop trick for models that do not fit VRAM; on unified memory it only slows things down. |
| Context Size | 2048 | The token window the model can see, prompt plus reply. The KV cache is sized from this at load time. | 8192. Enough for a pasted document and a long answer; matches the fixed 8K ceiling other phone apps use, so results line up. Watch the memory estimate when you raise it. |
| Advanced: Batch Size | 512 | How many prompt tokens llama.cpp processes per step (n_batch). Caps prefill chunking; also the ceiling for the benchmark's prompt-processing size. | 512. Matches the pp512 shape of our Mac llama-bench runs. |
| Advanced: Physical Batch Size | 512 | The micro-batch actually dispatched to the GPU (n_ubatch). Lower values trade prefill speed for less peak memory. | 512. Drop to 256 only if a large model is memory-limited during prefill. |
| Advanced: CPU Threads | 4 of 6 | Threads for the CPU side of the work. With all layers on Metal, mostly tokenization and sampling. | 4. More threads on a phone means more heat for little gain. |
| Advanced: Image Max Tokens | 512 | Token budget for an attached image. Lower is faster and sees less detail. | 512, and only matters with Vision on. |
| Advanced: Flash Attention | Auto (on for Metal) | Fused, memory-efficient attention kernel. | Auto. |
| Advanced: Key / Value Cache Type | F16 / F16 | Precision of the attention cache. q8_0 halves its size with negligible quality cost; q4_0 saves more and starts to cost quality. | F16 for the first pass so nothing but the model is quantized; q8_0 if memory forces it, noted in the run. |
| Advanced: Speculative Decoding | toggle (experimental, added Aug 2026) | Runs a small draft model ahead of the main one and verifies its guesses in bulk. Can raise decode speed on bandwidth-bound hardware; output is unchanged in theory, the speedup is model-pair dependent. | Off for the baseline. A separate speed experiment: Unsloth ships 98 MB MTP draft files alongside the Gemma 4 E4B GGUFs that are built for exactly this. |
| Use Memory Lock | off | Pins model pages in RAM so iOS cannot compress or evict them. | Off. See the memory section. |
| Memory Mapping | on | Maps the GGUF file instead of reading it into a buffer. Faster loads, lower peak. | On. |
2. Model Settings (per model; reload to apply)
| Setting | Default for Gemma 4 E4B | What it does | Bench profile |
|---|---|---|---|
| BOS / EOS | off / off | Whether the app manually inserts beginning- and end-of-sequence tokens. With Jinja templating on, the chat template handles these; forcing them doubles up. | Leave off while Jinja is on. Turn on only for a model whose template is missing or broken. |
| Add Generation Prompt | on | Appends the assistant-turn opener so the model knows it is its turn to speak. | On. |
| System Prompt | empty | Standing instructions sent before every chat. | Empty for the baseline cells; the test prompts carry their own instructions. A fixed system prompt is a separate cell. |
| Template | from the GGUF | The chat format. Read from the file's metadata; editable. | Do not edit. A template edit is a different instrument. |
| Stop words | <turn|> | Strings that end generation. Gemma 4's end-of-turn marker is pre-filled. | Leave as detected. A missing stop word shows up as the model talking to itself; a wrong one truncates answers. |
| Vision toggle and projection model | on, mmproj-F16 | Loads the image encoder alongside the language model. | Off for text cells (saves ~1 GB). On, recorded, for the vision cell. |
| Reasoning model | on (shows the thinking control) | Exposes the thinking on/off control in chat for this model, even when the app did not auto-detect it. | Leave the toggle on so the control is visible. Our Mac baseline for Gemma 4 E4B was scored with thinking on (the model's default), so the comparable cell runs with thinking on. A second cell with thinking off is the phone-realistic one: thinking tokens dominate latency on a phone, and the two together show what the switch costs. |
| Supports graded effort | off | For models that take a low/medium/high effort hint. | Off for Gemma 4; it has a binary thinking switch. |
3. Chat Generation Settings (preset; applies to the next message)
| Setting | Default | What it does | Bench profile | Daily use |
|---|---|---|---|---|
| N Predict | Unlimited | Maximum reply length in tokens. | Custom, 2048. A runaway reply then ends in a measured place instead of a thermal throttle. | Unlimited is fine. |
| Include thinking in context | on | Whether earlier thinking blocks are re-sent on later turns. | Off. Our cells are single-turn; for multi-turn tests this doubles context use for no measured gain. | Off, unless you rely on the model remembering its own reasoning. |
| Temperature | 0.7 | Randomness. 0 picks the most likely token every time. | 0. Removes one source of variance; it is how the Mac baseline was scored. | 0.7 for chat, 0.2 to 0.4 for anything structured. |
| Top K | 40 | Only the K most likely tokens are eligible. | Irrelevant at temperature 0; leave at 40. | 40. |
| Top P | 0.95 | Nucleus sampling: smallest set of tokens whose probability sums to P. | Irrelevant at temperature 0. | 0.95. |
| Min P | 0.05 | Drops tokens below 5% of the top token's probability. The modern replacement for fiddling with Top P. | Irrelevant at temperature 0. | 0.05. If you raise temperature, this is the guardrail that keeps it coherent. |
| XTC threshold / probability | 0.1 / 0 | "Exclude top choices" sampler for creative writing. Probability 0 means off. | Off. | Off unless you write fiction. |
| Typical P | 1 (off) | Locally typical sampling. | Off. | Off. |
| Penalty last N / repeat / frequency / presence | 64 / 1 / 0 / 0 | Repetition controls. Repeat 1.0 means no penalty. | Defaults. A repeat penalty changes what the model would have said; we want to see the loop if there is one. | Defaults. Raise repeat to 1.1 only if a model loops in normal use, and know it costs accuracy on code and lists. |
| Mirostat | Off | Adaptive sampler that targets a perplexity instead of fixed Top K / Top P. | Off. | Off. |
| Seed | -1 (random) | Random number seed. A fixed value makes a run repeatable at a given temperature and build. | Fixed (we use 42) and recorded. | -1. |
| Jinja | on | Uses the GGUF's own Jinja chat template rather than a legacy format guess. Required for current models' tool and thinking tags. | On. | On. |
Save the bench profile as a preset once and reuse it across models, then record the preset name in the run. The export carries the values anyway, which is the point of exporting.
The built-in benchmark 2026-10-02
Easy to miss in the side menu: a Benchmark screen that reports the device (Apple iPhone, iOS 27.0.1 on ours), then runs the loaded model through a fixed prompt-processing and text-generation loop and reports tokens per second for each. It is llama-bench in a phone UI, and its advanced settings expose the same three knobs:
| Knob | Default profile | Fast profile | What it measures |
|---|---|---|---|
| PP, prompt processing tokens | 512 | 128 | Prefill: how fast the model reads a prompt. Compute-bound; this is where GPU matrix accelerators show up. Capped by the Physical Batch Size above. |
| TG, text generation tokens | 128 | 32 | Decode: how fast it writes. Memory-bandwidth-bound; this is the number people quote as "tokens per second". |
| NR, repetitions | 3 | 3 | Runs averaged. Three is enough to see a thermal step between repetitions on a phone. |
The Default profile is pp512 / tg128, the exact shape our Mac bench records with llama-bench and the shape most vendors quote, so a Default run on the phone sits next to the Mac number for the same GGUF with no conversion. Fast is for a quick sanity check after a settings change. How we use it: Default profile, run once from a cool phone, then again straight after the ten-minute sustained generation, and record both. The screen also offers to share results to a community device leaderboard; we keep our runs in the export.
What we are measuring 2026-10-02
- Quality against the Mac baseline. A fixed six-prompt set drawn from our suite families (reasoning, extraction, formatting, an ops task, a JSON function call against a given schema, a long summary of a pasted document), same prompts on the phone and the Mac, same quant where the app allows it. Scored by the same pinned judge from the exported transcripts.
- Burst speed. The Benchmark screen's Default profile (pp512 / tg128) for the Mac-comparable number, plus tokens per second and time to first token from the per-message footer on the summary prompt, cold and warm.
- Sustained speed. One ten-minute continuous generation per model, tokens per second at minute one and minute ten. Published iPhone 17 Pro data shows GPU runtimes losing half their throughput over ten minutes while the Neural Engine path holds about two thirds; we want the A20 Pro number.
- Fit. Which quant of each model loads, which trips the memory warning, which gets killed. A kill is a result, not a retry.
- iPad versus iPhone. The same cells on the iPad Pro M5, which isolates the memory ceiling and the thermal envelope from the model.
Models on the list 2026-10-02
PocketPal's own "available to download" shelf is current: Gemma 4 E2B (4.09 GB with projector), LFM2.5 8B-A1B (5.16 GB), Qwen3.5 4B (3.41 GB), Spark-X2.5 4B (2.6 GB), plus Gemma 3 4B and 1B. Through the Hugging Face search we add Gemma 4 E4B (the Mac baseline pairing), LFM2.5-2.6B (tied first on the only independent iPhone bake-off we know of, Artificial Analysis, August 2026) and, if the bundled llama.cpp build loads it, Bonsai 2 27B, the ternary build that puts 27B-class knowledge in 5.9 GB. Results land in this section as dated updates.
Open questions 2026-10-02
- Where is the iOS kill line on a 12 GB iPhone? PocketPal's estimate flagged 7 GB and the model ran; the estimate is static, so it cannot answer this. Next probes: the Q5_K_M (5.48 GB) and Q6_K (7.07 GB) files of the same model, then the 5.9 GB ternary 27B.
- Does speculative decoding with Unsloth's MTP draft file move the tg128 number on the A20 Pro, and by how much?
- Do the A20 Pro's GPU Neural Accelerators change prefill through llama.cpp's Metal path, as the M5's did on the Mac, or does that wait for a runtime update?
- How much of the iPad's advantage, if any, is the memory ceiling versus sustained thermals?