Download
04 · AI Models

Three tiers. You pick.

Everyday drafting uses the Apple foundation model at the core of Apple Intelligence — on your iPhone, built in, instant, even offline. With iOS 27, Apple's Private Cloud Compute steps in when a turn needs more room than the phone has: the same Apple Intelligence stack, on Apple silicon servers, with a much larger context window and deeper reasoning. And when the work is genuinely long and multi-step, Claude takes over on prepaid AI Credit or your own Anthropic key. You switch in the chat itself, or per feature in Settings.

On your iPhone when…

It's an everyday reply, you want it instantly, you're offline, or nothing should leave your phone. This is the default — most messages never need more.

Private Cloud Compute when…

The thread is longer than the phone can hold at once, or you want your Knowledge Base documents searched — still Apple Intelligence, still no key, just more room to think.

Hand it to Claude when…

The task has steps, you need research with sources, a document produced, or a nuanced tone for a message that matters.

CapabilityApple IntelligenceClaude
Where it runsTwo tiers, both Apple. On-device, powered by the foundation model at the core of Apple Intelligence — no network required, works offline. Or Apple's Private Cloud Compute on iOS 27, the same stack on Apple silicon servers, for the turns that need more room. You pick which in the chatAnthropic cloud — with prepaid AI Credit, or direct to Anthropic on your own API key
What you payBuilt in — no API key needed. The Free plan runs unlimited for your first week, then 5 Apple Intelligence actions a day (shared across Smart Reply, Analyze, Summary, Skills, and AI Writer); every paid plan makes them unlimitedTwo ways in: AI Credit — prepaid top-ups, no subscription and no API key needed — or the Pro plan plus your own key, paying Anthropic directly per use with no markup
ModelsApple Foundation Models — the ~3B on-device model, or Apple's larger server-side model on Private Cloud ComputeOpus 5.5 (most capable for ambitious work), Sonnet 5.5 (most efficient for everyday tasks), Haiku 4.5 (fastest for quick answers)
Context window4,096 tokens on device (input + output combined, ~3,000 words); about 32,000 on Private Cloud Compute. Longer threads are condensed automatically either wayUp to 1,000,000 tokens (Opus/Sonnet) or 200,000 (Haiku) — handles entire conversation histories. The same on AI Credit and on your own key
Input typesText only — images, PDFs, and documents are handled by ClaudeMultimodal — understand and analyse a wide range of visual formats, including photos, charts, graphs, and technical diagrams
Extended thinkingNot on the on-device model, which is tuned for direct, single-pass responses. Private Cloud Compute brings the deeper reasoning the larger model allowsEnhanced reasoning capabilities for complex tasks — Claude's step-by-step thought process is visible before it delivers its final answer, with adaptive thinking and adjustable effort (Low, Medium, High, Extra, Max)
Reasoning and logicOn device, best for simple focused tasks like everyday drafting rather than heavy logic or math. Private Cloud Compute takes the heavier ones — more context held at once, and reasoning to go with itStrong complex reasoning, multi-step problem solving, code generation, and mathematical tasks
Tool useNearly all of them, running through Apple's own frameworks — web search, URL fetching, contacts, messages and chat discovery, calendar, reminders, location, files, and image generation. Knowledge Base document search needs Private Cloud Compute; on the on-device model only saved links can be opened. The Claude-native ones — Task Tracking, Memory, Sub-agents, Code Execution, and Agent Skills — sit outside the Apple setThe full set. Claude decides when to call a tool from your request and the tool's description, then AI Reply Assistant executes it on your device. It adds the agent-only tools on top of everything the on-device set already has. Files and Agent Skills file creation (PPTX, XLSX, PDF, DOCX) need your own Anthropic key
Agentic behaviourSingle-turn — one prompt, one responseClaude dynamically directs its own processes and tool usage, maintaining control over how it accomplishes tasks — planning, calling tools, observing results, and refining across multiple iterations
Document creationNot availablePre-built Agent Skills let Claude run code in a sandboxed container to create PowerPoint presentations, Excel spreadsheets with charts, Word documents, and PDF reports — downloaded and delivered as shareable attachments. Requires your own Anthropic key
Message draftingGenerates up to 3 drafts in different tones — you pick the one to use before anything is sentHuman-in-the-loop — Claude drafts messages as cards, you review, choose recipients, and confirm before anything is sent
Context managementWhen a thread outgrows the context window, the transcript is automatically condensed and the conversation continues — no error, nothing lostCompaction extends effective context length by automatically summarising older context, keeping the active context focused and performant. Prompt caching reduces processing time and costs for repetitive tasks. On-demand tool loading and real-time token tracking
Languages9 languages (English, French, German, Italian, Portuguese, Spanish, Japanese, Korean, Chinese)Broad multilingual support across dozens of languages
Factual knowledgeA compact on-device model — tuned for everyday drafting rather than broad factual recall or current eventsBroader world knowledge with more current training data — stronger on factual and domain-specific questions
PrivacyNothing leaves your phone — maximum privacyYou control when Claude is used. On your own key, drafts go directly to Anthropic; with AI Credit they pass through Flowbie's relay, which stores nothing.
Best forEveryday replies, quick suggestions, full privacy, offline useEven more expertise — longer drafts, nuanced tone, complex reasoning, multi-step tasks, document creation, and research

Which Claude, specifically

All three, on either route into Claude — models are never gated by plan or by how you pay, and the context column holds either way, so AI Credit runs every model at its full native window exactly as your own key does. Switch mid-conversation and the next response uses your selection. All three support Extended Thinking. We move to Anthropic’s newest release in each tier as it ships, so these are the current models rather than a fixed list.

ModelStrengthContextBest for
Haiku 4.5Fastest, most cost-efficient200KQuick drafting
Sonnet 5.5Best balance of speed + intelligence1MEveryday and complex tasks
Opus 5.5Most capable for deep reasoning1MComplex, long-running work

Switching is one setting away

Settings → Chat → AI Features. Smart Reply and Skill each pick Apple Intelligence or Cloud AI independently, and a model row chooses the Claude model used for cloud drafts. Cloud AI only becomes selectable once Claude can serve a request — either a prepaid AI Credit balance or your own Anthropic key — and any feature switches back to Apple Intelligence anytime. On the Apple side there is nothing to configure: a turn runs on Private Cloud Compute when it can serve and on device when it cannot, and in the agent chat you can pin the tier yourself from the model picker.

AI Features
Smart Reply
Apple Intelligence or Cloud Model
Cloud AI
Skill
Apple Intelligence or Cloud Model
Cloud AI
Cloud Model
Model used for cloud drafts
Sonnet 5.5
Off by default. Cloud Model runs on prepaid AI Credit or your own Anthropic key, stored in iOS Keychain — switch any feature back to Apple Intelligence anytime.

Writing for each engine

The table above is what each one can do. This is what that means for the instructions you write in a Skill: the same prompt does not get the best out of both.

When you writeApple IntelligenceClaude
Prompt lengthKeep it short — 2–4 sentences. The on-device model has a 4,096-token context window (input + output combined), so every word counts.Can handle longer, more detailed prompts — up to 1M tokens on Sonnet 5.5 and Opus 5.5, 200K on Haiku 4.5, whichever way you pay. Add examples, edge cases, and nuanced instructions.
ExamplesUse 1–2 simple examples. Too many or too complex will consume the limited context and cause parroting.Use 3–5 diverse examples for best results. Claude generalises well from richer input.
Conditional logicKeep conditionals simple. The on-device model handles if-this-then-that branching poorly — reach for the tone sliders instead of writing rules.Can handle moderate conditional complexity — "If the message is a complaint, be empathetic; if it's a question, be direct."
ReasoningDon't ask it to think step by step. At this size the reasoning tends to leak into the reply itself, and the logic is unreliable anyway.Enable Extended Thinking for complex tasks. Claude handles chain-of-thought reasoning natively with visible, adjustable thinking.
Persona depthA single sentence for role + persona is usually enough. Longer persona prompts eat into the 4K token budget.You can build richer personas with backstory, communication preferences, and relationship dynamics.
Images and documentsNot supported — the on-device model is text-only. No images, PDFs, or file attachments.Include screenshots, photos, charts, and PDFs. Claude analyses visual content alongside text for richer context.

Which entry uses which is set per Skill, so a quick everyday Prompt and a research-heavy Skill can run on different engines — templates and the mistakes to avoid are on the Skills page.

What it costs

Apple Intelligence runs on every plan with no per-message fees — the Free plan runs unlimited for your first week, then 5 Apple AI actions a day (Smart Reply, Analyze, Summary, Skills, and AI Writer), and every paid plan makes them unlimited. That covers both Apple tiers: Private Cloud Compute costs nothing extra here, and is metered by Apple against your own device's allowance rather than by us — if you reach it, the app says so and keeps working on device. Claude comes two ways. AI Credit works on any plan: buy a prepaid top-up (from US$4.49), sign in with Apple, and spend it as you go — no subscription, no API key, and it never expires. Or subscribe — Solo US$19.99/year or US$2.99 monthly, Duo US$32.99/year or US$4.99 monthly, Pro US$49.99/year or US$7.49 monthly, each with a 7-day free trial — and bring your own Anthropic API key: you pay Anthropic's standard usage rates directly with no markup, can cap spending in your Anthropic console, and add Managed Agents, Agent Skills, and Files. Bringing a key works on all three plans; what they differ on is how many messaging accounts you connect at once. Context is the same either way — each model runs at its full native window, so Sonnet 5.5 and Opus 5.5 give you 1M tokens whether you paid with credit or a key. Subscribing also puts AI Credit in your balance once your first bill goes through — US$1.00 on Solo, US$1.50 on Duo, US$2.00 on Pro, once per Apple Account — so you can try Claude without buying a top-up. Either way, a typical drafted reply uses only a few thousand tokens, and you choose the model per task — Haiku for fast, inexpensive drafting up to Opus for the hardest work.

Compare Free, Solo, Duo and Pro →

Your personal AI assistant for messaging.

Download AI Reply Assistant on iOS and try it free.