Download
03 · The agent chat

A full conversation,
on the model you pick.

The three buttons above the keyboard work inside one thread, on one reply. For work bigger than that, the app has a dedicated agent chat — a standing conversation that runs on whichever model you choose. Apple Intelligence handles it on your iPhone, with no key and no credit to set up; Private Cloud Compute takes a turn that needs a bigger context window; Claude takes the long multi-step work, on prepaid AI Credit or your own Anthropic key. Ask it anything, with your own documents, the web, and your message history already loaded, and send the answer straight back into the chat it came from.

Everything is set per conversation, right in the chat: the model — and, on Claude, how hard it thinks — which tools the agent may use, the active Skill, and, most importantly, the scope. You pick which of your conversations the agent may read, and everything outside that selection is invisible to it. Locked chats never appear at all.

Apple Intelligence or ClaudeSees only the chats you scopeLocked chats: never

Know your way around

Everything in the agent chat is one tap away. Tap a feature to see where it lives and what it opens.

Sonnet 5.5
Apple IntelligenceRuns privately on this device
Private Cloud ComputeApple's private cloud — larger context and deeper reasoning
Opus 5.5Most capable for ambitious work
Sonnet 5.5Most efficient for everyday tasks
Haiku 4.5Fastest for quick answers
Adaptive thinkingClaude decides when to think deeply
Thinking EffortDeep reasoning for complex problemsHigh
Chat with Claude
🧭The agent, control by control

What it adds over an in-thread draft

Most of it is the chat itself, and holds on any engine you pick.

A standing thread, not one replyTopics, each with its own contextScoped to the chats you chooseYour Skills and your tools“Use as Reply” from any chat

The rest is the Claude path. There is no on-device counterpart for any of it, so while Apple Intelligence is selected the app removes those controls from the chat rather than greying them out — the sheets are simply shorter.

Extended Thinking, visibleSub-agents in parallelPowerPoint, Excel, PDF & Word creation — own keyCited sourcesMemory and attachmentsUp to 1M tokens of context on Sonnet 5.5 and Opus 5.5

The model is yours to pick, in the chat itself, and the chat remembers which one you used last — with no Claude connection set up it opens on Apple Intelligence. Once a conversation has started on Claude it stays there; start a new topic to go back to Apple Intelligence. What each one costs, and how they compare →

It proposes. You approve.

Wherever the AI runs — the Skill button or the agent chat — it can't touch your conversations or your device directly. When it wants to act, it must produce an action card — a structured preview of exactly what would be sent or created. A drafted message lands in the composer of the chat it is addressed to, so the send stays yours; the other messaging cards open a recipient picker; device cards file into iOS with a single tap; Agent Skills cards carry the documents Claude creates — and nothing moves until you say so. Nobody on the other end ever receives something you have not read.

Renders on deviceNothing sent until you confirmPer-recipient delivery status
Let Sarah know I'll be 15 minutes late to the standup.
Here's a draft — it goes nowhere until you send it.
Draft Message
Sarah Chen
Running ~15 minutes late to standup — start without me and I'll catch up on the notes. 🙏
Use as Reply
✅The draft lands in your composer — you send it

Every card it can put in front of you

Four groups, by what a card does when you tap it. Every one is a full preview before anything happens, every one shows its own status as it works, and every one stays in the thread afterwards, so you can look back at what was sent. Tap a card to see what it holds.

Something to send. A draft opens the composer of the chat it is for; the rest open the recipient picker.

The agent chat — questions, answered.

The context window is all the text a model can reference when generating a response, including the response itself — a working memory for the model. A larger window handles longer, more complex prompts, but more context isn't automatically better, so AI Reply Assistant actively manages it for you:

  • Compaction — automatically summarises older context as you near the limit, keeping the active context focused and performant
  • Prompt caching — resumes from cached prefixes, significantly cutting processing time and cost on repetitive tasks
  • On-demand tool loading — tool definitions load only when needed rather than all at once, preserving working memory
  • Attachment deduplication — avoids re-sending the same file

When a conversation approaches its context limit, AI Reply Assistant prompts you to compact or start a fresh topic.

This describes the Claude path, where the window runs to 1M tokens on Sonnet 5.5 and Opus 5.5. Apple Intelligence works in a much smaller window on your iPhone — Private Cloud Compute is the tier to switch to when a turn needs more room than that.

MCP (Model Context Protocol) is an open standard for connecting AI assistants to the systems where data lives — content repositories, business tools, and development environments — replacing fragmented integrations with a single protocol. Think of MCP like a USB-C port for AI: just as USB-C standardises how you connect electronic devices, MCP standardises how AI applications connect to external systems.

MCP is designed primarily for desktop and server environments. AI Reply Assistant takes a complementary approach: instead of connecting to external MCP servers, the agent runs natively on iOS and executes tools directly through Apple's own frameworks (EventKit, Contacts, and others) inside the iOS app sandbox.

The result: your data benefits from iOS-level permissions, every action requires your explicit approval, and tools run with the same security guarantees as any native app.

Extended Thinking gives Claude enhanced reasoning for complex tasks, with visibility into its step-by-step thought process before the final answer. It is not a different model — it lets the very same model give itself more time and expend more effort to reach an answer.

You set the effort level to control how much time Claude spends on a problem — low, medium, high, or extra high; higher effort can improve quality through more thorough analysis. Opus and Sonnet use adaptive thinking that automatically adjusts effort to the task; Haiku uses a fixed budget.

A collapsible "Thinking" block with an elapsed timer shows a summary of the reasoning as it happens, making Claude's answers easier to understand and check.

When Claude searches the web or reads one of your documents, superscript markers appear in the answer where it drew on something. Tap the Sources pill under the response for the full list — favicon, title, domain, and the passage that was actually cited — and tap a web source to open the page.

The pointers come back with the response rather than being written into the text by the model, so a citation always resolves to a real passage in a real document.

Each completed response has action buttons: Copy (to clipboard), Share (via the iOS share sheet), and Speak (reads the response aloud, with a stop toggle). In overlay mode, a "Use as Reply" button lets you send the response directly back to your messaging conversation. You can also select any text natively for Look Up, Translate, or Writing Tools. None of this depends on which model produced the answer.

The agent chat opened from inside a conversation rather than from the tab bar — which is what the Skill button above the keyboard does. It arrives pre-scoped to that one chat and pre-loaded with the Skill you picked, and Use as Reply drops the answer into that chat’s composer.

The difference from a standing agent chat is scope and starting point, not capability: it is the same chat, already pointed at the thread you were reading.

Yes. Multi-agent orchestration lets one agent coordinate with others to complete complex work. Agents can act in parallel with their own isolated context, which helps improve output quality and time to completion. For example, one sub-agent can search the web while another analyses your messages. Each sub-agent has its own conversation history and restricted tool access — no message sending or further delegation. Only one level of delegation is supported: the coordinator can call other agents, but those agents cannot call agents of their own. Sub-agents are capped at 15 turns per task.

The agent loop runs on Anthropic’s servers instead of on your iPhone. Claude keeps the conversation history and directs its own multi-step work there, while your device still executes the tools locally and streams the results back — so a long job survives your phone locking, the app being backgrounded, or you picking it up on another device.

Toggle it per topic from the model picker. It needs Opus or Sonnet and your own Anthropic key: server sessions are organisation-scoped, so they are not offered on AI Credit.

Yes. The agent supports VoiceOver with labels and hints on every interactive element, Dynamic Type for scalable text, native text selection with Look Up, Translate, and Writing Tools, and haptic feedback on key actions. It also adapts fully to dark mode.

Your personal AI assistant for messaging.

Download AI Reply Assistant on iOS and try it free.