Your AI assistant, explained.
You’ve seen what AI Reply Assistant can do. This is how it’s organised — one library of Skills, two places to use them, two brains to run them, and you confirming every action.
One library. Two kinds.
Everything the AI does starts in your Skill library. Every entry there is one of two kinds: a Prompt — always-on instructions that shape how every reply sounds — or a Skill — an on-demand task that carries its own tools and knowledge base. Build either in the same editor, and the library, the editor, and everything in them live on your device.
Always-on instructions added to every reply — tone, style, and approach. A system prompt plus tone sliders, and nothing else to configure.
Examples: Family Circle, Work Team, Dating — nine are built in.
Loaded on demand when relevant — a task the AI runs with its own tools and knowledge base, for work instructions alone can’t do.
Examples: Plan a Meetup, Travel Research, Birthday Check — ten are built in.
Build a Skill, start to finish
The same six screens you'd tap through in the app — your library, the marketplace, a one-tap import, the editor, its tools, and its Knowledge Base. Step through them, or jump to any screen.
Step 1 of 6: Open your Skill library
What a Skill can reach: the tools
A Skill enables its own set of tools — your calendar, your location, the web, and more. Here's the mechanic: the model reads each tool's description and decides when to call it; AI Reply Assistant then executes the call natively on your device through Apple's own frameworks, inside the iOS sandbox and behind iOS permission prompts. No external servers, no MCP plumbing — and every call appears inline as a tappable pill.
Give your Skill access to live information and your own documents.
Web Search · Knowledge Base · URL Fetcher
Tap ↻Direct access to real-time web content, allowing Claude to answer questions with up-to-date information beyond its knowledge cutoff. Includes intelligent result selection, concurrent page fetching, and per-page summarisation with cited sources.
Only the search query leaves your device — your messages are never sent.
Hybrid semantic + keyword search across documents, URLs, and files you attach to a Skill. Results are AI-reranked for relevance.
Documents stay on your device. Nothing is uploaded.
Fetch and cache content from your Skill's configured URLs. Refreshes automatically with rate limiting and cache management.
Fetches the configured URL only — no message data is sent.
Search across your contacts, messages, and conversations — with deep cross-platform visibility.
Contacts · Platform Contacts · Messages · Message Context · Unread Summary · Chat Discovery · Forum Topics · Group Members · Mutual Chats
Tap ↻Look up names, phone numbers, birthdays, relationships, work info, and group members from your device contacts.
Reads your local address book. Nothing leaves your device.
Search contacts on a specific platform by name, username, or phone number — finds people who may not be in your device address book.
Reads your local platform contact cache. Nothing leaves your device.
Search messages with 15+ filters — by sender, platform, date, type, starred, forwarded, media, and more. Falls back to semantic search when exact matches are not found.
Searches your local message database only.
Get surrounding messages around a specific message for full conversation context — useful for understanding replies in thread.
Reads your local message database only.
Get a summary of unread messages across all platforms — see what you missed at a glance.
Reads your local message database only.
Find and filter conversations by platform, unread status, pinned, archived, favourite, participant count, and activity date.
Reads your local conversation list. Nothing leaves your device.
List and browse forum topics within Telegram supergroups — see topic names, message counts, and pinned status.
Reads your local conversation data. Nothing leaves your device.
List participants of a group chat with roles, join dates, and admin status.
Reads your local group membership data. Nothing leaves your device.
Find groups and conversations you share in common with a specific contact.
Reads your local conversation list. Nothing leaves your device.
Check your calendar and to-do lists without leaving the chat.
Calendar · Reminders
Tap ↻Search events by title, date, or location — up to 90 days back and a year ahead. Detects all-day events and shows calendar names.
Reads your local calendar. Nothing leaves your device.
Access your task lists, reminders, and to-dos.
Reads your local reminders. Nothing leaves your device.
Contextual awareness of where you are.
Location
Tap ↻Get your current location, search for nearby places, or get directions. Location is cached for 10 minutes for efficiency.
Uses iOS Core Location. Location data stays on your device.
Browse your files, generate images, create documents, and use pre-built Agent Skills for spreadsheets and presentations.
Files · Image Generation · Document Creation · Agent Skills
Tap ↻Search and read files from locations you grant access to. Reads content from PDF, Word (DOCX), Excel (XLSX), PowerPoint (PPTX), Pages, Keynote, Numbers, RTF, and HTML; notes and data files (TXT, Markdown, reStructuredText, JSON, XML, CSV, TSV, YAML, LOG, PLIST, INI, TOML, ENV, CONF, CFG, .properties); and source code (Swift, JavaScript, TypeScript, JSX/TSX, Python, Java, Kotlin, Ruby, Go, Rust, C, C++, Objective-C, C#, PHP, Shell/Bash/Zsh, SQL, R, Scala, Dart, Lua, Perl, Gradle, Groovy). Also integrates with the Files API, which lets you upload and manage files without re-uploading content with each request — upload once and reference across multiple conversations.
Local file reads stay on-device. Files API uploads go directly to Anthropic.
Create fun, original images on-device using Apple Intelligence Image Playground. Choose from three styles: Animation, Illustration, and Sketch. All images are created on device.
Generated entirely on-device by Apple Intelligence.
Generate or edit documents in TXT, Markdown, CSV, JSON, HTML, and Microsoft Word (.docx) formats. Claude writes the content with full formatting — headings, bold, italic, tables, lists, code blocks, and hyperlinks — then registers it as a sendable attachment.
Documents are created and stored locally on your device. Nothing is uploaded.
Modular capabilities that extend Claude's functionality. Pre-built Agent Skills let Claude run code in a sandboxed container to generate professional documents: PowerPoint presentations (.pptx), Excel spreadsheets (.xlsx), PDF reports (.pdf), and Word documents (.docx) — downloaded automatically and presented as previewable, shareable file cards.
Files are generated in Anthropic's code execution container using your API key. The finished document is downloaded straight to your device.
Persistent memory and parallel research for smarter conversations.
Memory · Sub-agent
Tap ↻Store and retrieve information across conversations, allowing Claude to build knowledge over time. Save names, preferences, or important details — Claude automatically checks its memory before starting tasks and references it to handle similar tasks more effectively.
Memories are stored locally on your device.
Spawn parallel research tasks — agents can act in parallel with their own isolated context, helping improve output quality and time to completion.
Inherits the privacy model of the active AI provider.
Working with tools in a chat
- Enable the tools a Skill can use in the Skill editor — the most relevant ones are chosen automatically per message. Knowledge Base and URL Fetcher auto-enable from your configuration.
- Tools that need iOS permissions (Calendar, Contacts, Location, Photos, Reminders) prompt for access the first time they are used.
- Every tool invocation is shown inline in your chat as a tappable pill — tap to see parameters, results, and citations.
- If a tool fails, the AI continues with a graceful fallback and shows the error inline.
What a Skill can know: the Knowledge Base
Attach documents and URLs to a Skill — up to 10 documents and 20 URLs — and it references them when it works: brand guidelines, a product sheet, your team wiki. Everything is processed and stored on your device, searched with hybrid semantic + keyword retrieval, and cited in replies. Adding even one source switches the knowledge-base tool on for you.
The built-in library
Nineteen entries ship ready to use — nine Prompts for everyday voice and ten Skills wired to tools. Expand either kind to browse them.
A voice for each relationship in your life — use them as they are, or open any one to see how its system prompt and sliders produce its personality.
Friendly all-rounder that helps craft engaging messages that build meaningful connections.
Strengthens bonds and keeps everyone connected.
Real, warm, no filter — for the people you can be yourself with.
Professional, but human.
Polite, precise, on-task.
Coordinates time, place, and people without the chaos.
Practical and excited — perfect for trip logistics.
Energetic and inclusive across larger circles.
Authentic, warm, and present — connection over cleverness.
Each pairs instructions with the right tools, so one plain-language ask researches, plans, and drafts in a single turn — no assembly required.
Checks your calendar and proposes available times for scheduling.
Plans outings by checking your schedule and finding nearby places.
Finds nearby restaurants, cafes, and venues based on your current location.
Checks upcoming birthdays from your contacts and helps craft greeting messages.
Checks your reminders and tasks to give a clear status update.
Researches travel destinations with web search for trip planning.
Checks your calendar for conflicts and crafts an appropriate RSVP reply.
Looks up contacts and helps draft personalised outreach messages.
Checks upcoming events and related tasks to give a preparation status update.
Or borrow one — over a million of them
The built-in Skills Marketplace imports community Skills with one tap. Importing copies only the Skill definition to your device — the marketplace is for browsing and downloading, and never sees a message.
Sales, marketing, finance, and project management.
Writing, design, documents, and content creation.
Knowledge bases, technical docs, and educational material.
LLM prompts, machine learning, and data analysis.
Wellness, writing, philosophy, arts, and culinary.
Academic, scientific computing, and lab tools.
Browse, import, make it yours
Open the Skills Marketplace from your Skill library and explore by category or keyword.
Read the Skill's description, instructions, and author before importing.
Tap Import — the Skill appears in your library instantly, ready to use.
Adjust tone, creativity, formality, or edit the system prompt to make it your own.
One tap, in the thread.
Above the keyboard in every conversation sit three actions — Smart Reply, Skill, and Analyze. The Skill button runs your library right where you're typing: pick any Prompt or Skill, add an instruction, set filters, and generate. It reads only the conversation you're in — never your inbox — and the draft lands in the composer for you to edit before it goes anywhere.


By default this all runs on Apple Intelligence, on your iPhone. Flip the Skill feature to Cloud AI in Settings and the same button drafts with Claude instead — more on choosing below.
A full conversation with Claude.
For work bigger than one reply, the app has a dedicated agent chat — a standing conversation with Claude, running on your own Anthropic key. Everything is set per conversation, right in the chat: the model and its thinking effort, which tools Claude may use, the active Skill, and — most importantly — the scope: you pick which of your conversations the agent may read, and everything outside that selection is invisible to it. Locked chats never appear at all.
Know your way around
Everything in the agent chat is one tap away. Hover over a feature — or tap it — to see where it lives and what it opens.
Two brains. You pick.
Everything above runs on one of two engines. Everyday drafting uses the Apple foundation model at the core of Apple Intelligence — on your iPhone, built in, instant, even offline. When a task needs more — long threads, research, documents, multi-step work — Claude takes over on your own Anthropic key: the agent chat always runs on it, and the in-chat features can switch to it per feature in Settings.
The full comparison
| Capability | Apple Intelligence | Claude |
|---|---|---|
| Where it runs | On-device processing — powered by the foundation model at the core of Apple Intelligence. No network required, works offline | Anthropic cloud — you bring the API key, direct to Anthropic |
| What you pay | Built in — no API key needed. The Free plan includes 10 on-device AI actions a day (shared across Smart Reply, Analyze, Summary, Skills, and AI Writer); Pro makes them unlimited | Pro plan, plus your own API key — you pay Anthropic directly per use, no middleman, no markup |
| Models | Apple Foundation Models (~3B parameters) | Haiku 4.5 (fastest, most cost-efficient), Sonnet 4.6 (best combination of speed and intelligence), Opus 4.8 (most capable, complex reasoning and agentic coding) |
| Context window | 4,096 tokens at a time (input + output combined, ~3,000 words) — best for shorter threads; longer ones are condensed automatically | Up to 1,000,000 tokens (Opus/Sonnet) or 200,000 tokens (Haiku) — handles entire conversation histories |
| Input types | Text only — images, PDFs, and documents are handled by Claude | Multimodal — understand and analyse a wide range of visual formats, including photos, charts, graphs, and technical diagrams |
| Extended thinking | Not available — the on-device model is tuned for direct, single-pass responses | Enhanced reasoning capabilities for complex tasks — Claude's step-by-step thought process is visible before it delivers its final answer, with adjustable effort (low/medium/high) |
| Reasoning and logic | Best for simple, focused tasks like everyday drafting — rather than heavy logic or math | Strong complex reasoning, multi-step problem solving, code generation, and mathematical tasks |
| Tool use | 20 on-device tools — contacts, calendar, reminders, location, nearby places, web search, and your messages, all running privately on your iPhone | 22 tools — Claude decides when to call a tool based on your request and the tool's description, then AI Reply Assistant executes it on your device. Richer native web search, file creation (PPTX, XLSX, PDF, DOCX), calendar, contacts, image generation, and more |
| Agentic behaviour | Single-turn — one prompt, one response | Claude dynamically directs its own processes and tool usage, maintaining control over how it accomplishes tasks — planning, calling tools, observing results, and refining across multiple iterations |
| Document creation | Not available | Pre-built Agent Skills let Claude run code in a sandboxed container to create PowerPoint presentations, Excel spreadsheets with charts, Word documents, and PDF reports — downloaded and delivered as shareable attachments |
| Message drafting | Generates up to 3 drafts in different tones — you pick the one to use before anything is sent | Human-in-the-loop — Claude drafts messages as cards, you review, choose recipients, and confirm before anything is sent |
| Context management | When a thread outgrows the context window, the transcript is automatically condensed and the conversation continues — no error, nothing lost | Compaction extends effective context length by automatically summarising older context, keeping the active context focused and performant. Prompt caching reduces processing time and costs for repetitive tasks. On-demand tool loading and real-time token tracking |
| Languages | 9 languages (English, French, German, Italian, Portuguese, Spanish, Japanese, Korean, Chinese) | Broad multilingual support across dozens of languages |
| Factual knowledge | A compact on-device model — tuned for everyday drafting rather than broad factual recall or current events | Broader world knowledge with more current training data — stronger on factual and domain-specific questions |
| Privacy | Nothing leaves your phone — maximum privacy | You control when Claude is used. Drafts go directly to Anthropic. |
| Best for | Everyday replies, quick suggestions, full privacy, offline use | Even more expertise — longer drafts, nuanced tone, complex reasoning, multi-step tasks, document creation, and research |
It's an everyday reply, you want it instantly, you're offline, or nothing should leave your phone. This is the default — most messages never need more.
The thread is long, the task has steps, you need research with sources, a document produced, or a nuanced tone for a message that matters.
Switching is one setting away
Settings → Chat → AI Features. Smart Reply and Skill each pick On-Device or Cloud AI independently, and a model row chooses the Claude model used for cloud drafts. Cloud AI only becomes selectable once you've connected your own Anthropic key — and any feature switches back to On-Device anytime.
What it costs
Apple Intelligence runs on every plan with no per-message fees — the Free plan includes 10 on-device AI actions a day (Smart Reply, Analyze, Summary, Skills, and AI Writer), and Pro makes them unlimited. Claude is a Pro feature: Pro is $49.99/year (billed annually, save ~45%) or $7.49 monthly, each with a 7-day free trial. You then bring your own Anthropic API key and pay Anthropic's standard usage rates directly — no markup, no per-message fees from us, and you can cap spending in your Anthropic console. A typical drafted reply uses only a few thousand tokens, and you choose the model per task — Haiku 4.5 for fast, inexpensive drafting up to Opus 4.8 for the hardest work.
It proposes. You approve.
Wherever the AI runs — the Skill button or the agent chat — it can't touch your conversations or your device directly. When it wants to act, it must produce an action card — a structured preview of exactly what would be sent or created. Messaging cards open a recipient picker, OS cards execute with a single tap, Agent Skills cards carry the documents Claude creates — and nothing moves until you say so.

WhatsAppMessaging cards9 types
Pick recipients from WhatsApp, Telegram, Instagram, Messenger, or iMessage and confirm before sending.
Text message to a chat
Photo with caption
Video with caption
Voice or audio file
Document or PDF
Share a contact card
Share a calendar event
Share a map pin
Interactive poll
OS cards6 types
Create items on your device with a single tap. Status updates in real time.
Add to Calendar
Add to Reminders
Add to Contacts
Save to Notes
Open a URL
Save generated image
Agent Skills cards1 type
Documents Claude generates with its pre-built Agent Skills — preview, share, or save.
PowerPoint, Excel, PDF, or Word document
Working with action cards
- Claude generates a card based on your request — you see a full preview before any action is taken.
- Messaging cards open a multi-select recipient picker. Send to one chat or fan out to many at once.
- Image, video, audio, and file cards reference attachments you provided — Claude never fabricates media.
- Every card shows real-time status: idle, sending, sent (with recipient count), or failed with a retry option.
- Cards persist across sessions so you can review what was sent later.
Where everything runs.
The whole page, audited in one view — each surface above, what runs where, and what it can see. Nothing in this table routes through Flowbie.
| Surface | Where it runs | What it can see |
|---|---|---|
| Skills & Knowledge Base | On device | Your prompts, sliders, and attached documents |
| Skill button in a chat | On device · Claude if you enable it | Only the conversation you’re in |
| Agent chat | Claude, on your own key | Only the chats you scope — locked chats never |
| Tools | On device, Apple frameworks | Only what each tool needs, behind iOS permissions |
| Apple Intelligence | On device | The scoped context; nothing leaves your iPhone |
| Claude | Anthropic cloud, your own key | Only what you send — direct to Anthropic |
| Action cards | On device | Nothing is sent or created until you confirm |
| Skills Marketplace | Browsing only | Skill definitions — never a message |
AI assistant — questions, answered.
Skills and AI
A Skill is a modular capability that packages instructions, context, and best practices to transform AI Reply Assistant into a specialist for any relationship or domain. Each Skill has its own voice, tone, system prompt, and Knowledge Base — giving you domain-specific expertise that loads on demand. AI Reply Assistant ships with nine built-in Skills — General Assistant, Family Circle, Best Friend, Work Team, Concierge, Event Planner, Travel Buddy, Social Group, and Dating. You can also create your own or import community Skills from the built-in Skills Marketplace.
Your library holds two kinds of entry, each badged with its type. A Prompt is always-on instructions added to every reply — a system prompt plus tone sliders, with nothing else to configure. A Skill goes further: it loads on demand when relevant and carries its own tools and Knowledge Base, so the AI can actually do things — check your calendar, search the web, look up contacts — before it writes. Both are built in the same editor; a Skill simply adds the Tools and Knowledge Base sections on top.
| Prompt | Skill | |
|---|---|---|
| What it is | Always-on voice and tone | On-demand task |
| You configure | System prompt + tone sliders | Adds Tools + Knowledge Base |
| Fetches live data | No | Yes — calendar, web, contacts |
| Best for | A consistent voice (Family, Work Team) | Replies needing real data or steps |
| Built-in count | Nine | Ten |
Tap the + button in your Skill library, then:
- Choose whether you’re making a Prompt or a Skill
- Give it a name and write a system prompt describing the personality you want
- Adjust the tone sliders
- For a Skill, enable any tools and attach a Knowledge Base
- Optionally pick a category and choose an AI model
Your custom entry is ready to use immediately.
Every Skill has five tone sliders — creativity, formality, directness, empathy, and length — plus sampling controls (Consistent / Balanced / Varied). You can edit the system prompt directly, switch between Apple Intelligence (built in) and Claude (even more expertise) per Skill, attach a Knowledge Base with your own documents, and enable Extended Thinking for enhanced reasoning with Claude.
The Skills Marketplace is a built-in browser powered by SkillsMP that gives you access to over a million community-created Skills. Browse by category — Business, Content, Documentation, Lifestyle, Research, and more — preview any Skill, and import it into your library with one tap. You can then customise it to fit your needs.
A Knowledge Base lets you attach documents, URLs, and custom instructions to a Skill. When drafting replies, the AI references this material to give more informed, contextual responses. For example, a "Brand Voice" Skill could include your company's tone guidelines, or a "Legal" Skill could reference compliance documents. You can attach up to 10 documents and 20 URLs per Skill.
A wide range of formats:
- PDFs
- Microsoft Office — Word (DOCX), Excel (XLSX), PowerPoint (PPTX)
- Apple iWork — Pages, Keynote, Numbers
- Text & markup — rich text (RTF), Markdown, HTML, plain text
- Data & code — JSON, XML, CSV, Swift, JavaScript, Python, and more
- Images — AI Reply Assistant extracts text from those too
Each file can be up to 50 MB, with a total limit of 100 MB per Skill.
RAG stands for Retrieval-Augmented Generation. Instead of relying solely on the AI model's built-in knowledge, RAG lets AI Reply Assistant search your Knowledge Base documents for relevant information and include it in the prompt before generating a reply. This means the AI's responses are grounded in your actual documents — your company guidelines, reference material, or personal notes — rather than generic training data.
It runs a small retrieval pipeline entirely on-device. When you add a document to a Knowledge Base:
- It’s broken into small chunks
- Each chunk gets a mathematical representation (an embedding)
- The embeddings are stored on your device
Then, when you ask for a reply:
- Your conversation is converted into a query
- The most relevant chunks are found using similarity search
- Those chunks are fed to the AI alongside the conversation
- The AI drafts a reply informed by both the thread and your documents — all on your device
RAG works with both. On iOS 26, Apple Intelligence powers the document processing pipeline — chunking, embedding, query optimisation, and retrieval all run on-device using Apple Foundation Models. The retrieved context is then passed to whichever AI provider your Skill uses (Apple Intelligence or Claude) for the final reply.
Yes. All document processing — text extraction, chunking, embedding, and similarity search — happens entirely on your device. Your documents are never uploaded to any server. The only time Knowledge Base content leaves your phone is if your Skill uses Claude, in which case the relevant retrieved chunks (not the full documents) are sent to Anthropic as part of the prompt.
Each slider works together with the system prompt to shape the Skill's overall voice:
- Creativity — how inventive vs predictable the language is (higher means more varied word choices)
- Formality — casual vs professional tone
- Directness — blunt and to-the-point vs softer and more diplomatic
- Empathy — how much emotional awareness and warmth the AI shows
- Length — how brief or detailed the reply is
Sampling controls how the AI picks its next word. You can set this per Skill:
- Consistent (greedy) — always picks the most likely word; great for factual, predictable replies
- Balanced (top-K) — samples from a curated set of likely words; good for natural-sounding conversation
- Varied (top-P) — samples from a wider probability range; best for creative or playful replies
Extended Thinking gives Claude enhanced reasoning for complex tasks. It is not a different model — it lets the very same model give itself more time and expend more effort to reach an answer.
When enabled, Claude reasons step-by-step through the conversation context, your Knowledge Base material, and the right tone before writing its reply. AI Reply Assistant makes this thought process visible in raw form — being able to observe how Claude thinks makes its answers easier to understand and check.
You can set a thinking budget to control precisely how long Claude spends on a problem; larger budgets can improve response quality by enabling more thorough analysis.
Yes. Apple Intelligence is built in through on-device processing. With Claude, you get even more expertise — choose between three models: Haiku 4.5 is the fastest, most cost-efficient model with near-frontier intelligence. Sonnet 4.6 is a hybrid reasoning model with the best combination of speed and intelligence — it can produce near-instant responses or extended, step-by-step thinking. Opus 4.8 is the most capable generally available model for complex reasoning and agentic coding. You can also enable Extended Thinking on any model for enhanced reasoning capabilities.
Both are built in, and you choose per chat or per Skill. Apple Intelligence runs entirely on-device — fast, private, and free — and is ideal for focused everyday replies. Claude brings far more expertise for complex requests; you control when it is used by bringing your own Anthropic API key.
| Apple Intelligence | Claude | |
|---|---|---|
| Runs | On-device | Cloud (your API key) |
| Context window | ~4,096 tokens (~3,000 words) | Up to 1M tokens |
| Understands | Text only | Text, images, charts, PDFs |
| Tools | Core set | 22 tools with agentic iteration |
| Extended thinking | No | Yes |
| Document creation | — | PPTX, XLSX, PDF, DOCX |
| Best for | Fast, private everyday tasks | Complex, multi-step requests |
Three different tools:
- Smart Reply — Apple’s native quick suggestions
- Skill — a full reply in the voice of your active Skill, the AI persona you’ve set for that chat
- Analyze — summarises the day’s thread so you can catch up in seconds
Use a built-in Skill when the relationship fits — family, work, dating, etc. Create a custom Skill when you have a specific context that repeats: a particular client, a community with its own jargon, a professional domain like legal or medical, or any situation where you want the AI to follow particular guidelines every time.
Claude Haiku 4.5 by default — the fastest, most cost-efficient model with near-frontier intelligence. You can switch to Sonnet 4.6, a hybrid reasoning model with the best combination of speed and intelligence that can compress multi-day tasks into hours. Or choose Opus 4.8, the most capable generally available model for complex reasoning and agentic coding — it handles complex, long-running tasks with rigor and consistency. All three support Extended Thinking. Switch in Chat Settings at any time.
Skills Marketplace
It takes four taps:
- Open your Skill library and tap the Marketplace button
- Browse or search for a Skill
- Tap it to preview
- Tap Import — it’s added to your library instantly
From there you can customise the tone, system prompt, and other settings before using it.
The marketplace hosts over a million Skills across categories such as:
- Business — sales, marketing, finance
- Content & Media — writing, design
- Documentation — knowledge bases, technical docs
- Data & AI
- Lifestyle — wellness, writing, arts
- Research — academic, scientific
Whether you need a brand-voice enforcer, a customer-service persona, or a creative writing coach, there’s likely a community Skill for it.
Yes. The marketplace is only used for browsing and downloading the Skill definition (a text-based instruction set). Once imported, the Skill runs locally in AI Reply Assistant just like any built-in Skill. No message data is ever shared with the marketplace.
Absolutely. An imported Skill is fully editable — you can change its name, rewrite the system prompt, adjust tone sliders, switch its AI model, or attach your own Knowledge Base. Think of it as a starting point you can make your own.
Claude Agent
The Claude Agent is a built-in agentic AI assistant that dynamically directs its own processes and tool usage, maintaining control over how it accomplishes tasks. It uses tool use — Claude decides when to call a tool based on your request and the tool's description, then AI Reply Assistant executes it natively on your device. Search your messages, check your calendar, create reminders, look up contacts, browse the web, generate images, and more. Tool access is one of the highest-leverage capabilities you can give an agent — it grounds Claude's responses in real-time context, not just the chat thread.
You are always in control. When Claude drafts a message, it appears as a card in your chat. You choose which conversations to send it to, review the content, and confirm with a final prompt before anything is sent. This human-in-the-loop design ensures every outgoing message reflects your intent.
Yes. Before using message-related tools, you select which conversations Claude can access using the scope picker. Claude can only search and read messages within those specific conversations — everything else remains invisible to the AI. You can change or clear this selection at any time.
Tool use lets Claude call functions that AI Reply Assistant defines — Claude decides when to call a tool based on your request, then your device executes it natively. The agent has 21 tools across categories you can toggle on or off:
- Calendar — view and create events
- Reminders — view and create tasks
- Contacts — search your address book
- Web Search — browse the internet
- Image Generation — create images on-device via Apple Intelligence
- Message Search — find messages in scoped conversations
- Files — cloud-based file management via the Files API
- Document Creation — Word, Markdown, CSV, JSON, HTML
- Agent Skills — server-side generation of PowerPoint, Excel, PDF, and Word
- Location, Memory, and more
Tool access is one of the highest-leverage capabilities you can give an agent — it grounds Claude's responses in real-time context, not just the chat thread.
The context window is all the text a model can reference when generating a response, including the response itself — a working memory for the model. A larger window handles longer, more complex prompts, but more context isn't automatically better, so AI Reply Assistant actively manages it for you:
- Compaction — automatically summarises older context as you near the limit, keeping the active context focused and performant
- Prompt caching — resumes from cached prefixes, significantly cutting processing time and cost on repetitive tasks
- On-demand tool loading — tool definitions load only when needed rather than all at once, preserving working memory
- Attachment deduplication — avoids re-sending the same file
When a conversation approaches its context limit, AI Reply Assistant prompts you to compact or start a fresh topic.
MCP (Model Context Protocol) is an open standard for connecting AI assistants to the systems where data lives — content repositories, business tools, and development environments — replacing fragmented integrations with a single protocol. Think of MCP like a USB-C port for AI: just as USB-C standardises how you connect electronic devices, MCP standardises how AI applications connect to external systems.
MCP is designed primarily for desktop and server environments. AI Reply Assistant takes a complementary approach: instead of connecting to external MCP servers, the agent runs natively on iOS and executes tools directly through Apple's own frameworks (EventKit, Contacts, and others) inside the iOS app sandbox.
The result: your data benefits from iOS-level permissions, every action requires your explicit approval, and tools run with the same security guarantees as any native app.
Yes. Every tool call is shown as a step in your conversation — you can see what tool was used, what parameters were passed, whether it succeeded or failed, and what it returned. Extended Thinking is also visible when enabled, so you can follow Claude's reasoning before it responds.
AI Reply Assistant tracks your context usage and shows a suggestion when the conversation approaches its limit. You can compact the conversation to keep only recent context, or start a new topic. Each topic has its own isolated conversation history, so switching topics gives you a fresh context window without losing previous work.
Three Claude models are available — switch mid-conversation and the next response uses your selection. All three support Extended Thinking.
| Model | Strength | Context | Best for |
|---|---|---|---|
| Haiku 4.5 | Fastest, most cost-efficient | 200K | Quick drafting |
| Sonnet 4.6 | Best balance of speed + intelligence | 1M | Everyday and complex tasks |
| Opus 4.8 | Most capable for deep reasoning | 1M | Complex, long-running work |
Extended Thinking gives Claude enhanced reasoning for complex tasks, with full transparency into its step-by-step thought process before the final answer. It is not a different model — it lets the very same model give itself more time and expend more effort to reach an answer.
You set a thinking budget to control how long Claude spends on a problem — low, medium, or high; larger budgets can improve quality through more thorough analysis. Opus and Sonnet use adaptive thinking that automatically adjusts effort to the task; Haiku uses a fixed budget.
The thought process is visible in raw form — a collapsible "Thinking" block with an elapsed timer shows the reasoning in real time, making Claude's answers easier to understand and check.
Yes. Pick a style before sending your message — it shapes how Claude formats its reply:
- Normal — balanced
- Learning — patient, step-by-step explanations
- Concise — short bullet points, no elaboration
- Explanatory — detailed and educational
- Formal — professional and structured
Yes. Claude's vision capabilities allow it to understand and analyse images, opening up exciting possibilities for multimodal interaction. Claude can process a wide range of visual formats, including photos, charts, graphs, and technical diagrams. You can attach up to 5 images (from your photo library or camera) and up to 5 PDFs per message. Multiple images can be included in a single request, which Claude will analyse jointly when formulating its response. For PDFs, you can ask Claude about any text, pictures, charts, and tables in the documents you provide — from analysing financial reports to extracting key information from legal documents. The Files API lets you upload files once and reference them across multiple conversations without re-uploading.
Claude is capable of providing detailed citations when answering questions about documents, helping you track and verify information sources in responses. Citations are guaranteed to contain valid pointers to the provided documents — significantly more reliable and more likely to cite the most relevant quotes compared to purely prompt-based approaches. When Claude uses web search or references documents, inline citation markers (superscript numbers) appear in the response text. Tap the "Sources" pill to see a full list with favicons, titles, domains, and the cited text from each source. Web sources are tappable to open the original page.
Action cards are structured outputs Claude creates when it wants to take an action on your behalf. Each card appears visually in the chat with a preview of what will be sent or created. There are two categories:
Messaging cards — send content to WhatsApp, Telegram, Instagram, Messenger, or iMessage. You pick recipients and confirm before anything is sent:
- Draft Message — text message
- Image Message — photo with caption
- Video Message — video with caption
- Audio Message — voice or audio file
- File Message — document or PDF
- Contact Message — share a contact card (vCard)
- Event Message — share a calendar event (ICS)
- Location Message — share a map pin
- Poll Message — create an interactive poll
OS cards — create items on your device with a single tap:
- Event — add to Calendar
- Reminder — add to Reminders
- Contact — add to Contacts
- Note — save to Notes
- Link — open a URL
- Image — save a generated image to Photos
Agent Skills cards — documents generated server-side by Claude:
- Skill File — preview, share, or save PowerPoint, Excel, PDF, and Word documents created by pre-built Agent Skills
Every card shows real-time status: working, done, sent (with recipient count), or failed with a retry option.
AI Reply Assistant tracks delivery per recipient. If some sends fail, you see a summary — "Sent to 2 of 3 chats" with the failure reason for the first error. The retry flow pre-selects only the failed recipients so you can resend without duplicating to those that succeeded.
Yes. A stop button appears while Claude is generating a response. Tap it to cancel immediately. You can also regenerate the last response if you want Claude to try again with a different approach.
Each completed response has action buttons: Copy (to clipboard), Share (via the iOS share sheet), and Speak (reads the response aloud, with a stop toggle). In overlay mode, a "Use as Reply" button lets you send the response directly back to your messaging conversation. You can also select any text natively for Look Up, Translate, or Writing Tools.
Topics let you organise separate conversations within the agent. Each topic has its own message history, context window, and persona. You can create named topics, pin important ones, close finished ones, reorder them, and customise their icons. Switching topics gives you a fresh context without losing previous work.
Overlay mode launches the Claude Agent from within a messaging conversation. The agent is pre-scoped to that chat and pre-loaded with the active Skill. When Claude drafts a response you like, tap "Use as Reply" to send it directly back to the conversation you came from. It is a quick way to get AI help without leaving your thread.
Yes. Multi-agent orchestration lets one agent coordinate with others to complete complex work. Agents can act in parallel with their own isolated context, which helps improve output quality and time to completion. For example, one sub-agent can search the web while another analyses your messages. Each sub-agent has its own conversation history and restricted tool access — no message sending or further delegation. Only one level of delegation is supported: the coordinator can call other agents, but those agents cannot call agents of their own. Sub-agents are capped at 15 turns per task.
Managed Agents provides the harness and infrastructure for running Claude as an autonomous agent. Instead of building and running the agent loop on your device, you get a fully managed environment where Claude dynamically directs its own processes and tool usage — orchestrating tool calls, maintaining persistent conversation history, and handling multi-step workflows server-side while your device executes tools locally and streams results back. Managed sessions, persistent history, and frontier models mean you get the agent, not the infrastructure. You can toggle this per topic from the model picker — available for Opus and Sonnet.
Yes. The agent supports VoiceOver with labels and hints on every interactive element, Dynamic Type for scalable text, native text selection with Look Up, Translate, and Writing Tools, and haptic feedback on key actions. It also adapts fully to dark mode.
AI Tools
AI Tools give your Skill real-time access to information on your device and the web. Instead of relying only on what is in the chat thread, a Skill with tools enabled can search the web, check your calendar, look up contacts, read files, generate images, and more — all while drafting a reply.
A Skill can enable as many of its selectable tools as you like — there is no fixed limit. When several are enabled, AI Reply Assistant uses embedding-based selection to pick the most relevant tools for each message automatically, so the model stays focused. Some tools — like Knowledge Base and URL Fetcher — auto-enable based on your Skill's configuration.
AI Reply Assistant offers 21 tools across six categories:
- Research & Knowledge — Web Search, Knowledge Base, URL Fetcher
- People & Messages — Contacts, Platform Contacts, Messages, Message Context, Unread Summary, Chat Discovery, Forum Topics, Group Members, Mutual Chats
- Schedule & Tasks — Calendar, Reminders
- Location — Location
- Files & Creation — Files, Image Generation, Document Creation, Agent Skills
- Agent Intelligence — Memory, Sub-agent
The web search tool gives Claude direct access to real-time web content, so it can answer questions with up-to-date information beyond its knowledge cutoff. It uses a four-stage pipeline:
- Search for results
- Intelligently select the most relevant ones
- Fetch those pages concurrently
- Summarise each page within a token budget
The response includes citations for the sources used. Only the search query is sent — your messages never leave your device.
When you attach documents (PDF, Office, iWork, text, and more) or URLs to a Skill, the Knowledge Base tool uses a hybrid search combining 80% semantic similarity with 20% keyword matching. Results are then reranked by an AI model for relevance. The tool auto-enables when you add documents — no manual toggle needed.
Only if you grant Calendar permission and enable the Calendar tool on a Skill. It can search events by title, date, or location — up to 90 days back and a year ahead. It reads events but never creates or modifies them. All calendar data stays on your device.
Yes — with Image Playground, you can create fun, original images entirely on-device. Choose from three styles: Animation, Illustration, and Sketch. All images are created on device, giving you the freedom to experiment with as many images as you want. The tool supports multiple concepts per image.
Agent Skills are modular capabilities that extend Claude's functionality. Each packages instructions, metadata, and optional resources Claude uses automatically when relevant, running code in a secure, sandboxed environment to analyse data, create visualisations, perform calculations, and generate files. Pre-built Agent Skills ready to use:
- PowerPoint (.pptx) — create presentations, edit slides, analyse content
- Excel (.xlsx) — build spreadsheets, analyse data, generate reports with charts
- Word (.docx) — create and edit documents, format text
- PDF (.pdf) — generate formatted PDF documents and reports
Files download automatically and appear as previewable, shareable file cards — open with QuickLook, share via the iOS share sheet, or save to Files. Agent Skills use your Anthropic API key directly.
Document Creation is a local tool — Claude writes content and AI Reply Assistant creates the file on your device. It supports TXT, Markdown, CSV, JSON, HTML, and Word (.docx) with full formatting. Agent Skills, on the other hand, use Claude's server-side code execution to generate more complex documents like PowerPoint presentations and Excel spreadsheets with charts. Use Document Creation for quick text-based files that stay on-device; use Agent Skills when you need richer formats like .pptx or .xlsx.
The Memory tool enables Claude to store and retrieve information across conversations, allowing it to build knowledge over time without keeping everything in the context window. Claude automatically checks its memory before starting tasks and can create, read, update, and delete facts — like "my wife's name is Sarah" or "I prefer window seats". Memories are stored locally on your device, eliminating the need to re-explain context.
The Messages tool searches your local message database with 15+ filters — sender, platform, date range, message type, starred, forwarded, media type, and more. When exact matches are not found, it automatically falls back to semantic search to find relevant results by meaning rather than keywords.
Most tools run entirely on-device — Calendar, Contacts, Reminders, Location, Files, Photos, Messages, Memory, and Image Generation never send your data anywhere. Web Search sends only the search query (not your messages) to DuckDuckGo. URL Fetcher fetches only the configured URL. Tool permissions are managed per Skill.
Knowledge Base and URL Fetcher turn on automatically when you add documents or URLs to a Skill, separately from the tools you select by hand. This ensures your Skill always has access to the context you have given it, without needing manual configuration.
Yes. Every tool invocation appears inline in your chat as a tappable pill showing the tool name and status. Tap it to see the full details — what parameters were used, whether it succeeded or failed, and what results it returned, including any citations or sources.
Tools are designed with graceful fallbacks. If a web page cannot be fetched, cached content is used. If a tool encounters an error, the AI continues without that data and notes the limitation. Errors are shown inline so you always know what happened.
Yes. The tool system is provider-agnostic — the same tools are available regardless of whether your Skill uses Apple Intelligence or Claude. This means you get calendar access, web search, and everything else with either AI provider.
Yes. You choose which tools each Skill can use in the Skill editor — toggle them on or off individually. When you enable several, the most relevant ones are selected per message automatically. Knowledge Base and URL Fetcher auto-enable based on your configuration but can be overridden. Tools that require iOS permissions (Calendar, Contacts, Location, Photos, Reminders) will prompt for access the first time they are used.
No. Each tool operates independently within the context of the active conversation. The Calendar tool cannot see your Contacts data, and the Web Search tool cannot read your Files. The AI model may use the results from one tool to inform a query to another — for example, checking your calendar after looking up a contact — but the tools themselves do not exchange data directly.
Message intelligence
Yes. AI Reply Assistant can classify every incoming message with a sentiment (positive, negative, neutral, mixed, or question) and an intent (informational, request, social, urgent, or transactional). It also scores urgency and formality on a scale, assigns a priority level (1–3), and extracts topics automatically. Tap any message to see a compact badge, or expand it for the full breakdown.
Yes. When someone mentions a date, time, or event in a message — "dinner tomorrow at 7" or "team standup Monday 10 am" — AI Reply Assistant extracts the event details (title, location, date range, notes) and shows an inline card. Tap "Add to Calendar" to save it in one step. Past events are highlighted so you can tell at a glance which ones have already happened.
Yes. AI Reply Assistant scans messages for action items and shows them as inline cards with status tracking (pending, in progress, completed, delegated), owner assignment, deadline indicators, and blocker flags. Overdue items are highlighted. Tap "Add to Reminders" to send any action item straight to Apple Reminders with the deadline pre-filled.
Day Summary is an AI-generated digest of the day’s conversation. It condenses the thread into key points and surfaces any events, reminders, and action items mentioned throughout the day — all in a collapsible card you can drag to resize. Extracted events and reminders appear as tappable inline cards so you can add them to your calendar or reminders without scrolling back through the chat.
Yes. When someone shares a phone number, email, company name, or job title in a message, AI Reply Assistant extracts the contact details and shows an inline card. Tap "Add to Contacts" to save the person to your address book with all the extracted fields pre-filled.
Yes. AI Reply Assistant detects foreign-language messages automatically and shows a compact translate icon on the bubble. Tap it to open the iOS translation sheet, then tap "Replace with Translation" to swap the original text for the translated version directly inside the bubble. The original is shown below in a smaller font, and a "Translated · Show Original" indicator lets you switch back at any time. Translations are persisted to your device, so they stay visible when you leave and return to the chat.
Translation works on any bubble with text content:
- Text and link messages
- Media captions — photos, videos, GIFs, documents
- Transcripts — voice messages, audio files, videos, and video notes
- Interactive messages — button, list, template, and interactive
Translation is powered by Apple’s on-device Translation framework — all processing happens locally on your device, and no text is sent to external servers.
Over 20 languages are supported, including Arabic, Bengali, Chinese (Simplified and Traditional), Czech, Danish, Dutch, English (UK and US), Finnish, French, German, Greek, Hebrew, Hindi, Hungarian, Indonesian, Italian, Japanese, Korean, Malay, Norwegian (Bokmål), Persian, Polish, Portuguese (Brazil), Romanian, Russian, Slovak, Spanish, Swedish, Tagalog (Filipino), Thai, Turkish, Ukrainian, and Vietnamese. Availability depends on your iOS version — newer releases add more.
You can download additional language packs in Settings → Apps → Translate → Downloaded Languages for offline use.
Yes — four media types, all on-device: voice messages, videos, video notes (Telegram’s circular videos), and audio files. All processing happens locally; nothing is sent to external servers.
Tap the transcription button on any voice, audio, or video bubble to start. The transcript streams in as it’s recognised and is cached to your device, so it loads instantly the next time you open the chat.
You can set a per-chat transcript language from the long-press context menu. If the transcript is in a foreign language, a translate button appears inline so you can view the translation without leaving the bubble — and those translations are persisted too.
Requirements
On iPhone and iPad — and on Apple Silicon Macs as the iPad app. No Android, no web.
iPhone 15 Pro and later, or any iPad or Mac with M1 or A17 Pro and later.
Pro plan plus your own Anthropic API key — pay Anthropic directly for usage, no markup.
Your personal AI assistant for messaging.
Download AI Reply Assistant on iOS and try it free.
Apple Intelligence
Claude