Shopify AI Chatbot App vs Building Your Own RAG Chatbot
An app is faster to start, and a custom build gives you control over answers, data and cost. Here is how to decide, what a build includes, and what to ask a developer if you hire.

Short answer: pick an app if you want something working quickly for common questions and your catalogue is simple. Build your own RAG chatbot, or hire a developer to build it, if answer quality, data ownership or cost at volume matter more than speed to launch. Most small stores should start with an app. The criteria, the cost arithmetic and a checklist are below, followed by what a custom build includes and what to ask a developer before you hire one. I build custom RAG systems, so I have tried to state the case for apps fairly.
What are the three kinds of Shopify chatbot?
Rules-based, general AI and RAG. They differ in where their answers come from, which decides whether they can be trusted with product questions.
A rules-based bot follows a script: pick an option, get a canned reply. It is predictable and cannot answer anything outside the script. A general AI bot sends the question to a language model that answers from what it learned in training, so it can sound right and still be wrong about your products. A RAG bot retrieves your own products and policies first, then answers from them, so its answers can be traced to your data. See what is RAG and why do LLMs hallucinate for the mechanism.
How do the two routes compare?
The table summarizes the main trade-offs. Neither route wins on every row.
| Criterion | Chatbot app | Custom RAG build | |---|---|---| | Time to launch | Hours to days | Weeks | | Upfront cost | Usually none | Development cost | | Ongoing cost | Recurring fee, often tied to usage | Model usage plus hosting | | Control over answers | Limited to the app's settings | Full: prompt, retrieval, checks | | Data ownership | Held by the vendor | Held by you | | Integrations | Often built in (helpdesk, orders) | Built as needed | | Page weight | App script on your pages | Small script you control | | Maintenance | Vendor | You or your developer |
Read the table against your own situation. A store with a few dozen products and common questions about shipping gains little from the right-hand column. A store with thousands of products and attribute-driven questions gains a lot.
When is an app the right choice?
An app is right when speed and low effort matter and the questions are routine. Typical cases are order and shipping questions, returns policy, store hours, and simple product questions on a small catalogue.
Apps also bring integrations that take real work to rebuild, such as connecting to a helpdesk, handing off to a human, or looking up an order. Evaluate an app on a short list: how it syncs your catalogue and how often, whether you can restrict it to your own data, what happens when it does not know, what it adds to page weight, and how pricing scales with conversations. Test it with your 20 hardest real questions before you pay.
When does building your own make sense?
Build when the quality or ownership of answers matters more than launch speed. Four situations justify it.
Large or complex catalogues. When shoppers ask by attributes, use cases or compatibility, retrieval design determines answer quality, and an app gives you little control over it.
Specific behavior. If the bot must follow your rules, such as only recommending in-stock items, never discussing certain topics, or applying your own ranking, a custom pipeline lets you enforce them in code and in output checks.
Data ownership. Your own index and logs stay in systems you control, which can matter for privacy.
Volume. If conversation-based pricing becomes large, a custom build with per-query model costs can be cheaper, but only above a break-even point you should calculate.
How do I compare the cost?
Work out your own numbers with a simple formula, since published prices change and vary by app. For a custom build, monthly running cost is roughly queries per month times cost per query, plus hosting, plus a hosted vector database if you use one, plus maintenance time. Add the one-off build cost on top.
In my own RAG project, one query cost about one cent, covering one embedding call and one GPT-4o call. That is a figure for that project, and yours depends on how much product text goes in each prompt and which model you use. As an illustration only: at one cent a query, 10,000 queries a month would be about $100 in model usage. Compare that with an app's quoted monthly price at your expected conversations, including any overage rules, and decide where the break-even sits. Ask any developer what happens to cost if your traffic doubles.
What about using Shopify's own catalogue tools?
There is a third route: let your agent query Shopify directly instead of building your own index. Shopify provides a Storefront Catalog MCP server that lets AI agents search a single merchant's catalogue, with search_catalog, lookup and get_product tools.
It saves you the sync work, and it has constraints. Shopify's usage guidelines say not to cache search results or re-use images, that queries are rate limited, and that keyless access cannot get rate-limit increases. Some fields may be inferred by Shopify's AI and vary in accuracy. It suits a live agent that always asks Shopify, and not a system where you want your own ranking, your own policy content or your own index. Check Shopify's current documentation, since endpoints and limits can change.
How do I decide? A checklist
Answer these five questions, and the pattern usually becomes clear.
- Are most questions routine (shipping, returns, order status)? That points to an app.
- Do shoppers ask by attribute, use or compatibility across a large catalogue? That points to a custom build.
- Do you have a developer or budget for a build and for upkeep? Without one, choose an app.
- Will you receive enough conversations that usage pricing becomes significant? Calculate the break-even.
- Do you need answers constrained by your own rules or checked in code? That points to a custom build.
If the answers are mixed, start with an app and keep a log of the questions it fails on. That log becomes the specification if you later build your own, and it tells you whether the build is worth it. If the checklist points to a custom build, the next sections cover what you would be paying for.
What does a custom Shopify chatbot build include?
A complete build has seven deliverables. Each one is a place a cheaper build tends to cut corners, so check that all seven are in the scope.
- Catalogue sync. A first full export, then product webhooks so the index follows your store when products change.
- Vector index. Product text, plus policy, shipping, returns and sizing content, stored for search by meaning.
- Retrieval. Finding the right products per question, ideally with a reranking step.
- Grounded prompt. Instructions that restrict the model to the retrieved products and define what to do when nothing fits.
- Safeguards. Product links and prices built from the retrieved results and live store data, never from the model's text, plus a check for invented items.
- Theme section. A custom Liquid section with an accessible form and streamed answers, with no performance hit on your pages.
- Operations. Rate limiting, an origin allowlist, an output token cap, and a log of questions and zero-result queries.
The technical detail behind each item is in how to add an AI chatbot to Shopify, and the grounding and output checks are in building a RAG chatbot that doesn't hallucinate. If you want search rather than conversation, the same index powers an AI product finder.
What should I prepare before hiring a developer?
Bring four things, and the estimate will be much more accurate. They also show whether a chatbot is the right tool.
- Your catalogue size and how it is described. Number of products, and whether descriptions, tags and metafields hold the details shoppers ask about.
- The 20 questions you most want answered. Real ones from email, chat and social messages, including the ones that currently cost you time.
- Your policies in text. Shipping, returns, sizing and care, because a chatbot cannot answer these from product data alone.
- Your rules. What the bot must never say or promise, such as medical claims, delivery guarantees or discounts.
If your product data is thin, expect to spend part of the budget on improving it. A retrieval system can only return what the catalogue contains.
What questions should I ask a developer?
Ask questions whose answers show how the bot behaves when it fails. A demo that answers easy questions proves little.
- How do you stop it recommending products that do not exist? Look for a restriction to retrieved items and a check on the output, not only a prompt line saying "do not hallucinate".
- Where do prices and links come from? The right answer is live store data and retrieved handles, since a model can write a plausible link that goes nowhere.
- How will we test it? Ask for a fixed set of real questions, including some your catalogue cannot answer, and a way to re-run them after each change.
- What happens when nothing matches? It should say so and ask a clarifying question, not guess.
- Who owns the code, the index, the prompts and the logs? The answer should be you.
- How is the endpoint protected? Rate limits and an origin allowlist protect your model bill.
- What does it do to page speed? The script should be deferred and small.
A developer who cannot answer these plainly is likely to deliver a demo, not a system.
What does a build process look like?
A sensible build runs in four stages, with something you can test at each one. I would work this way.
- Discovery. Review the catalogue, the 20 questions and the policies, and decide whether a chatbot or a finder fits.
- Prototype on your data. Index the catalogue or a sample, and answer your real questions in a test page, outside your live theme.
- Evaluation. Run a fixed question set, fix the failures, and agree what counts as good enough before launch.
- Launch and monitor. Add the section to a duplicate theme, test on phones, publish, then review the logs for questions it misses.
Timelines depend on catalogue size and the integrations you want, so I give an estimate after seeing your catalogue and questions rather than quoting a number up front. The first stage is where most of the risk is found.
What proof of experience should I look for?
Look for retrieval systems that work on real data, and for Shopify theme work you can inspect. My RAG projects are the K-Drama RAG engine, over 838 dramas and about 6,234 indexed chunks with two-stage retrieval and a grounded prompt, and the Manhwa Recommender, with hybrid search over a cleaned dataset. My Shopify work includes custom-coded sections and variant selectors for client stores such as BAOBAB, Curet and Naturabites.
I have not published a Shopify chatbot case study yet, so the retrieval pipeline and the Shopify theme work are separate pieces of evidence that a chatbot project would combine.
When should I not hire for this?
Skip a custom build when your store is small, your questions are routine, and you have no budget for upkeep. A chatbot app, or a good FAQ page and clear policy pages, will serve you better. A custom bot also needs someone to watch the logs and fix the questions it misses, so budget for that, not only for the build.
If you are unsure, send me your store URL, your catalogue size and your five most common customer questions through the contact page, and I will tell you whether a custom build, an app or a simpler fix is the right fit. You can also see my Shopify development services.
Related
- How to add an AI chatbot to Shopify (custom RAG)
- How to build an AI product finder for Shopify
- Building a RAG chatbot that doesn't hallucinate
- Hire a RAG developer for a custom chatbot
Sources
Frequently asked questions
Should I use a Shopify chatbot app or build my own?
Use an app if you need simple questions answered quickly and want to launch this week. Build your own when the catalogue is large or complex, when you need control over how answers are produced, or when you want to own the data.
What are the three kinds of chatbot?
Rules-based bots follow scripted flows, AI bots answer from a language model's general knowledge, and RAG bots retrieve your own data first and answer from it. A RAG bot is the one that can stay grounded in your catalogue.
Is a custom RAG chatbot cheaper than an app?
It depends on volume. An app usually has a recurring fee, and a custom build has an upfront cost plus per-query model charges and hosting. At low volume an app is usually cheaper overall, and the balance can shift as volume grows.
Can a chatbot app use my own product data?
Many do, by syncing your catalogue. Check how often it syncs, what fields it reads, and whether it can be told to answer only from your data before you commit.
What does a custom Shopify AI chatbot build include?
A catalogue sync, a vector index, retrieval, a grounded prompt, safeguards against invented products, a chat or finder section in your theme, rate limiting, and logging of what shoppers ask.
How do I know a custom chatbot will not make up products?
Ask the developer how answers are restricted to retrieved products, how links and prices are produced, and for a test set that includes questions your catalogue cannot answer. Look for an output check, not only a prompt instruction.
Who should own the code and the data?
You should. Agree in writing that the repository, the vector index, the prompts and the logs are yours, and that the backend can be handed to another developer.
Can I use Shopify's own catalogue tools to build an agent?
Shopify provides a Storefront Catalog MCP server for AI agents that search a single store. Its guidelines say not to cache results or images, and requests are rate limited.
Want an AI that actually knows your catalogue?
I build Shopify chatbots and product finders that answer from your real products instead of guessing — custom-coded into your theme, with no monthly app fee. Tell me what your store needs and I'll tell you what it takes.


