AI Knowledge Base for Magento 2: How Our RAG Engine Connects, Chunks, and Vectorizes Your Data



Every AI feature we’ve built, the chatbot that answers shoppers, the reporting tool that answers sales questions, the support agent that drafts email replies, runs on the same foundation underneath: a retrieval-augmented generation (RAG) engine that connects to your store’s real data and only ever answers from what it actually finds there. We get asked a lot of the same questions about it: where does the data actually come from, what happens to a document after you upload it, and what does “vectorization” even mean in practice. This post on AI Knowledge base for Magento 2 is the honest, detailed answer, written the way we’d explain it to another engineer, not a sales page.

What “RAG” Actually Means, in Plain Terms

A generic AI model answers from what it was trained on, which is fixed, general, and often out of date. Retrieval-Augmented Generation flips that around: before the model writes an answer, the system first searches your own data for the pieces that are actually relevant to the question, then hands only those pieces to the model and says, in effect, “answer using this, and only this.”

That single design choice is what separates a trustworthy AI feature from one that quietly makes things up. If nothing relevant is found in your data, a properly built RAG system says it doesn’t know, instead of guessing.

Data connectors feeding the AI knowledge base: Magento, Confluence, site crawl, document upload

Where the Data Comes From: Four Real Connectors

The engine is only as good as what feeds it, so we built dedicated connectors for the kinds of content a store actually has, rather than expecting you to reformat everything into one file type.

1. The Magento Connector

This is the core connector for any store using our AI Suite. It syncs directly from your Magento REST API on a recurring schedule: products (including live stock and pricing), CMS pages, and, if you opt in, anonymized order and customer trend data for reporting. Nothing here requires manual exporting. Once connected, it stays current on its own.

2. The Confluence Connector

If your team keeps technical documentation, product specs, or support knowledge in Confluence, the engine can sync directly from it using the Confluence Cloud API. You can point it at an entire space, a specific list of pages, or a single parent page and everything nested underneath it. Each of those can be marked independently as customer-facing or internal-only, so a public spec sheet and a staff-only escalation guide can live in the same Confluence space and still end up treated completely differently by the engine. It re-syncs on a schedule and only reprocesses pages that actually changed since the last run, so a large space doesn’t mean a slow or expensive sync.

3. Website Crawling

For content that lives on your site but isn’t in Magento’s own CMS module, the engine can crawl your sitemap directly and pull in those pages on a schedule, keeping them fresh as your site changes.

4. Document and URL Upload

For everything else, PDFs, Word documents, plain text files, or a specific URL, you can upload or link it directly. PDFs are handled properly here too: text is extracted natively, and image-based or scanned pages are still processed rather than silently skipped. A document can also be attached to a specific product, so a spec sheet or installation guide only ever surfaces when it’s actually relevant to that SKU.

Every one of these connectors feeds the same pipeline below, so it doesn’t matter whether a piece of knowledge came from Magento, Confluence, a crawled page, or an uploaded PDF: it gets treated with the same rigor.

We’re not stopping at Magento. Shopify is next, and more connectors are on the way.

The Pipeline: From Raw Content to a Grounded Answer

AI knowledge base pipeline: extraction, chunking, vectorization, indexing, retrieva

Step 1: Extraction and Cleaning

Each connector’s raw format gets normalized into plain text. A Confluence page’s storage-format HTML is parsed so its actual heading structure is preserved rather than flattened into one wall of text. A PDF’s text layer is extracted directly; where there isn’t one, the page is rendered and processed so the content still makes it in. This matters more than it sounds: a chunker that can’t see where a heading or section actually starts has no way to cut sensibly, and produces chunks that split concepts in the wrong place.

Step 2: Chunking

Long content isn’t useful to a retrieval system as one giant block, and it isn’t useful as arbitrarily-sliced fragments either. The engine supports several chunking strategies, from simple fixed-length slicing to a recursive, structure-aware splitter (built on LangChain’s `RecursiveCharacterTextSplitter`) that tries paragraph breaks first, then sentence breaks, and only falls back to raw character cuts as a last resort, with a configurable overlap between chunks so a concept split near a boundary isn’t lost entirely in either piece.

Step 3: Vectorization

Each chunk is converted into a numerical embedding, a vector that captures its meaning rather than just its exact wording, using a dedicated embedding model (`all-MiniLM-L6-v2`, 384 dimensions) that runs locally rather than depending on a third-party API for this step. This is what allows a shopper’s question like “does this hinge work on a fire door” to match a chunk about fire-rated compliance even if the exact phrase never appears in either one. A live chat query is also given priority over background bulk vectorization jobs internally, so answering a shopper’s question never waits behind a large catalog re-index.

Step 4: Indexing and Storage

Every vector, along with metadata about where it came from (its source, entity type, SKU if relevant, and whether it’s internal or customer-facing) is stored in an index built specifically for that store, backed by Postgres with the pgvector extension by default. Every store’s index lives in its own isolated schema. There is no shared table with a “tenant ID” column to filter by; the data genuinely lives in a separate space, which is a meaningfully stronger isolation guarantee than row-level filtering.

Step 5: Retrieval and Reranking

When a question comes in, the engine doesn’t just grab the single closest vector match and stop there. It runs a wider similarity search, then reranks the results with a blended scoring formula: 40% vector similarity, 25% literal keyword overlap with the question, and 35% a source-type trust weighting (an official CMS or Confluence page is trusted more than a generic linked asset, for instance). Results are also capped per source document, so one long changelog or manual can’t crowd out everything else in the answer. If a question is about a specific product, chunks attached to that exact SKU are reserved a guaranteed share of the results, so a specific, relevant attachment doesn’t get buried under generic catalog noise.

Step 6: Generation

Only after all of that does the retrieved context get handed to the language model, with explicit instructions to answer from that context and nothing else. If the retrieval step found nothing relevant, the system says so honestly instead of generating a plausible-sounding guess. If the first answer looks weak, it can automatically retry with a stronger model rather than shipping a shaky response.

Why This Is a Foundation, Not a Single Feature

This engine isn’t built for one product. It’s the shared base that the AI Chatbot, AI Advanced Reporting, and AI Support Agent all run on. Connect it once, and every service built on top inherits the same grounded answers, the same connectors, and the same store-by-store data isolation. Add a new Confluence space or upload a new spec sheet, and it becomes available to whichever of those services needs it, without repeating any setup.

Frequently Asked Questions

Does this mean the AI is trained on our data, in the machine learning sense?

No, and that’s actually the point. Nothing is retrained. Your content is indexed so it can be searched and retrieved at answer time, which means updating a document takes effect immediately rather than requiring a costly retraining cycle.

What happens if we update a document or a Confluence page?

On the next sync, only what changed gets reprocessed. Confluence pages are compared by version number, so an unchanged page is skipped entirely rather than re-embedded for no reason.

Can we keep some knowledge internal and some public?

Yes. Confluence subtrees and uploaded documents can each be marked internal-only or customer-facing, so the same engine can safely serve a shopper-facing chatbot and an internal support tool without either one seeing the other’s material by accident.

Is our data mixed with any other store’s data?

No. Each store’s vectors are stored in a completely separate schema, not a shared table filtered by an ID column.

What if the AI can’t find a real answer in our data?

It says so. The system is deliberately built to admit when nothing relevant was retrieved rather than generate a confident-sounding guess.

The Bottom Line

The chatbot, the reporting tool, and the support agent are the parts you actually see. Underneath all three is the same disciplined pipeline: real connectors for the data you actually have, careful chunking that respects structure, vectorization that captures meaning, a properly isolated index per store, and a retrieval step that would rather admit uncertainty than fabricate an answer. That’s the difference between an AI feature that feels like a gimmick and one your team can actually rely on.

Similar Posts