Busan AI development · RAG

RAG chatbots fail on data before they fail on models

For a chatbot that answers from internal documents, where the documents are and what state they are in decides most of the outcome. Here are six things to check before starting and why each one shows up as answer quality.

Direct answer

In one paragraph

Before building an internal-document RAG chatbot, check six things: which documents exist and which version is current, who may see what (permissions), whether duplicates and old versions are mixed in, whether personal or contract data is present, whether there is a question set to judge answers against, and who updates documents when they change. SMK Labs runs this check as the data-connection step, connects documents and working knowledge into a searchable, citable answer flow, and then validates wrong answers, omissions, permission errors, cost, and latency — on site across Korea including Busan, or remotely.

No specific model, accuracy figure, or deployment environment is promised up front; they are set in the proposal after confirming security and operating conditions.

Checklist

Six things to check before starting

The second sentence of each item is the symptom you actually get when it is missing.

  1. 01

    Document inventory and latest versions

    Which documents will ground answers and where the current version of each lives. Without this, the chatbot confidently cites an outdated policy.

  2. 02

    Access rights

    Whether the rules for who may see which document already exist in the repository. Without them, the chatbot answers beyond permissions — or cannot answer at all.

  3. 03

    Duplicates, versions, and formats

    Whether the same content exists in several copies, and whether key facts sit inside scanned images, tables, or attachments. Duplicates produce conflicting answers; tables in images produce empty ones. OCR and document processing may come first.

  4. 04

    Personal and contract data

    Whether documents contain personal data or contract terms, and whether they may be sent to external model APIs. This condition changes architecture and cost the most and must be settled before starting.

  5. 05

    An evaluation question set

    Twenty to fifty real user questions with the documents that hold the correct answers. Without them, only a feeling that "it seems to work" remains, and validation cannot measure wrong answers or omissions.

  6. 06

    Who updates, and how often

    Who reflects document changes and when. Left undefined, the chatbot starts going stale in its first month of operation. This belongs in the operations and maintenance scope.

How we work

How SMK Labs handles RAG

Problem framing (which questions must be answered) → data connection (the six checks above; quality and security design) → build (a searchable, citable answer flow) → validation (not only correct answers but wrong ones, omissions, permission errors, cost, and latency) → operations (updates, observation, handover). RAG is often combined with workflow automation, AI agents, and API integration, so scope is set together in problem framing.

Common questions about preparing data for RAG

Q01What data do we need for an internal-document RAG chatbot?

An inventory of grounding documents with current versions, per-document access rules, the state of duplicates and old versions, whether personal or contract data is present, twenty to fifty real questions with their correct sources, and who updates documents. Describe as much of the six as you know and we check the rest together in the data-connection step.

Q02What if our documents are scanned PDFs or images?

OCR and document processing become a preceding step that extracts and reviews text and tables. That scope is listed as a separate item in the quote.

Q03How much accuracy do you guarantee?

No figure is promised in advance. Instead we build the evaluation question set together before starting and, in validation, measure wrong answers, omissions, permission errors, cost, and latency against it to agree on acceptance criteria.

Q04Is it possible where external model APIs cannot be used?

The deployment environment and model choice are set in the proposal after confirming security and operating conditions. Whether external APIs may be used changes architecture and cost the most, so it helps to state it at the inquiry stage.

Start with the questions to answer and the documents you have

Scattered documents are fine. Tell us which questions must be answered and where the documents live, and we will guide you from the data-connection checks.

Send a paid AI development inquiry