RAG Assistants
The quality of an assistant is decided by retrieval quality and how it shows its evidence, not by which model it runs on.
Answers grounded in your own documents. Sources shown, and no answer at all when the evidence is thin.
What usually goes wrong
The document exists but nobody finds it
Policies, manuals and past quotations are somewhere on the drive, but asking a colleague is still faster.
The same question, answered again
Support, administration and sales assistance answer the same questions every day.
A general model invents answers
A model that does not know your internal rules produces plausible wrong answers, which is worse than no answer.
How we approach it
- 01Collect and normalisePDF, HWP, Office, wiki and database content is normalised into a text layer, tables and images included.
- 02Chunk and embedChunked on meaning at paragraph level and embedded with metadata for department, version and validity period.
- 03Hybrid retrievalVector similarity plus BM25 keyword search plus reranking. This is essential for Korean proper nouns and part numbers.
- 04Grounded generationThe answer is written only from the retrieved passages and shown with its source. When the evidence is thin it does not answer.
- 05Evaluate and fall backA 100-question golden set measures accuracy and evidence fit; below the threshold the conversation is handed to a person.
What you receive
| Retrieval pipeline | Ingestion, chunking, embedding, retrieval, reranking, generation and evaluation end to end |
|---|---|
| Admin console | Document upload, version replacement, index status, blocked keywords |
| Conversation log analysis | A report of unanswered questions, fed back as content priorities |
| Widget and API | Web widget, internal messenger integration, REST API |
| Evaluation report | Accuracy, evidence rate and refusal rate against the 100-question golden set |
How it runs
| Document review | 1 week | Inventory of target documents, formats and permission structure |
| Pipeline build | 2 weeks | Ingestion, indexing and retrieval implemented |
| Golden-set QA | 1-2 weeks | Write 100 questions and iterate on accuracy |
| Integration and launch | 1 week | Widget and permission integration, log dashboard |
Frequently asked
What is a RAG assistant?
RAG stands for retrieval-augmented generation. The assistant searches your own documents before it writes anything and generates only from the passages it retrieved, so the source of an answer is your organisation documents rather than the model memory.
How do you stop it hallucinating?
It is constrained from writing anything outside the retrieved evidence, and when similarity falls below the threshold the response switches to an offer to check and come back. Every answer shows its source document and page.
Will our documents be used to train a model?
No. Customer data is never used for training and each customer index is isolated. We use only APIs that guarantee no training use, and we can run an on-premise model on request.
Can we limit which documents each department sees?
Yes. Permission tags are attached to document metadata and the search space is filtered by the user permissions at query time, so documents outside their scope never enter the results at all.
Can you handle HWP files and scanned PDFs?
HWP files are processed by text extraction and scanned PDFs through OCR. Documents with many tables are indexed with the table structure preserved so numeric questions stay accurate.
What happens when a document is revised?
When a new version is uploaded in the admin console the previous version drops out of the index, and documents past their validity date are excluded from search automatically.
How long does it take, and what do we need to prepare?
Typically four to six weeks. From your side we need the list of target documents, the permission rules, and about 30 of the questions you receive most often.
A 30-minute call first
We do not quote a rate before the scope is defined. We map your current flow on a call and reply with a draft architecture and a staged cost range within a week.