SPACESJ AI Work Layer
Home / Solutions / RAG Assistants
S/02 · RAG Chatbot & Agent

RAG Assistants

The quality of an assistant is decided by retrieval quality and how it shows its evidence, not by which model it runs on.

Solutions

Answers grounded in your own documents. Sources shown, and no answer at all when the evidence is thin.

Duration
1 week + 2 weeks + 1-2 weeks + 1 week

What usually goes wrong

Pain 01

The document exists but nobody finds it

Policies, manuals and past quotations are somewhere on the drive, but asking a colleague is still faster.

Pain 02

The same question, answered again

Support, administration and sales assistance answer the same questions every day.

Pain 03

A general model invents answers

A model that does not know your internal rules produces plausible wrong answers, which is worse than no answer.

How we approach it

  1. 01Collect and normalisePDF, HWP, Office, wiki and database content is normalised into a text layer, tables and images included.
  2. 02Chunk and embedChunked on meaning at paragraph level and embedded with metadata for department, version and validity period.
  3. 03Hybrid retrievalVector similarity plus BM25 keyword search plus reranking. This is essential for Korean proper nouns and part numbers.
  4. 04Grounded generationThe answer is written only from the retrieved passages and shown with its source. When the evidence is thin it does not answer.
  5. 05Evaluate and fall backA 100-question golden set measures accuracy and evidence fit; below the threshold the conversation is handed to a person.

What you receive

Retrieval pipelineIngestion, chunking, embedding, retrieval, reranking, generation and evaluation end to end
Admin consoleDocument upload, version replacement, index status, blocked keywords
Conversation log analysisA report of unanswered questions, fed back as content priorities
Widget and APIWeb widget, internal messenger integration, REST API
Evaluation reportAccuracy, evidence rate and refusal rate against the 100-question golden set

How it runs

Document review1 weekInventory of target documents, formats and permission structure
Pipeline build2 weeksIngestion, indexing and retrieval implemented
Golden-set QA1-2 weeksWrite 100 questions and iterate on accuracy
Integration and launch1 weekWidget and permission integration, log dashboard
Stack
Python 3.12 / FastAPI pgvector / Qdrant BGE-m3 embeddings Cross-encoder reranker LangGraph Langfuse Redis Docker
Related reference
National dealer-network distributor
-61%first-line CS volume
FAQ

Frequently asked

What is a RAG assistant?

RAG stands for retrieval-augmented generation. The assistant searches your own documents before it writes anything and generates only from the passages it retrieved, so the source of an answer is your organisation documents rather than the model memory.

How do you stop it hallucinating?

It is constrained from writing anything outside the retrieved evidence, and when similarity falls below the threshold the response switches to an offer to check and come back. Every answer shows its source document and page.

Will our documents be used to train a model?

No. Customer data is never used for training and each customer index is isolated. We use only APIs that guarantee no training use, and we can run an on-premise model on request.

Can we limit which documents each department sees?

Yes. Permission tags are attached to document metadata and the search space is filtered by the user permissions at query time, so documents outside their scope never enter the results at all.

Can you handle HWP files and scanned PDFs?

HWP files are processed by text extraction and scanned PDFs through OCR. Documents with many tables are indexed with the table structure preserved so numeric questions stay accurate.

What happens when a document is revised?

When a new version is uploaded in the admin console the previous version drops out of the index, and documents past their validity date are excluded from search automatically.

How long does it take, and what do we need to prepare?

Typically four to six weeks. From your side we need the list of target documents, the permission rules, and about 30 of the questions you receive most often.

A 30-minute call first

We do not quote a rate before the scope is defined. We map your current flow on a call and reply with a draft architecture and a staged cost range within a week.

Talk to us