Why does a RAG assistant still invent answers?
In most cases the cause is retrieval rather than the model. When an irrelevant passage enters as evidence, the model faithfully uses that passage and produces a plausible wrong answer. The fault therefore sits in the retrieval stage, not in generation.
Which single measure helps most?
Setting a similarity threshold for the evidence and refusing to answer below it. A configuration that leaves a refusal rate of 8-12% earned the most trust in practice. A refusal is designed behaviour, not a failure.
Why is vector search alone not enough in Korean?
Many queries turn on an exact string rather than a meaning - part numbers, model names, clause numbers. Vector similarity blurs those tokens, so it has to run alongside BM25 keyword search with a reranker to reorder the candidates.
How should document revisions be handled?
Attach effective and repeal dates as metadata at clause level rather than document level, and filter the search space by the query date before retrieval runs. A withdrawn document left in the index will be cited sooner or later.
How do you measure quality?
Build a 100-question golden set and watch accuracy, evidence fit and refusal rate together. Tuning on accuracy alone pushes you to remove refusals, which produces a more dangerous assistant.