Every enterprise now has the same idea: a system that answers questions over our own documents, accurately, with sources. The idea is right. The budget conversation usually is not, because it is anchored on the demo. A retrieval-augmented demo over fifty clean PDFs genuinely takes a week, and it works well enough to raise expectations that the production system then has to survive. We have built these systems for research, insurance, and travel organisations across Southeast Asia and Europe, and the pattern is consistent: the demo is one tenth of the work, and the other nine tenths are decided before any model is called.
The demo takes a week. The system takes a quarter. The difference is everything your documents are hiding.
Why the demo misleads
Demo corpora are curated: recent, text-first, uniform, and free of the questions that make knowledge management hard. Real enterprise corpora are none of those things. They hold scanned documents with no text layer, tables whose meaning lives in their structure, figures and charts that carry the actual evidence, five versions of the same policy with no marker for which is current, and documents that certain roles must never see. Retrieval quality has a hard ceiling set at ingestion: what your pipeline throws away during parsing, no model downstream can recover. Teams that budget for the model and treat ingestion as plumbing get a system that answers fluently from the fraction of the corpus it actually understood.
The seven decisions that shape the build
Costs and timelines in a RAG engagement are downstream of seven decisions, most of which are organisational rather than technical:
- Corpus scope and ownership: which document sets are in the first release, who owns their correctness, and who arbitrates when two documents disagree. Undefined ownership becomes hallucination blamed on the model.
- Ingestion depth: text-only parsing, or full multimodal handling of tables, figures, and scans. This single choice moves the answerable share of a technical corpus dramatically.
- Access control inheritance: the retrieval layer must respect the source systems' permissions, so the answer to a question never leaks a document its asker could not open. Retrofitting this is far harder than designing for it.
- The evaluation harness: a golden set of questions with verified answers, built with the business before the system exists. Without it, quality arguments are anecdotes.
- Freshness pipeline: how new and updated documents flow in, how stale ones retire, and how fast a correction propagates to answers.
- Citation policy: whether every answer must show its sources. For enterprise trust the answer is yes, which constrains architecture more than most teams expect.
- Failure behaviour: what the system says when it does not know. A confident wrong answer costs more than an honest refusal.
Notice that only one of the seven is a model choice. This is why quotes for the same project vary so widely: vendors who price the model price the demo, and vendors who price the seven decisions price the system. The gap between those two numbers is not margin. It is scope one of them has not noticed yet.
The architecture that survives contact with users
Version one is usually a single retrieval chain, and for one well-scoped corpus that is correct. The ceiling appears when corpora multiply and question types diverge: a single agent stops scaling and a routed, specialised architecture takes over, with its own coordination costs. The practical guidance is to design for the ceiling without building past it on day one: hybrid retrieval (keyword plus semantic) from the start, a reranking stage once the corpus is meaningfully large, citations wired through every path, and instrumentation on retrieval quality, not just answer quality, so you can tell which half of the system is failing.
Governance is a feature, not a tax
For regulated organisations in Singapore and across Southeast Asia, the knowledge base is also a data-protection surface. PDPA and GDPR obligations follow the documents into the vector store: personal data needs redaction or lawful basis at ingestion, deletion requests must propagate through embeddings and caches, and every answer needs an audit trail of what was retrieved and shown to whom. Built in from the start, this is engineering. Retrofitted after a regulator asks, it is a rebuild. It is also, quietly, a sales feature: the enterprises most worth serving are the ones that will ask about it in procurement.
What a realistic engagement looks like
- Discovery, 2 to 3 weeks: corpus audit (formats, quality, permissions, versioning), use-case prioritisation, and the golden-question set agreed with the business. This is where the real scope emerges.
- Pilot on one high-value corpus, 6 to 12 weeks to a working system: full ingestion pipeline, retrieval and generation, citations, and evaluation gates that must pass before anyone calls it done.
- Hardening: access control inheritance verified role by role, monitoring and drift alerts, load behaviour, and the failure modes rehearsed rather than discovered.
- Rollout with adoption measurement: usage, answer acceptance, and escalation rates by team. A knowledge base nobody queries is a very expensive index.
- Steady state: freshness pipeline running, evaluation set growing with real user questions, and a quarterly review of what the system still cannot answer.
The pattern behind all of it: treat the knowledge base as a production data system with an LLM at the end, not an LLM project with some documents attached. That framing is what separates the builds that reach a second year from the demos that die in procurement. It is also how we run AI consulting engagements: the model is the easy part, and everything around it is the work.



