2 wks
Knowledge-base assistant
Employees spent hours hunting for answers in policies, manuals and the internal wiki. We built an assistant that answers questions across 12,000 documents and always shows where the answer came from.
- Timeline
- 2 weeks from brief to production
- Solution
- RAG assistant with a web interface
- Team
- Orchestrator and 6 AI agents
- Stack
- FastAPI · PostgreSQL · pgvector
The challenge
Company knowledge was spread across thousands of documents in different formats: PDFs, wiki exports, spreadsheets, scans. Keyword search missed the right page whenever a question was phrased differently from the document.
The assistant had to be trustworthy: answer only from internal documents, cite the source and say so honestly when the answer is not in the knowledge base.
- 12,000 documents in mixed formats, some of them scans
- keyword search does not understand rephrased questions
- an answer without a source cannot be used at work
The solution
- 01
Ingestion and labelling
A pipeline parses documents, runs OCR on scans, splits text into meaningful chunks and keeps metadata: section, date, document type.
- 02
Hybrid search
Chunks live in PostgreSQL with pgvector. Search blends semantic similarity with keywords, so it finds the answer however the question is worded.
- 03
Answers with sources
The LLM composes an answer only from the retrieved chunks and attaches links to the documents, so every answer can be checked.
- 04
Hallucination control
A separate check matches every statement against its source. When there is no support, the assistant says so instead of making things up.
Results
Employees get an answer in seconds instead of digging through folders and the wiki, and every answer can be verified by its link. The first working version was ready two days after the brief; the assistant went live in two weeks.
Metrics are a reference: we check the result on your data with a prototype in 1–2 days.
Common questions
Can you connect our documents and systems?
Yes. We connect files, wikis, databases and internal APIs. We clarify formats and volume in the brief and build the prototype on your real data.
Where is the data stored?
Wherever you need it: in your cloud or on your own servers. Open models can run inside your perimeter so no data leaves it.
How much does a project like this cost?
It depends on the number of documents and sources and on security requirements. After the brief we estimate timeline and budget, and a 1–2 day prototype shows the result you will get.
More cases
Ticket processing pipeline
Classification, enrichment and routing of inbound requests by a fleet of agents with no operator in the loop.
Voice assistant
A phone bot with speech recognition, scenarios and handover of complex dialogues to a live operator.