A document assistant for a construction company that answers in Slovak or English, cites the exact page and runs on the company's own machine.

Case Study / 15 of 15

Work context
Built for a construction company
Status
Pilot since 2025
Scope
Retrieval-Augmented Generation / Knowledge Graph / Local LLMs

Local Graph RAG

Local Graph RAG project visual. A question in English answered from Slovak contracts, with numbered citations and each claim checked against the passage it cites

01 / Overview

Answers that point to the page

Contracts, budgets, site diaries and inspection protocols arrive in Slovak and English, some of them only as scans. Finding a deadline, a penalty or the subcontractor on a roof meant opening PDFs one by one.

The assistant reads those files into a search index and a graph of the concepts and quantities they share. Every answer carries numbered citations, and a citation opens the PDF on the cited page with the passage highlighted. I designed and built the system end to end: parsing, retrieval, the Python API, the React interface and the evaluation suite.

02 / Challenge

A wrong number is worse than no answer

Two projects with the same set of documents and similar figures sit side by side in the archive. A model that mixes them up, rounds a penalty or invents a deadline does real damage. Every figure has to come from a source the reader can open, and the system has to say when the documents hold no answer.

The documents could not leave the company. Everything runs on one Apple machine with open models, so accuracy had to come from retrieval and verification, with a 4-billion-parameter model writing the final answer.

Trace of one question with dense, keyword and reranking stages and the model call on a timeline. Every question is one trace, so a slow or wrong answer can be followed to the stage behind it
Knowledge graph of passages, concepts and quantities across the document archive. Passages link the concepts and quantities they mention, which ties together documents about the same sites and subcontractors

03 / Approach

Retrieve carefully, then write

Docling turns each file into text and tables, with OCR for scanned pages. Every chunk gets a sentence of context about its document before bge-m3 embeds it into LanceDB, next to a BM25 keyword index. A question runs both searches, merges them with reciprocal rank fusion, follows the knowledge graph when it asks how documents relate and reranks the candidates with a cross-encoder. A local qwen3 model writes the answer from the top passages only. A checker then marks any sentence its cited passage does not support, and figures from tables are computed in code.

04 / Decisions

Why this stack

Most of these choices were tested against an alternative on the evaluation set. The rest follow from where the system has to run.

  • Instead of a hosted API

    Local models via Ollama

    Contracts and budgets never leave the machine. qwen3 4B writes the answers and a 1.7B model extracts the graph, both beside the embedding model on a 48 GB M4 Pro. The instruct build replaced the hybrid thinking model, which leaked its reasoning into answers even with thinking switched off.

  • Instead of a vector database server

    Embedded LanceDB

    The index is a folder next to the app, with nothing extra to run, back up or secure on the client's machine. Vectors, chunk text and the access-group filter live in one table, so a user never retrieves a document outside their groups.

  • Instead of dense vectors alone

    Hybrid search, then rerank

    Dense search alone put the right page in the top ten for 90.2% of questions. A context sentence on every chunk gave the largest gain, BM25 keeps exact names and codes findable, and a cross-encoder reranker moved the right passage to first place far more often. Together they reach 99.2%, with the reranker costing about 0.85 s per query.

  • Instead of graph expansion on every query

    Graph on demand

    Expanding every query through the graph pulled in neighbours that pushed the right passage down, and mean reciprocal rank fell from 72.9 to 51.4. A router now sends only questions about how documents relate, such as which sites one subcontractor works on, through the graph.

  • Instead of letting the model read tables

    Numbers computed in code

    Table rows are parsed into typed records. The model only maps the question onto fields and values, while filtering, sums and lists run in code. In an early test the model listed one of eight matching rows and invented a limit by subtracting two numbers from the question.

  • Instead of LangChain or LlamaIndex

    No framework, full traces

    Each stage is a small Python function with its own tests and its own OpenTelemetry span. Spans follow the GenAI semantic conventions, so every model call records its model, tokens and latency, and a bad answer can be traced to the stage that caused it. Traces stay local and can be exported to Langfuse or Phoenix.

  • Instead of trusting uploaded files

    Guarded retrieval

    A document can carry instructions aimed at the model. Sources enter the prompt marked as data, passages that read like instructions are flagged and stripped of links, and images in answers are never loaded. On a set of poisoned documents the attack success rate fell from 16.7% to 0%.

05 / Results

Scored on 128 questions

The evaluation set is synthetic on purpose: 12 linked construction documents in Slovak and English, and 128 questions with the exact page that holds each answer, ten of them unanswerable. It can be shared without exposing client data, and every configuration was compared on the same questions.

  1. 99.2%

    Recall@10

    The page that holds the answer is among the first ten passages for 99.2% of questions, up from 90.2% with dense search alone.

  2. 100%

    Valid citations

    Every citation points to a passage retrieved for that answer, and 95.8% of answers contain all of the expected facts.

  3. 0%

    Prompt injection success

    Poisoned documents tried to plant a canary word, an exfiltration image and a fake price. None succeeded, and every question was still answered correctly.

Next Project

01 / 15

Web Design and Development / 3D Configurator / CRM