RETRIEVAL · PLATFORM
Faberiq.ai

The retrieval pipeline behind an AI agent platform's knowledge base

Built the pipeline that turns a customer's uploaded documents into answers with reliable citations, for an AI agent platform, delivered as one piece slotted into the client's existing product.

The situation

An AI agent platform's core promise is that its agents answer from a customer's own uploaded documents, with a citation on every claim. That promise is kept or broken entirely in how documents get processed before an answer is ever generated, and customers upload real-world PDFs, scans, and exports full of tables and nested clauses, not clean text.

The hard part

We were building one piece inside a live product already serving many customers, matching the client's existing interfaces, while their own team kept working on other parts in parallel. Every citation had to point to an exact spot in the source document, which constrains every earlier step in the pipeline.

What we did

We worked backwards from the citation requirement. The document, page, and exact location are attached the moment text is pulled out, and never dropped along the way. Documents are split up following their actual structure, headings, sections, rather than a fixed chunk size, with tables kept together as whole units. We combined keyword search with meaning-based search, and built in an honest "I don't know" response for when the match is too weak, rather than handing the model a thin, unreliable answer to work with.

What we delivered

A production pipeline matching the client's existing interfaces, citations that point to the exact document, page, and passage, a calibrated "I don't know" fallback, and a test set that scores citation accuracy separately from answer accuracy.

What didn't go to plan

Our first attempt split documents by a fixed length. Overall accuracy looked fine, but citations on tables and detailed clauses pointed to fragments that didn't actually back up the claim, and tables carry most of the precise content customers care about. Looking at average results across everything hid a failure concentrated exactly where the product's promise mattered most. We rebuilt the splitting logic to follow document structure instead, which cost about a week.

The result
Document · page · span
the exact location every citation points to, tested on documents the model had never seen, and scored separately from answer accuracy
Industry

AI platform / developer tools

Stack
Pythonpgvector (a database for meaning-based search)Hybrid retrievalOCR with positional output
Duration

7 weeks

Engagement

Component delivery

Contact

Have a project you want built properly?

Start a project →