The chatbot they asked for, and the rulebase it needed first
A consultancy asked us for a chatbot to answer eligibility questions. Their rules weren't written down in one place, so we built that first. The assistant now answers only from the official rules, and shows which rule it used and when that rule was in force, including for past cases.
The practice wanted a chatbot for eligibility criteria. But those criteria lived across spreadsheets, saved prompt text, and consultants' memory, with no single current version and no record of what the rules had been in the past. A chatbot on top of that would have sounded confident and fluent while being wrong in ways no one could check, and the fluency is what makes it dangerous, because it removes the friction that would prompt someone to double-check.
The rules change often, sometimes retroactively, and the people who update them aren't engineers. Under Canadian licensing rules, only licensed representatives can tell an applicant whether they qualify, so the assistant had to be internal-only, and built so it's structurally incapable of giving an eligibility determination, not just told not to.
We flipped the order of the ask: build the official rulebook first, the chatbot second. Every rule is stored bitemporally, meaning we record both the period the rule applied to, and separately, the date we learned about it, because rules here are often announced retroactively. The assistant can't make anything up from general knowledge; every answer is a lookup against that store, and if a rule doesn't exist yet, it says so instead of guessing.
A single official rulebook the practice now treats as its source of truth, an editing process their own team runs with a review step built in, a test suite, a tool that can replay any past decision against the exact rules in force on its date, and an internal assistant that answers in plain language with the rule, its version, and the dates it applied, including questions about any date in the past.
We first designed the system to record only when a rule took effect. Then the rules changed mid-build, retroactively, and our design couldn't tell the difference between what applied at the time and what we only learned about later. That broke the assistant's whole premise, since answering "what applied as of this past date" only works if you track both dates. We rebuilt it to track both, at a cost of about a week. The problem that broke our first design was also the proof it needed fixing: in this line of work, a rule change mid-project is the normal case, not the exception.