1. No Vectors, On Purpose
Ask anyone how to build an AI assistant over company documents and you’ll get the same recipe. Chunk the docs, embed each chunk, put them in a vector database, pull the nearest neighbours for every question. It’s the default basically everywhere.
Our assistant for the factory floor, Ebook Genba, does none of that. No embeddings, no vector DB. When an operator asks why the packaging is bloated, it finds the relevant cases with keyword search: tokens, a stopword list, a weighted score. And it works well.
The knowledge isn’t a pile of PDFs. It’s a structured library of cases, real problems from the production lines, each written up with a title, a summary, a 5M root-cause analysis (man, machine, method, material, measurement), a why-why analysis, and corrective and preventive actions.
The assistant itself is a tool-calling agent. Instead of stuffing retrieved text into a prompt, the model gets a handful of tools (search_cases, get_case, ask_clarification and a few more) and decides when to use them. So someone asks “kenapa kemasan kembung di packaging?” and the model searches, reads the most relevant cases, and answers with a 5M breakdown and corrective actions, linking every case it used.
2. Why Keywords Fit This Job
Mostly it’s the vocabulary. The factory floor talks in abbreviations and codes: WH FG (warehouse finished goods), QC, VM, machine names, standard codes like CRE-012. Those are exactly teh tokens embeddings are bad at. To an embedding model two machine codes look almost identical, to an operator they’re two different machines.
The language is mixed too, Indonesian and English in the same sentence (“gland packing bocor di mixer”), and keyword matching doesn’t care where a word comes from.
Precision also matters more than cleverness here. Vector search always returns something close, and for an operator about to adjust a machine, a confident answer built on a loosely related case is worse than “there’s no case for this yet”.
And the corpus is small. Tens of cases per ebook, hundreds in the whole library. At that size scoring every case in scope is cheap, and a vector index would add infra and a new way to be wrong to solve a scaling problem we don’t have. Plus when search misses something I can see exactly why, which tokens matched and how each case scored. Try doing that w/ nearest-neighbour results.
3. How the Search Works (and How It Got There)
It’s deliberately plain Go.
The query is lowercased and split on anything that isn’t a letter or digit, and tokens of two or more characters are kept. Then a hand-written list of Indonesian and English filler words gets dropped (dan, yang, pada, tidak, di, ke, of, in…). That turned out to matter a lot. Words like tidak (not) show up in nearly every case, and leaving them in made every case match every question.
Each case then gets points per matching token, depending on where it matched:
| Where the token matches | Points |
|---|---|
| Title | 3 |
| Summary | 2 |
| Root cause, why-why, action plan, tags | 1 |
Results are sorted by score with a stable sort, so ties keep a predictable order. There’s also a pass across the whole library, in case the answer sits in an area the user didn’t think of.
None of these rules came from a design doc. Each one came from a reported miss.
The first version fetched a fixed number of cases and then filtered them, so anything past that limit was just invisible (“the case exists, but the AI says not found”). Now it scores the whole scope. The minimum token length was originally three, which silently dropped WH, FG and QC; lowering it to two and adding two-letter filler words to the stopword list fixed a whole class of misses.
Stemming went in and came back out. We added light stemming and put area names into the searchable text to improve recall, and precision dropped. a question about gland packing, which is a seal on a machine, matched every case in the packaging area. Both got reverted. More recall isn’t always better.
4. Keeping the Answer Honest
Retrieval is only half of it. The system prompt requires at least one link to a case, and if nothing relevant is in scope the answer has to say so instead of filling the gap. Early on the model sometimes made up document links that looked right but didn’t exist, so now every case and SOP link is checked against the catalogue on the server and fake ones get removed. Scope is enforced outside the model too. The model isn’t trusted with that.
For measuring, there’s a gold test suite of twenty real questions against a fixed copy of the library, plus an LLM judge where anything under 4 out of 5 fails. Every conversation turn is also one OpenTelemetry span in the same ClickHouse and HyperDX setup as our other services (see From logs to spans). And then there’s operator feedback. In one review cycle 96 pieces came in, 43 were actionable, and 33 of those got fixed and deployed, most of them search misses like the ones above.
We did build something fancier at one point (an automatic judge on every live answer, among other things) and it added moving parts faster than it added value. A simple feedback inbox that a person actually reads turned out to be the better loop.
Would we add vectors eventually? Probably, if we start seeing real synonym misses (so far it’s been tokens and scope, not vocabulary) or the library grows into many thousands of cases. The plan then is hybrid, keywords for codes and abbreviations and vectors for meaning, and the test suite will tell us if it actualy helped.
For now though, tokens and a stopword list are doing just fine.