Before you put AI on top, build information worth putting it on. By the end of the previous chapter the parts are in place — foundation, gate, documents, code, mail, meetings, web, API; the tools stand. But the information itself that flows over them is still scattered. Files sunk to the bottom of a shared folder, paper and scanned PDFs, and the heaviest of all — tacit knowledge that lives only in someone's head. Here we move those three into a written, structured state.
Preparation is the main body, AI the last move
Do not get the order wrong. RAG, and your own AI, stand only on prepared information. Put the cleverest model you like on scattered, unwritten information, and what comes out is just as unprepared (garbage in, garbage out).
So this chapter comes before the AI. The work is three things — read the paper (OCR), align the scatter (classification and structuring), write out what is in people's heads (codifying tacit knowledge).
And this preparation has two properties.
- It is judgment only people can make. What to keep, what to discard, how to structure it — this is the judgment of people who know the work, and it cannot be handed off wholesale to AI (1-04). AI helps with the draft, but people decide.
- It is a no-regret investment. Structure the information and you recover it even if you put no AI on it at all. Personnel lock-in dissolves, handover gets easier, the business gets healthier. AI is merely the last move laid on top.
Prepared information beats a cleverer model. Preparation is the main body. AI is the last move.
Turn paper into text with OCR
The first barrier is information a machine cannot read. Paper, scans, image PDFs, handwriting.
- For fixed-form, typeset text, Tesseract and other OSS OCR turn it into text. Tesseract and
ocrmypdfare both in Debian 13's apt. - For complex forms and figure-laden layouts, have an open-weight vision model (the local AI you stand up in 2-16) read it and render it to Markdown.
# example: give a scanned PDF a text layer (OSS OCR)
ocrmypdf --language eng input.pdf output.pdf # searchable PDF + text
Aim the output at text (AsciiDoc or Markdown) and plain text. Do not lock it into a proprietary format, so that later anyone, and any AI, can read it (the principle of 2-07).
Align the scatter by classifying and structuring
Next, align the scattered files.
- One place — gather them into the Forgejo repository decided in 2-07. Files that arrived stay as the files they are.
- Add metadata — type, department, date, version, held in the folder and the filename.
- Structure in Markdown — headings, bullets, tables, into a form both machine and human can read. Let AI draft, and people fix.
Here AI is a powerful assistant. "Classify these 200 files by type and add a summary" — classification and summarization both run on the local model (2-16). But the axis of classification is decided by people. What counts as "the same type" for the business is something only those who know the business can tell.
Write out what is in people's heads
The heaviest, and most valuable, is unwritten knowledge. The vague parts of a spec, the exception handling, the reasons things are the way they are — all in the head of a veteran in charge. When this disappears, the system becomes "it runs, but no one understands it" (the human dependency of 3-05).
The method is the same one used for core logic in 2-12 — interview the people on the ground, have AI draft, and have the ground confirm.
- Ask and record the person in charge, then AI drafts the transcription and the structuring
- The person reads the draft and fixes the errors (this is the verification, and only people can do it)
- Put the finalized version, in Markdown, into the document store
The moment tacit knowledge is written down, it becomes a transferable asset. Independent of whether you ever put AI on it, the company gets stronger right here.
Mount the AI on prepared information
Only once the prepared information is in place comes the next chapter. Embed the written, structured documents into the pgvector of 2-03 and mount RAG on them (2-16).
Skip the preparation and build RAG, and the sources are vague and the answers unreliable. With the preparation done, an AI that answers from your own real data, with citations, stands up cleanly.
RAG quality is decided not by the model's cleverness but by how well the information you mounted is prepared.
How to check you are done
This chapter is done when these five hold.
- A paper form has become searchable text — open the PDF, search for a word, and it is found
- A figure-laden form has been rendered into readable Markdown
- There is one document store, and type, department, date, and version are visible in the folder and the filename
- A procedure only one person knew is now in Markdown, and the version that person read and corrected is in the store
- Documents in the store open as Markdown and plain text, not as a proprietary format
What the human holds
Values the human supplies
- The axis of classification — what counts as "the same type" for the business
- The decision of what to keep and what to discard
- The metadata fields (type, department, date, version) and how they are attached
- The language to run OCR in (the value passed to
--language) - Which people to interview, and permission to record them
- The location of the document store
Actions the AI states before performing
- Deleting or moving the original paper or scans
- Renaming or moving files in the store in bulk
- Sending an interview recording to an outside service
- Filing a draft into the store as final before the person in charge has confirmed it
Versions checked, and when
- Tesseract 5.5,
ocrmypdf16.7 (Debian 13 packages), an open-weight vision model (the AI server of 2-16), pgvector (2-03) - This procedure was written on 2026-07-16 and reviewed on 2026-10-06
- If a version has moved, have the AI confirm the official procedure before proceeding
Summary
Before the AI, prepare the information.
- OCR — paper and scans into Markdown, with Tesseract or a vision model
- Classification and structuring — one place, aligned with metadata and Markdown (the axis decided by people)
- Codifying tacit knowledge — interview, AI draft, the ground confirms (same as 2-12)
- AI is the last move — mount the prepared information on pgvector; RAG is the next chapter (2-16)
Preparation pays off even without AI on top — personnel lock-in dissolves, handover gets easier. Preparation is the main body, and prepared information beats a cleverer model.
In the next chapter, on top of this prepared information, we set up our own AI.