Products
Documents
Problems
What we solve.
Scans, forms and handwritten pages become a transcript that scrambles the page.
Digitise the whole page and keep the reading order, or extract only the fields you name.
A table or a column is dropped into a paragraph and called done.
Layout stays. Every block can be checked against the place it came from on the source.
A missing field is filled with a plausible value.
We do not invent a value. A person corrects the block before anyone else uses it.
Tech stack
What we build on.
Language and vision models read the page. The review against the source is the product, not an extra step.
Llama
Qwen
OpenAI
Anthropic
Gemini
PyTorch
Python
Llama mark via selfh.st, CC BY 4.0.
What it does
Documents has two jobs, and you pick the one in front of you. Digitise keeps the whole page: paragraphs, headings, tables, captions, reading order. Extract returns only the fields you name, as structured data.
Both jobs keep a link from the result back to the place on the source page. A person can open that place, correct a cell or a line, and send the corrected text on. A transcript with nowhere to look is not a delivery.
Digitise the page
Digitise is for when the page itself is the work. A chapter, a policy, a letter, a set of minutes, a page of a catalogue. The result keeps the order a person would read: heading, then the paragraph under it, then the table, then the note under the table.
Columns stay columns. A footnote stays a footnote. A caption stays with its figure. We do not pour the page into one paragraph and call the reading order a detail for later.
The digitised page is source material. Education Apps can turn a chapter into a lesson after a person has checked it. Content Studio can draft from a page you trust, not from a guess about what the scan said.
Extract the fields
Extract is for when you need specific facts, not the page around them. Invoice number, date, total, line items. Student name, class, mark. A form field, a stamp, a signature present or absent.
You write the fields before we run the set. Each field has a meaning, an example, and what “missing” looks like. The result is a record, not a story. Empty stays empty. We do not invent a value to make the row look complete.
Tables and mixed pages
A table is a table. Rows, columns, headers, merged cells where the page actually merged them, and a note when a cell was unreadable. A page that mixes a paragraph, a table and a handwritten margin is returned as those three things, not as one blur.
Multi-page documents keep their page numbers. A row that starts on one page and finishes on the next is joined only when the join is visible, and the review shows both pages.
Handwriting and scans
Scans, photographs of pages, and handwriting are ordinary input, not an exception. Quality varies. Where a word or a figure is uncertain, the result says so, and the review opens on that spot.
We do not silently “clean up” a bad scan into a confident sentence. Uncertainty is part of the output. A person decides whether to re-scan, to type the line, or to leave it marked.
Review against the source
Review is the screen where the block and the source sit together. A heading, a cell, a line of handwriting. The person corrects it there. The correction is the version that moves on. The original block stays visible, so you can see what changed.
Who reviews is named before the first batch: an editor, a teacher, a clerk, an operator. A batch is not finished because the software finished. It is finished when that person has cleared it, or has marked what they will not clear.
A chapter for a lesson
Education uses Documents to read a syllabus or a chapter before the lesson model touches it. The teacher, or the person who owns the source, checks the reading. Then Education Apps builds the explanation, the example and the practice from that checked text.
A student does not see the raw scan as the lesson. They see what a person approved, at the reading level the studio set.
Forms and records
A form becomes the fields you listed, checked against the page. The record can then go to the system you already use, as a draft a person confirms. Documents does not write into that system until the review for that form type is in place.
The first release is one form, or one record type, that you already handle by hand. The second type is the next increment, with its own fields and its own review.
Archives
An archive is many pages of uneven quality. We start with a sample you choose: the clear pages, the bad scans, the handwriting, the tables. The sample tells us what “done” means for that archive. We do not promise a whole room of boxes from a look at one clean page.
Pages that fail the sample’s bar are flagged, not forced. You can re-scan them, set them aside, or accept a partial reading with the gaps marked.
Who it is for
Teams sitting on scans, forms, textbooks and archives that a plain transcript would scramble. Schools that need a chapter in a form a lesson can use. Operations teams that retype the same form. Editors who need a page they can trust before they draft from it.
What you get
- Digitise: the page, with layout and reading order kept.
- Extract: the fields you named, as structured data, with blanks left blank.
- Tables and mixed columns that do not collapse into a wall of text.
- A mark where a word or a figure was uncertain.
- A review view that opens the source beside the block.
- Corrections that become the version you export.
- A first document type in production, with the review named, before the rest of the set.
What we will not ship
We will not ship a one-shot transcript with no way to fix it. We will not drop a table into a paragraph and call it done. We will not fill a missing field with a plausible value.
We will not train a shared model on your documents. Pages you upload stay in the engagement.
How it meets the other work
A checked chapter feeds Education Apps. A checked page can be the source Content Studio drafts from. A field set can be what Conversation reads back on a call. If the reading itself has to be a model you own, that is Model Training, and Documents still supplies the review against the page.
The first release
One document type you already handle. The fields or the layout rules for that type. The review, with a named person. A sample that includes the bad pages, not only the clean ones. Export into the place the text needs to go, as a draft until that person clears it.
Further types, languages of the source, and larger batches follow once the first type is being corrected in real use.
Your documents
You decide which pages are in scope, where they are stored, and who may open the source and the correction. Discovery records that before we read a production set. Your pages are not used to train a model we share.
Commercials
Fixed-scope first release: one document type, review in place, a sample that matches the real pile. The date is set after we have seen that sample. A second document type is a separate increment.
Questions
Digitise or extract?
Digitise when you need the page, in reading order. Extract when you need specific fields, such as a form or a record. A document can need both, as two steps.
What happens when a word is unreadable?
The result marks it. Review opens on that spot. We do not guess a confident word to hide the gap.
Can this feed a lesson?
Yes. A chapter becomes source for Education Apps after a person has checked the reading. A teacher still approves what the student sees.
Can a draft in Content Studio start from a scan?
Yes, after the page has been digitised and reviewed. Content Studio drafts from the checked text.
Do you take a whole archive in the first release?
No. The first release is one document type, proved on a sample that includes the difficult pages. The rest of the archive follows that pattern.
Will our documents train a shared model?
No. They stay in this engagement.
Who corrects the output?
A person you name: a teacher, an editor, a clerk. The software prepares the page. That person clears it.
Can the fields go into a system we already use?
Yes, as a draft the reviewer confirms. We connect the export in the first release when that system is part of the job.
Next step
Start with two weeks.
A fixed-fee discovery gives you a ranked use-case portfolio, a target architecture, a cost model and a build proposal you can take to your board. If we don't find a case worth building, we tell you.