Products
Model Training
Problems
What we solve.
The task has to stay on your data. A shared model is the wrong place for it.
We train for one task. You keep the weights and can run them on your machines.
A demo looked convincing, so the model is treated as ready.
Evaluation is agreed before training. The model does not ship until that evaluation is met.
A small set of examples is forced through a full training run it does not need.
Adaptation when the set is small. A fuller run when the task has to live on your machines.
Tech stack
What we build on.
PyTorch and Nvidia for the run. Open weights as the start. The result is yours, not a model we keep.
PyTorch
NVIDIA
Llama
Mistral
Qwen
Python
Docker
Your weights
Llama mark via selfh.st, CC BY 4.0.
What it does
Model Training produces a model for a task you can describe and test. A way of answering inside one curriculum. A way of reading one kind of page. A voice or a phrasing that has to match a set of examples you own. The result is weights you receive, not a feature that only runs inside our account.
Conversation, Documents, Content Studio, Education Apps, Graphics Studio and AI Film Studio can call a model like this. They do not have to. Training is its own engagement, with its own scope.
The task comes first
We do not start from a pile of files and a wish to “have AI”. The task is a sentence a person outside the project can understand, plus the examples that count as right and wrong. If we cannot write the wrong cases, we are not ready to train.
- The task, in one paragraph.
- The inputs the model will see in real use.
- What a good answer is, shown on examples.
- What a bad answer is, including the tempting wrong one.
- What the model must refuse.
The data you are allowed to use
Training uses data you clear, and only that data. You say what it is, where it lives, who it is about, and whether it may be used for this task. A folder someone forwarded is not clearance.
If the set is small, we say so, and we shape the work as an adaptation on a modest set of good examples. If the task is narrow and the set is large enough for a fuller run, we say that too. We do not pad a small set with material you did not clear.
Evaluation before the run
The checks are agreed before training, and they are written so a later person can rerun them. A score with no definition is not a check. The set used to judge the model is not the set used to train it.
We look at the failures, not only the average. A model that is right on average and wrong on the case you care about — a grade, a dose, a price, a name — is not ready. You see those failures before anyone calls the run a success.
Adaptation or a fuller run
Adaptation fits when you have a modest set of strong examples and the base behaviour is already close to the task. It is the smaller piece of work. A fuller training run fits when the task is narrow, the data supports it, and the model has to behave that way even when it runs on your machines, away from a general service.
Which of the two we recommend is a written decision, with the reason. You can refuse it. We would rather stop before a run than train the wrong shape of model.
Weights you keep
When the run meets the checks, the weights are delivered to you, under the terms written before training started. We do not keep a private copy as the product. How you store them, who may load them, and where they run is yours.
Deployment — putting the model where Conversation, Documents or another product can call it — is a separate increment. Delivery of the weights is the close of the training engagement.
On your machines
“On your machines” means the model can run in an environment you control: your servers, a private network, a machine in the room where the work happens. Discovery names that environment before we promise it. If the model is too large for that place, we say so before the run, and we change the task or the shape of the model.
A demo that only runs in our notebook is not the delivery.
What your data is not used for
Your data stays out of any shared model. It is not mixed with another customer’s material. It is not used to improve a model we offer generally. People who are not on this engagement do not get access because they work at Incubics.
If a file was included by mistake, the response is to remove it from the run and to tell you, not to quietly keep the checkpoint.
A tutor, a reader, a voice
A curriculum-specific tutor can be a Model Training engagement beside Education Apps. The studio still has a teacher approving what a student sees. The model is how the tutor stays inside that curriculum when a general model will not.
A reader for one kind of document can sit beside Documents, with the same rule: a person checks the page. A voice or a phrasing that must match your examples can sit beside Conversation or Content Studio. In each case the product or studio remains the place the person works. The model is the part you own.
Who it is for
Teams whose curriculum, records, voice or pages should not sit in a general model, and whose task is narrow enough to test. It is the wrong engagement if the task cannot be judged, or if the data cannot be cleared.
What you get
- A written scope: task, dataset, evaluation, data handling, and where the model must be able to run.
- A decision, in writing, between adaptation and a fuller run.
- Training on infrastructure you can account for.
- The failures as well as the score, before anyone calls it done.
- Weights delivered to you.
- Your data kept out of any shared model.
- A separate, optional increment to deploy the model where your other products call it.
What we will not ship
We will not train on data you did not clear. We will not keep the weights. We will not skip the evaluation and call a demo a model. We will not promise a general model “for the whole company” when the task we can actually test is one job.
How it meets the other work
Conversation can speak through a model you own. Documents can read with one. Content Studio can draft with one. Education Apps can tutor with one. Graphics Studio can picture with one. None of them require it for a first release. When you do train, the person-approves rule in that product or studio stays.
The first release
One task. One cleared dataset. One evaluation you have read. One delivery of weights. A note of where those weights run, even if the wiring into another product waits. We do not stack three tasks into the first run to make the engagement look larger.
Commercials
A training engagement. Scope, dataset and evaluation are agreed, and priced, before the run. Delivery of the weights is the close. Deployment is the next increment if you want it. If the data or the task will not support the checks, we say so before the run and we do not spend the fee pretending.
Questions
Is this the local-model offering?
Yes. The model is trained for your task, you keep the weights, and it can run on your machines. Conversation, Documents and Content Studio can call it. They do not have to.
Who owns the weights?
You do, under the terms written before training starts. We do not keep them as our product.
What if we only have a small set of examples?
Then we look at adaptation, or we tell you the set is too small to train. We do not pad it with data you did not clear.
When is the evaluation written?
Before the run. The judging set is separate from the training set. You see the failures, not only an average.
Will this data train a model you offer to others?
No. It stays in this engagement.
Can a student-facing tutor be a trained model?
Yes, beside Education Apps. A teacher still approves what the student sees. The model does not replace that approval.
Is deployment included?
Delivery of the weights is included. Putting the model where another product calls it is a later increment.
What do you need from us before a price?
The task, a look at the data you are allowed to use, and the place the model has to run. Without those, a number would be a guess.
Next step
Start with two weeks.
A fixed-fee discovery gives you a ranked use-case portfolio, a target architecture, a cost model and a build proposal you can take to your board. If we don't find a case worth building, we tell you.