Incubics

Products

Conversation

Conversation is the product for work that has to be spoken. A lesson, a support call, a follow-up. It listens, it answers from the job you set, and a person can take over.
Illustrative footage, Mixkit. Not a piece we have released.

Problems

What we solve.

  • The job is spoken, but the only tool on offer is a chat box beside it.

    Conversation holds a live turn: it listens, it answers from the job you wrote, and a person can take over.

  • Text to voice and speech to text live in different places, so the words and the audio drift apart.

    Both sit in this product. A person can correct a name, a number or a term before it is stored or acted on.

  • A voice updates a record without the person hearing what will be written.

    A lookup or an update happens only when the job allows it, and the confirmation is part of the turn.

Tech stack

What we build on.

Open weights, Nvidia, and paid models. We pick the one that holds the voice, the transcript and the cost of this job.

  • Llama
  • Mistral
  • Qwen
  • NVIDIA
  • OpenAI
  • Anthropic
  • Gemini
  • Python

Llama mark via selfh.st, CC BY 4.0.

What it does

Conversation holds one spoken job. You name who the voice is for, what it may say, which records it may open, and the moment a person steps in. Text to voice and speech to text live here, because a spoken job needs both.

A finished piece that is written, translated and then read aloud belongs in Content Studio. A live turn — someone speaking, the system answering, a person able to join — belongs here.

The job you set

Before anyone hears a voice, the job is written down. The audience. The purpose. The facts it may use. The facts it must not invent. The systems it may read. The systems it may change. The phrases that mean a person should take over.

That written job is what we test against. A pleasant voice with no job is not a release. If the job changes, the tests change with it.

  • Who is speaking, and who they are speaking to.
  • What a successful turn sounds like, in plain language.
  • Which records may be opened, and which may be updated.
  • What the voice must refuse, even if asked.
  • The exact moment a person takes the conversation.

Text to voice

Text to voice turns approved words into speech. A lesson line, a status, a short answer, a confirmation of what will happen next. The words are written first. The voice reads those words. It does not improvise a claim you have not allowed.

You choose the voice for the job: a lesson voice, a support voice, a follow-up voice. One job does not need a library of voices. If a second voice is required, it is a later increment, with its own listen.

Anything a student or a customer will hear gets a review listen. A person hears the line in context: the question that came before it, the record it refers to, the action it is about to take. A line that sounds fine on its own can be wrong in the turn.

Speech to text

Speech to text turns a recording or a live turn into words. Transcripts, captions, notes from a call, the student’s spoken answer. The transcript is a working document. A person can correct a name, a number or a term before it is stored or acted on.

A live conversation uses the transcript as the turn the system answers. A recorded conversation uses it as the record you keep. Both stay tied to the audio they came from, so a correction can be checked against what was actually said.

A live turn

A turn is one exchange. The person speaks or the system speaks. The next turn follows the job, not a general chat. If the person asks for something outside the job, the voice says so and offers the path you defined: a person, a form, a later call.

You test the turns that matter before the line goes live. The happy path. The refusal. The moment the record is missing. The moment the person is upset. The moment the voice should stop talking and hand over. We do not treat a demo of the happy path as a finished conversation.

Looking something up

During a turn, Conversation can read a record you already keep: an order, a lesson, a case, a balance, a timetable. It can confirm what it found, in words the person can check. If the record is missing or stale, it says that, rather than filling the gap.

An update — a note on the case, a booked slot, a marked lesson — happens only when the job allows it, and only after the person has heard what will be written. The confirmation is part of the turn, not a silent side effect.

A person takes over

Handoff is designed with the conversation, not added after the first complaint. You name the signals: the person asks for someone, the topic leaves the job, a payment or a grade is about to change, the voice is unsure, a student is distressed.

When the handoff happens, the person who takes over receives the turns so far, the record that was open, and what the voice was about to do. They do not start from a blank line. The voice stops. It does not keep talking beside them.

The review listen

A review listen is a person hearing the voice do the job, on examples you chose. Names, numbers, the next action, and the refusal. For a lesson, a teacher hears it. For a support line, the person who owns the queue hears it.

Lines that fail the listen do not ship. We change the words or the job, then listen again. A voice nobody has heard is not ready for a student or a customer.

A spoken lesson

A spoken lesson uses Education Apps for the teaching and Conversation for the voice. The tutor, the lesson and the check decide what may be said. Conversation speaks it, listens to the student’s answer, and can hand the moment to a teacher.

The teacher still approves what the student hears. A spoken explanation is not a loophole around that rule. If the lesson source changes, the spoken lines are reviewed again.

Support and follow-up

A support conversation is one queue, one job. Status, a simple change, a question the knowledge you approved can answer. It is not a place to discuss anything the company does.

A follow-up is a call or a message you start: a reminder, a confirmation, a check that something arrived. The person can stop it. The voice says who it is and why it is calling, in the first line.

Who it is for

Support teams that want a spoken front to a queue they already run. Schools and education products that want a lesson heard, not only read. Product companies that need speech in one feature, with a person still accountable for what is said.

It is the wrong product if you want an open chat on the website with no owner. That is a thing we will not build.

What you get

  • One conversation, one job, written down and tested before anyone outside the team hears it.
  • Text to voice for the lines that job is allowed to say.
  • Speech to text for the live turn and for the recording you keep.
  • A lookup, and an update only where the job allows it, on systems you already use.
  • A handoff that carries the turns and the open record to a person.
  • A review listen, and a record of what was heard and what changed.
  • A way to stop the voice without taking the rest of the product down.

What we will not ship

We will not ship an open chatbot with no owner and no way to step in. We will not publish a voice nobody has listened to. We will not let the voice invent a price, a grade, a policy or a medical or legal claim.

We will not train a shared model on your calls. Recordings and transcripts stay in the engagement. They are not material for a model we offer to someone else.

How it meets the other work

Education Apps decides the lesson. Conversation speaks it. Documents can turn a scanned chapter into the source the lesson uses, before any voice reads it. Content Studio translates and narrates a finished written piece. Model Training is the path when the voice or the understanding of the job has to be a model you own, running on your machines.

A first Conversation release does not require the other products. They are there when the job needs them.

The first release

The first release is one conversation. One audience. One voice. One set of records. One handoff. A review listen on the turns you named. Further channels, languages and voices are later increments, each with their own listen.

We would rather ship a short conversation that a person can trust than a long one that has not been heard.

Your calls and recordings

You decide where audio and transcripts are stored, who can open them, and how long they are kept. Discovery writes that down before build. A recording is not copied into a shared training set.

If a turn is used to improve this conversation, that is a decision you make, on examples you clear, for this job. It is not a default.

Commercials

Fixed-scope first release: one conversation, one voice, the lookups that job needs, a handoff, and a review listen. The fee and the date are set after we have seen the job, the records and the queue or the classroom it sits in.

A second channel, a second voice, or a model you own are separate increments. We do not fold them into the first date and hope.

Questions

Where do text to voice and speech to text live?

In Conversation. Translation of a written piece, and narration of a finished article or script, live in Content Studio. A spoken lesson uses Education Apps and Conversation together.

Can a person take the call?

Yes. The handoff is part of the first release. The person receives the turns so far and the record that was open.

Is this an open chatbot with a voice?

No. It answers inside one job. Questions outside that job go to the path you defined, usually a person.

Who listens before it goes live?

A named person on your side: a teacher for a lesson, the owner of the queue for support. We do not publish a voice they have not heard.

Can it read and update our systems?

It can read the records the job names, and update only what the job allows, after the person has heard what will be written.

Will our calls train a model you sell to others?

No. Calls and transcripts stay in this engagement. They are not used to train a shared model.

Can the voice run on a model we own?

Yes, when that is the requirement. That model is a Model Training engagement. Conversation can call it. The first release does not have to wait for it.

What does the first release leave out?

Extra channels, extra voices, and extra jobs. Each of those is a later increment with its own listen.

Next step

Start with two weeks.

A fixed-fee discovery gives you a ranked use-case portfolio, a target architecture, a cost model and a build proposal you can take to your board. If we don't find a case worth building, we tell you.