Incubics

Insights

Evaluation before the first user

If you cannot test it, you cannot ship it — evaluation harnesses belong in sprint one, not after launch.

The most expensive place to learn that a model fails is in front of a customer. The second most expensive is in front of a regulator. The cheapest is in CI on a Tuesday afternoon, against a suite your domain experts helped define.

Evaluation before the first user is not perfectionism. It is how production AI stays production AI when prompts change, models upgrade and retrieval corpora drift.

What "evaluation" means in production

Not a one-off benchmark score. Not a demo script that always passes.

An evaluation harness is an automated suite that runs against your system — model, prompts, retrieval, tools, guardrails — before release. It includes:

  • Representative tasks from real workflows
  • Edge cases your operations team already knows about
  • Adversarial inputs: prompt injection, off-topic requests, attempts to exfiltrate data
  • Compliance scenarios where policy applies
  • Regression checks so a fix for one failure does not reopen another

Results are pass/fail or scored against thresholds you agreed in discovery. Releases do not ship when the suite fails.

That is the same contract you have with any other business-critical software.

Why pilots skip it

Pilots optimise for speed to impress. Someone writes twenty test questions by hand, runs them once, saves the screenshots.

That is not a harness. It does not run in CI. It does not cover tool failures. It does not update when the corpus changes. It does not block a deploy.

When the pilot goes live without a harness, quality becomes a matter of hope and user complaints. Engineering learns about failures from support tickets. Compliance learns from audit findings. Neither is fast.

Building the harness with the product

We start evaluation design in discovery and implement it in the first sprint — alongside authentication, logging and the first integration.

Domain experts contribute cases. Not hundreds on day one. Start with thirty to fifty that represent volume and pain: the invoice layout that always breaks OCR, the policy question that has a wrong answer in the old FAQ, the account type that must never be auto-updated.

Engineers make them runnable. Fixed inputs, expected outputs or acceptable output classes, clear assertions on tool calls where agents act.

CI runs the suite on every change. Prompt edit, model version bump, retrieval index refresh — all gated.

Humans review what automation cannot judge. Tone, subtle policy nuance, new case types. Sampling review complements the harness; it does not replace it.

Metrics that matter

Accuracy alone misleads. A system can be accurate on easy cases and fail silently on high-risk ones.

We track suites by category: routine workflow, edge case, adversarial, compliance. We track tool success rates separately from text quality. For agents, we track whether the right tool was called with the right arguments — not only whether the final message read well.

Thresholds are agreed with sponsors. A drop in compliance scenarios is a hard stop even if average accuracy rises.

Drift is why the harness keeps running

Production changes underneath the model. New products in the catalogue. Updated policy PDFs. Seasonal volume. Users phrasing questions differently.

Drift detection combines harness runs on a schedule, monitoring on live traffic samples and alerts when metrics move outside bands.

Managed run includes ongoing evaluation — not only uptime pings. A model can be "up" and wrong.

Red-team versus regression

Regression suites protect against known failures. Red-teaming searches for unknown ones: jailbreaks, indirect injection via retrieved documents, tool chains that exfiltrate data.

Both belong in the programme. Red-team findings become regression cases so fixes stick.

Common objections

"We do not have golden answers for generative outputs." You often have acceptable classes: must cite source, must refuse, must escalate, must not contain PII. Evaluation tests structure and behaviour, not only exact strings.

"Evaluation slows us down." Incidents slow you down more. A suite that runs in minutes saves weeks of reputational repair.

"The model will keep changing." Exactly why versioning and suites matter. Upgrade decisions become data: new model passes or it does not ship yet.

What you should ask any build partner

  1. Where does the evaluation harness live — repo, CI, ownership?
  2. Who writes and approves test cases from the business?
  3. What blocks a production deploy today?
  4. How are retrieval corpus changes tested before they go live?
  5. What happens to evaluation after go-live?

Vague answers mean you are buying a demo path again.

Maintaining the harness after launch

Evaluation is living infrastructure. New case types from support become cases. Regulatory changes add scenarios. Model upgrades rerun the full suite.

Assign an owner — often product or ops with engineering support — for harness maintenance. Without ownership, suites rot and releases slip back to manual spot checks.

Quarterly review of categories and thresholds catches drift in what "good" means as the business changes.

Working with subject matter experts

SMEs do not write code. They review cases, label acceptable outputs and sign off compliance scenarios. Sessions are structured: show failure, capture expected behaviour, prioritise by risk.

Start with high-risk, high-volume paths. Perfect coverage is not the goal on day one. Repeatable process to add cases is.

Non-determinism and acceptable variance

Generative outputs vary. Evaluation defines acceptable classes — must cite section 4.2, must refuse, must escalate — rather than exact string match except where required.

Flaky tests erode trust in CI; invest in robust assertions.

Performance regression

Harness tracks latency p95 on standard cases. Model upgrades that slow workflows fail release even if accuracy improves — ops cares about both.

Sharing results with sponsors

Monthly eval summary in steering: pass rates by category, new cases added, failures fixed, red-team highlights. Transparency beats surprise at audit.

Evaluation as part of Perceive

Discovery defines the evaluation approach for the first release: categories, thresholds, adversarial scope, who owns cases after handover.

Engineer implements it in week one of the build. Deliver includes a passing suite as part of the definition of done. Run keeps it current.

If you are planning an AI release without a harness, you are not planning a release. You are planning an experiment on users.

Write to hello@incubics.com if you want discovery to map what evaluation should look like for your use case — before the first user finds out for you.

Evaluation tooling choices

Tooling ranges from custom scripts to platforms — fit team skills and integration needs. Avoid tooling that cannot run in CI headless.

Store eval datasets with same access controls as source data — eval leaks are still leaks.

Version datasets when sources change — tag eval set v3 with corpus version reference.

Human eval sampling complements automation — weekly sample review for categories hard to assert automatically.

Benchmark public leaderboards rarely predict your production behaviour — invest in private suites.

Eval runtime budget — full suite should complete in practical CI time or split smoke versus nightly full.

Flaky eval erodes culture — fix or quarantine flaky cases immediately.

Eval results dashboard for non-engineers — sponsors see pass rates without reading JSON.

Evaluation is the contract between product, engineering, compliance and ops — enforce it before users do.

Read How we work for where evaluation fits in Perceive, Engineer, Deliver and Run. Discovery defines categories and thresholds before sprint one — not as a post-launch audit exercise.

Next step

Start with two weeks.

A fixed-fee discovery gives you a ranked use-case portfolio, a target architecture, a cost model and a build proposal you can take to your board. If we don't find a case worth building, we tell you.