Capabilities
Evaluation, guardrails and audit
What it is
Evaluation is structured testing of models and agents against scenarios that matter to your business: factual accuracy, tool correctness, refusal behaviour, latency and cost. Guardrails enforce policies on inputs and outputs — PII redaction, topic restrictions, tool allow lists, injection detection. Audit records who triggered what, with which model version, and what data was touched.
Together they form the control plane around generative systems. Application features sit on top; evals and guardrails gate every release and run continuously in production.
Layers
- Offline evals — golden datasets, red-team prompts, regression on every change.
- Online evals — sampled production traffic scored for quality and safety.
- Input guards — injection detection, rate limits, authentication context.
- Output guards — schema validation, policy classifiers, human review queues.
- Tool guards — argument validation, spending caps, write approval.
- Audit store — append-only logs with retention and access controls.
When it pays back
This capability pays back the moment you have more than a toy user base — when a bad release would reach customers, corrupt records or create regulatory exposure. It also pays back internally when teams argue about whether the model got worse: evals replace opinions with trends.
Organisations with model risk management, SOX-adjacent controls or customer data obligations often require audit trails before production approval. Building evals early is cheaper than retrofitting after launch.
Symptoms you need it now
- Prompt changes ship without tests.
- Nobody can answer what the model did on a specific ticket last Tuesday.
- Security raised prompt injection and you have no plan.
- Vendor model upgrades break behaviour silently.
How Incubics engineers it
Discovery: two weeks to define metrics, policies and evidence requirements. By week twelve you have eval gates on the release path and audit logging in production for the scoped system.
Perceive — weeks 1–2
We workshop with risk, legal, security and operations to translate policies into testable rules. We collect scenarios from historical failures and near-misses. Output: eval catalogue, guardrail specification, audit schema, reporting cadence and build estimate.
Engineer — weeks 3–10
We implement harnesses integrated with CI and staging, policy middleware on agent paths and centralized logging with correlation IDs across tools. Red-team exercises run before deliver. Dashboards show pass rates and guardrail triggers.
Deliver — by week 12
Release blocked unless eval thresholds pass. Production emits audit events to your SIEM or warehouse if required. Runbook covers incident classification, model rollback and comms templates.
Run — ongoing
We expand eval coverage from production failures, tune guardrails to reduce false positives and produce monthly evidence packs for governance forums. Model upgrades go through the same gates as day one.
Failure modes
- Vanity metrics — BLEU scores while users get wrong answers.
- Guardrails that block everything or nothing — no calibration.
- Logs without identity — cannot attribute actions.
- Eval sets that never update — pass while reality drifts.
- Audit as afterthought — storage costs explode, fields missing.
Over-trust in automated judges is its own failure mode. We combine model graders with human-labelled subsets and hard assertions on tools and schemas.
What you get
- Eval harness integrated with your delivery pipeline.
- Golden and red-team datasets versioned with the codebase.
- Policy middleware for inputs, outputs and tools.
- Audit logging schema and retention configuration.
- Dashboards and scheduled reports for governance.
- Incident runbooks for quality and safety regressions.
- Documentation mapping controls to tests — for your reviewers, not ours.
What we refuse to ship
We refuse production agents without attributable logging. We refuse shipping prompt changes that skip regression evals. We refuse guardrails that silently drop violations without alerting.
Can this bolt onto a pilot built elsewhere?
Often yes, after an assessment. Retrofit takes longer than building alongside features. Discovery sizes the hardening work.
Do you certify compliance frameworks?
We implement controls and evidence your auditors review. We do not sell compliance certificates or invent certification badges.
How often do evals run?
On every build for offline suites; sampled continuously in production for online checks. Frequency scales with traffic and risk tier.
What about third-party models?
Evals are model-agnostic. Upgrades from providers trigger the same regression gates. Contracts cover data handling; technical controls cover injection and logging.
Who owns the golden dataset?
You do. We help build and maintain it; it lives in your repositories and warehouses.
Next step
Start with two weeks.
A fixed-fee discovery gives you a ranked use-case portfolio, a target architecture, a cost model and a build proposal you can take to your board. If we don't find a case worth building, we tell you.