Incubics

Capabilities

Retrieval and tool integration

Retrieval and tool integration is the work that makes assistants and agents useful: finding the right facts and performing the right actions in your systems. Done poorly, you get hallucinations and brittle scripts. Done well, you get grounded answers and reliable automation with clear failure behaviour.

What it is

Retrieval spans vector search, keyword search, graph traversal and live queries against warehouses — unified behind policies that respect tenant, role and data residency. Tool integration wraps your APIs and databases as typed functions agents can call, with validation, retries and circuit breakers.

This capability is rarely visible to end users but determines whether production systems survive real load and real permissions.

Design decisions

  • Chunking and metadata strategy per content type.
  • When to embed versus when to query structured data directly.
  • Reranking and citation assembly for user trust.
  • Tool schema design — small, composable functions beat mega-endpoints.
  • Auth propagation — user context versus service accounts.
  • Caching and freshness SLAs for expensive queries.

When it pays back

Every assistant and agent investment depends on this layer. Payback shows up as higher answer accuracy, fewer tool errors and lower inference cost when retrieval replaces long context windows.

It is also the enabler for incremental rollout: add a new tool, add evals, promote to production without rebuilding the whole agent.

Common triggers

  1. Pilots worked on static exports but fail on live data.
  2. Agents timeout calling slow legacy APIs.
  3. Security blocked production because service accounts were over-permissioned.
  4. Duplicate retrieval stacks sprawl across teams.

How Incubics engineers it

Discovery: two weeks mapping systems, access patterns and latency. Week twelve: production retrieval and tool layer for the first use case, reusable for the next.

Perceive — weeks 1–2

We catalogue sources, API maturity, rate limits and identity models. We prototype retrieval quality on sample questions and tool reliability on non-production sandboxes. Output: integration architecture, security model, eval plan and fixed scope.

Engineer — weeks 3–10

We build ingestion and index pipelines, hybrid retrieval services and tool adapters with OpenAPI or internal SDK patterns. Observability traces each retrieval and tool call. Load tests match expected peak concurrency.

Deliver — by week 12

Production endpoints wired to the first agent or assistant, documented schemas, secret rotation procedures and runbook for index rebuilds and API degradation modes.

Run — ongoing

We monitor retrieval precision proxies, tool error budgets and cache hit rates. New sources and tools onboard through the same eval and security checklist.

Failure modes

  • Embedding stale PDFs while live pricing sits in SQL nobody connected.
  • Tool schemas that accept free-text where enums are required.
  • Retrieval returning cross-customer chunks — tenancy bug, not model bug.
  • No timeout strategy — one slow ERP call freezes conversations.
  • Logging full payloads including secrets.

Teams sometimes over-index on vector databases when a simple indexed query would be faster and exact. We choose mechanisms per question type.

What you get

  • Retrieval service with hybrid search and reranking.
  • Ingestion pipelines and index lifecycle automation.
  • Tool gateway with auth, validation, rate limits and audit.
  • Developer documentation and contract tests for each tool.
  • Tracing integrated with your APM or OpenTelemetry stack.
  • Performance benchmarks and capacity guidance.
  • Patterns for adding sources without fork-lifting the agent.

What we refuse to ship

We refuse tools with blanket admin credentials. We refuse retrieval without tenant filters when data is multi-customer. We refuse undocument tool behaviour that agents depend on.

Do we need a vector database?

Not always. Discovery matches store choice to query patterns. Many systems combine warehouse SQL, search engines and vectors.

How do you handle API rate limits?

Queuing, backoff, caching and pre-fetch where business rules allow. Agents receive graceful degradation messages, not infinite retries.

Can tools call mainframe systems?

Through stable API layers — see legacy integration. Direct terminal scraping is not a production strategy we endorse.

How is retrieval tested?

Question-answer pairs with expected source IDs, plus manual review samples. Metrics track recall of correct chunks on held-out sets.

Is this reusable across multiple agents?

That is the design intent. Shared tool and retrieval services reduce duplication and centralise security review.

What about real-time data?

We define freshness requirements per field. Some answers pull live; others use indexed snapshots with visible timestamps in the UI.

Next step

Start with two weeks.

A fixed-fee discovery gives you a ranked use-case portfolio, a target architecture, a cost model and a build proposal you can take to your board. If we don't find a case worth building, we tell you.