Built from parts you can name

We work with mainstream model providers, ordinary vector and keyword search, and your own systems through their APIs. Nothing here is exotic, and that is deliberate: the parts have to be supported by your team after we leave.

How a request is handled

How a request is handledA question from a person passes access control, then retrieval over the permitted documents, then answer generation with citations, then a check that the sources support the answer. Answers that fail the check, or that fall outside the permitted material, are escalated to a person.Questionfrom a personAccess checktheir own rightsRetrievalpermitted passagesAnswerwith citationsEscalationto a personSource checkcitations must holdunsupported or outside scopeEvery step is logged: prompt, passages retrieved, model version, outcome.
Retrieval decides what is possible; generation only decides how it reads.

Access is applied before retrieval, so the assistant can never return a passage the person asking is not allowed to open. The source check is a separate step from generation: if the citations do not support the answer, it is not shown.

Server racks in a data centre
Server racks in a data centre

Where it runs

Deployments are containerised and described as code, so the same build runs in your cloud tenant, on your own hardware or in ours. Test and production are separate environments, and the evaluation suite runs against both before anything is promoted.

Where a model has to run inside your perimeter, it is deployed the same way — with the accuracy comparison from the pilot attached to the decision.

The parts

Models

Models from the main providers for hosted work, open-weight models inside your own environment where data may not leave it. The choice follows the constraint.

Retrieval

Vector and keyword search over your documents, hybrid ranking, permission-aware filtering, citation back to the source page.

Orchestration

Explicit pipelines and tool-calling agents with step logging, retries, timeouts and a defined hand-over point to a person.

Evaluation

Labelled question sets from your own operation, retrieval and answer metrics tracked per release, failure cases kept as tests.

Integration

APIs and web services against the systems you already run; no rip-and-replace, and no second copy of your data where it can be avoided.

Delivery

Containerised services in your cloud account or ours, infrastructure as code, separate test and production environments.

Observability

Per-request logs of prompt, retrieved passages, model version and outcome, retained according to your policy rather than ours.

Access control

Your identity provider, your groups and roles; the assistant never has wider access than the person asking.

Where the data lives

  • In your cloud tenant, on your hardware, or in ours — decided before any code is written.
  • Documents are indexed where they already are, with no second copy unless the architecture requires one.
  • External model calls only for the material your policy allows; everything else stays inside the perimeter.

What is logged

  • Per request: the question, the passages retrieved, the model version, the answer and the outcome.
  • Retention set by your policy, not ours; the log is exportable for an audit.
  • No training on your data, and no use of it to improve a vendor model.

When data may not leave

Open-weight models run inside your perimeter, with the accuracy trade-off stated in writing before the work starts. This is a real cost, not a checkbox — the pilot measures both options on your own cases so the decision is made on numbers.

What you get at hand-over

  • Infrastructure as code and a deployment procedure your team can run.
  • The evaluation suite, so quality can be re-measured after any change.
  • Runbooks, monitoring and the data-flow description used for review.