Applied AI

Language models, trained on your business.

Copilots, retrieval and agents grounded in your own data — measured before they meet a customer, and running on infrastructure you own.

Start a conversation

How We Work With Models

We train models. We don’t just prompt them.

Grounded in your data

A model trained on the open web knows everything except your business. We start with your documents, tickets and records — and every answer cites where it came from.

Measured before it ships

We score against cases your team already answered correctly. The evaluation set is yours, runs on every change, and outlives whichever model is fashionable next quarter.

Right-sized, not oversized

The largest model is rarely the right one. We test the small self-hosted option against the frontier API and tell you which tasks are worth owning — including when the answer is no model at all.

Owned by you

Weights, prompts, pipelines, evaluations and infrastructure, in your own accounts. No orchestration layer only we can maintain, no licence on your own data.

Evaluation

The scoreboard comes before the model.

We agree what “good” means and how it gets counted before anything is trained. This is the shape of the report you get back.

94.2

Answer accuracy

Matched the human answer on held-out cases

0.91

Brand-tone match

Scored against your own published writing

99.1

Safety / policy

Refused what it should refuse

96.5

Citation coverage

Answers traceable to a source document

An example of the format, not a result. Yours come from your own cases — and we bring them to you whether they flatter us or not.

How a Build Runs

  1. 01

    Curate

    Your exports, documents and transcripts, cleaned and labelled. Most of the quality of the finished system is decided here.

  2. 02

    Fine-tune

    Supervised fine-tuning, adapters, preference tuning — or none of them, when retrieval alone already clears the bar. We try the cheap thing first.

  3. 03

    Evaluate

    Accuracy, tone, safety and citations on held-out cases, plus prompts written to make it misbehave. Nothing ships on a number we can’t show you.

  4. 04

    Deploy

    Your cloud account, your keys, your bill — sized to actual traffic. Where a hosted API is genuinely the better answer, we say so.

  5. 05

    Monitor

    Tracing on every call, drift watched against the evaluation set, and a cost model that accounts for retries rather than the headline token price.

Capabilities

Fine-tuning & training

Teaching a base model your domain, your formats and your tone.

SFT · LoRA / PEFT · DPO

Retrieval & search

Answers grounded in your documents, with permissions enforced at retrieval — not at the prompt.

Embeddings · pgvector · Rerankers

Agents & automation

Systems that take actions in your tools, bounded by what they’re allowed to touch.

Tool use · Planning · Guardrails

Evaluation & safety

The scoreboard, the red-team set, and the tracing that shows why an answer happened.

Evals · Red-team · Tracing

Inference & deployment

Models running on your own infrastructure at a cost you modelled in advance.

vLLM · Quantisation · Autoscaling

Data curation & labelling

The unglamorous work that decides whether any of the above is worth doing.

Cleaning · Labelling · Synthesis