Applied AI
Language models, trained on your business.
Copilots, retrieval and agents grounded in your own data — measured before they meet a customer, and running on infrastructure you own.
How We Work With Models
We train models. We don’t just prompt them.
Grounded in your data
A model trained on the open web knows everything except your business. We start with your documents, tickets and records — and every answer cites where it came from.
Measured before it ships
We score against cases your team already answered correctly. The evaluation set is yours, runs on every change, and outlives whichever model is fashionable next quarter.
Right-sized, not oversized
The largest model is rarely the right one. We test the small self-hosted option against the frontier API and tell you which tasks are worth owning — including when the answer is no model at all.
Owned by you
Weights, prompts, pipelines, evaluations and infrastructure, in your own accounts. No orchestration layer only we can maintain, no licence on your own data.
Evaluation
The scoreboard comes before the model.
We agree what “good” means and how it gets counted before anything is trained. This is the shape of the report you get back.
94.2
Answer accuracy
Matched the human answer on held-out cases
0.91
Brand-tone match
Scored against your own published writing
99.1
Safety / policy
Refused what it should refuse
96.5
Citation coverage
Answers traceable to a source document
How a Build Runs
01
Curate
Your exports, documents and transcripts, cleaned and labelled. Most of the quality of the finished system is decided here.
02
Fine-tune
Supervised fine-tuning, adapters, preference tuning — or none of them, when retrieval alone already clears the bar. We try the cheap thing first.
03
Evaluate
Accuracy, tone, safety and citations on held-out cases, plus prompts written to make it misbehave. Nothing ships on a number we can’t show you.
04
Deploy
Your cloud account, your keys, your bill — sized to actual traffic. Where a hosted API is genuinely the better answer, we say so.
05
Monitor
Tracing on every call, drift watched against the evaluation set, and a cost model that accounts for retries rather than the headline token price.
Capabilities
Fine-tuning & training
Teaching a base model your domain, your formats and your tone.
Retrieval & search
Answers grounded in your documents, with permissions enforced at retrieval — not at the prompt.
Agents & automation
Systems that take actions in your tools, bounded by what they’re allowed to touch.
Evaluation & safety
The scoreboard, the red-team set, and the tracing that shows why an answer happened.
Inference & deployment
Models running on your own infrastructure at a cost you modelled in advance.
Data curation & labelling
The unglamorous work that decides whether any of the above is worth doing.