Approach

How we work, from the first rule to the handover

This page is for anyone who wants to know how we'll treat their systems, their data and their customers. It explains the order we work in, and why each step is there.

Write the rules first

Before anything is built, we write down what the system may and may not do. For an agent, that means what it may say, what actions it may take, what data it can touch and who can see what.

The rules are in plain words, so the people who answer for the system can read them, question them and agree to them. Each rule names who's accountable for it, and what should happen when a situation isn't covered.

An example rule and the test it becomes
Rule
The agent only discusses an account after the customer has confirmed who they are.
Test
When an unverified person asks for an account balance, the agent declines, offers to verify them and shares no account details.
If it breaks
The build fails, and the change doesn't ship until the rule holds again.

Turn every rule into a test

Each rule becomes an automated check that runs whenever something changes, whether that's code, a prompt, a setting or the model underneath. If a rule breaks, the build fails and the change stops there.

For agents, the tests include scripted attempts to talk the agent out of its rules. Others ask in unexpected ways, or hide an unfair request inside a fair one. Those tests stay in place for as long as the agent runs.

Keep decisions about data in plain code

A language model is good at conversation, but it can give different answers to the same question. That's fine for choosing words, but it's the wrong way to decide what gets stored, what gets sent or who can see it.

So those decisions live in plain, testable code that gives the same answer every time. The model can ask for something to happen, and the code decides whether it does.

This keeps the decisions that matter most easy to read, easy to test and easy to audit.

Evaluate before and after launch

Testing before launch tells you the rules hold. Evaluation after launch tells you whether the system is still doing a good job as real people use it.

We build a set of realistic questions and tasks with your team, score the agent against them, and run them again whenever something changes. We look at four things:

  • whether answers are correct and based on your own data
  • whether the agent stays inside its rules
  • whether it hands over to a person when it should
  • whether it's getting better or worse over time

We also agree where a person stays in the loop, reviewing, approving or taking over when the stakes are high.

Check accessibility with tools and by hand

Many people will use what we build with a keyboard, a screen reader, zoom or voice control. We aim for the Web Content Accessibility Guidelines (WCAG) 2.2 at level AA on every screen people use.

Automated tests run on every change and catch many problems early. Some barriers only show up when a person tries the real thing, so we also test by hand with a keyboard and a screen reader.

This site is held to the same standard, and its accessibility statement says how it was checked.

Build, measure and hand over

With the rules and tests in place, we build the features, in small releases that real users can try early. We measure what matters to your team, such as time saved, questions resolved or errors caught, and we report it plainly.

From the start, we write the notes, tests and runbooks your people will need. When the work ends, your team owns it and can run, change and extend it without us.

Choose open standards

We prefer open standards wherever they fit: documented APIs, events, open protocols such as the Model Context Protocol, and WCAG for accessibility.

They let you change one tool without rebuilding everything around it. Between now and 2050 you'll probably replace several tools, and open standards make each of those changes smaller.

Ask us how this would work for you

Tell us about the system you're planning or the one you already run, and we'll walk you through how we'd approach it.