Before you switch on an AI agent: a guardrails checklist

This checklist is for the team that has to say yes before an AI agent works with real customers and real data. That's usually a product owner, an architect, a security lead and whoever will run the agent once it's live. An agent speaks in your organisation's name and acts on your systems, so the protections that matter most have to be in place before launch. We've set them out in the order we work through them.

The checks don't depend on any particular platform or model. They apply whether your agent answers customers, drafts replies for staff or updates records quietly in the background.

1. Write the rules first

Before anyone writes a prompt, write a short document that says what the agent may say, what it may do and what data it can touch. Name who can see each kind of information, and list the topics the agent must pass to a person. Keep it in plain language, so a lawyer, a support lead and an engineer can all read it and disagree early.

Each rule should be specific enough to test. A rule like "never quote a price that isn't in the current price list" is one you can check automatically. "Be helpful and accurate" is too vague to check at all. We give every rule a number, because each later step points back to it.

This is the same move the NIST AI Risk Management Framework makes with its Govern function. NIST describes governance as a cross-cutting function that informs the other three, which are Map, Measure and Manage. The framework is voluntary and isn't tied to any sector, so it gives a mixed team a shared vocabulary for its rules. If you'd like a formal management system around this work, ISO/IEC 42001:2023 sets out requirements for establishing, running and continually improving an AI management system.

2. Give the agent its own identity and the least access it needs

An agent should run under its own identity, with permissions you've chosen on purpose. List every record type, field, file store and tool it can reach, and whether it can read, create, change or delete each one. Then remove everything the rules from step 1 don't need.

When the agent acts for a signed-in person, it should act with that person's rights and no more. The OWASP Top 10 for LLM Applications 2026 ranks Excessive Agency third and traces it to excessive functionality, excessive permissions or excessive autonomy. Its guidance is to minimise the tools an agent can call, minimise each tool's permissions and run tools in the user's context. It also asks you to carry the original user's context through chained agent calls, so one agent can't borrow another's wider access.

OWASP's own example is worth checking for. A team needs an agent to read documents, but the tool it picks can also edit and delete them. Review the tools themselves, as well as the role they run under.

3. Keep storing and sending decisions in code

Some decisions are too important to leave to a model's judgment. Whether a message goes out, whether a record is saved and who receives a copy should be decided by ordinary code you can read and test. The model can draft, summarise and suggest. The code then decides whether that draft is stored or sent, based on the rules you wrote down.

OWASP's 2026 guidance makes the same point in security terms. It advises implementing authorisation in logic rather than relying on the model to decide whether an action is allowed. For prompt injection, it recommends keeping credentials and the power to change state in application code. Privileged calls should then pass a deterministic policy check at the moment they run.

The reason is practical. OWASP says there's no reliable way to prevent prompt injection today, because models don't draw a firm line between instructions and data. So we design as if someone will eventually talk the model into something, and we make sure the code around it will refuse.

4. Decide where a person stays in the loop

Write down which actions need a person's approval before they happen. Good candidates are anything irreversible, anything that moves money, anything visible outside your organisation and anything your rules mark as sensitive. OWASP's guidance asks for explicit human confirmation before privileged, irreversible or externally visible actions. It also suggests showing the reviewer the exact action, because a summary can hide what will really run.

Handoff needs the same care. Decide when the agent passes a conversation to a person, who picks it up and what context travels with it. The customer shouldn't have to repeat themselves, and the person taking over should see what the agent said and did. Test the handoff in both directions, including what happens outside working hours.

NIST's Map function includes defining, assessing and documenting the processes for human oversight. Its Generative AI Profile treats the human and AI arrangement as a risk of its own, including over-reliance and automation bias. An approval step only protects you when the reviewer has the time and the information to say no.

5. Tell people they're talking to an AI

People should know when they're dealing with an AI system, and they should know it from the start. Say so in the first message, in plain words, and make the route to a person easy to find. Keep the notice in real text, so screen reader users meet it at the same moment as everyone else.

In the European Union this is now a legal duty for many providers of AI systems. Article 50 of the EU AI Act says AI systems that interact directly with people must be designed so those people are told they're dealing with an AI. That applies unless it's already obvious from the context. The information must be clear and distinguishable, given at the latest at the first interaction, and it must meet applicable accessibility requirements. The European Commission's timeline shows these transparency rules applying from 2 August 2026. If your agent reaches people in the EU, check your exact obligations with counsel.

6. Test answers and refusals before launch

Build a test set straight from your rules. For each numbered rule, write questions the agent should answer well and questions it should refuse or hand to a person. Include ordinary requests, awkward edge cases and the way real people write when they're upset or in a hurry.

Then add adversarial prompts. Try to talk the agent out of its rules, and hide instructions inside the documents or web pages it reads. Ask it to reveal its own instructions or someone else's data. OWASP warns that defences which look strong against a fixed list of attacks can fail against attackers who adapt. So bring in testers who know exactly how your defences work.

Run the whole set automatically, and make the build fail when a rule breaks. NIST's framework says AI systems should be tested before deployment and regularly while they're in operation. We treat a failed refusal test with the same seriousness as a failed payment test.

7. Evaluate after launch

Launch day is when the real questions start arriving. Log what the agent was asked, what it answered, which tools it called and what the code decided, and keep personal data within your retention rules. Review a sample every week at first, with the person who owns the rules in the room.

Track a small set of measures tied to your rules: correct answers, correct refusals, handoffs and actions a person had to reverse. NIST's Generative AI Profile suggests monitoring and documenting the cases where people override a system's decisions, then evaluating those cases. When a review finds a gap, add a test before you change the prompt, so the same mistake can't quietly return.

8. Plan for incidents and rehearse the rollback

Decide before launch what counts as an incident, who gets told and who is allowed to switch the agent off. NIST's framework calls for mechanisms to supersede, disengage or deactivate AI systems whose results don't match their intended use. Its Generative AI Profile adds protocols to make sure a system can be deactivated when necessary.

Make that switch real. Keep the previous version of the agent's instructions, tools and settings ready to restore, and practise the rollback in a test environment. Write down how you'll tell the people affected and how you'll record what happened. NIST's Manage function asks that incidents and errors are communicated to the relevant people, including affected communities, and that the process is documented.

When you can answer all eight

If you can answer every item above in writing, you're ready for a careful launch. If a few answers are still blank, you now have a clear list of what to build next.

This is how we build every agent at Havihi Digital. We write the rules first, turn them into tests and keep the important decisions in plain code. Then we build, measure and hand over something your team can own and run.

Sources