AI Agent

Agent guardrails: approved claims, banned phrases and disclosures

Set what AI agents may say when they publish for a project, check any draft for wrong or unsupported claims, and decide whether a broken rule warns or blocks publishing.

AI agents can write and publish for you: the Reaudit desktop agent, the Reaudit assistant in the web app, and Claude or ChatGPT connected over MCP. Guardrails are the rules they follow for a project: the claims your team has approved, the phrases nobody may use, and the disclosures that must appear. Agents check their drafts against them, and every agent publish is checked before anything goes live.

No guardrails, no flags

Until a project has at least one approved claim, banned phrase or disclosure, nothing is checked or flagged anywhere, and no AI claim check is run.

Where to find it

  1. 1Open Project settings for the project.
  2. 2Choose the Guardrails tab.
  3. 3Add your rules, choose what happens when one is broken, and press Save guardrails. Agents use them from their next publish.

Who can change them

Everyone on the project can see the guardrails and test drafts. Only the project owner and project admins can change them.

The kinds of rule

  • Approved claims: statements your team has signed off, such as prices, guarantees, certifications and results ("Professional costs €269 per month"). Pick a type for each (pricing, policies, company facts…) and, if it is only true in some countries, list those markets. When an agent checks a draft, approved claims count as the strongest evidence, above your knowledge base.
  • Banned phrases: words or claims that must never appear, such as "guaranteed" or "#1 in Europe". Add the reason (agents see it) and, if you like, the wording to use instead. Phrases match whole words in any language and any capitalisation, also inside headings and formatted text, so "guaranteed" does not match "unguaranteed".
  • Required disclosures: text that has to be in the copy, for example "#ad" on sponsored social posts or a regulatory notice in one country. Choose the channels (articles, landing pages, social posts, ads, email), markets and languages it applies to; leave them empty to apply it everywhere. Matching ignores capitals and extra spaces.
  • Brand rules: plain-language guidance every agent follows when it writes (“Call it a platform, never a tool”). They are never checked or flagged. [More on brand rules](/dashboard/help?article=brand-rules)
Markets are two-letter country codes (gb, us, gr) and languages are language codes (en, el). When an agent does not say which market or language a draft is for, the project’s country and language are used.

Warn or block

  • Off: agents publish without checks.
  • Warn (the default): agents publish, and are told what broke a rule so they can fix it and tell you.
  • Block: an agent cannot publish while a blocking rule is broken. It gets the list of problems, fixes them and tries again. The attempt appears in Agent activity as *Blocked by guardrails*.
Each banned phrase and disclosure is either Blocks publishing or Warning only, so you can block the serious ones and only flag the rest. A claim that contradicts your own evidence always counts as blocking.

What is checked at publish time

When an agent publishes to WordPress or your React site, or publishes, schedules or cross-posts on social media, banned phrases and disclosures are checked instantly. A WordPress draft is checked but never blocked, because it is not live. Saving a draft in the desktop app’s Studio is never blocked either; the agent is told about problems straight away so it fixes them while it writes.
  • Ads: when the assistant creates a Google responsive search ad, or switches a Google or Meta campaign on, the ad text (headlines, descriptions, body, call to action of the campaign’s creatives) is checked first. Pausing a campaign is never blocked.
  • When you publish yourself: Broadcast Review shows guardrail warnings under each caption, and Content Studio shows them above the publish options. They update as you type, and Check claims too runs the claim check. You are never blocked; the warnings are there so you can fix things before they go out.

The claim check

Agents run a fuller check on their drafts before they save or publish. It finds every factual statement the draft makes about your brand (prices, features, policies, figures, company facts; opinions like "easy to use" are left alone) and checks each one against your approved claims and your knowledge base.
  • Approved claim: matches one of your approved claims.
  • Supported: your knowledge base backs it.
  • Contradicted: your approved claims or knowledge base say something different, for example "£29 a month" when the approved price is £49. The check shows what the evidence actually says. This blocks publishing when enforcement is Block.
  • No source: nothing you have addresses it. For the claim types you mark as sensitive (pricing and policies by default) this is a warning: remove it, reword it, or add it to your approved claims once you have checked it.

A missing source is never “wrong”

A claim is only marked contradicted when your evidence clearly says otherwise. If your knowledge base is thin, claims show as No source, not as mistakes. Adding your pricing page, policies and product facts to the Knowledge Base makes the check much more useful.

Test a draft yourself

  1. 1In Project settings → Guardrails, scroll to Test a draft.
  2. 2Paste a post, article or ad, and choose the channel (and a market, if it is not the project’s country).
  3. 3Leave Check factual claims with AI on to include the claim check, then press Check draft.
  4. 4You see what an agent would be told: each broken rule with the passage and a suggested fix, and each claim with its result. The test uses the saved rules, so save changes first.

Change guardrails by asking an agent

You can also change the rules in a conversation, with the desktop agent, the Reaudit assistant, or Claude or ChatGPT over MCP: “ban the phrase *best in class* in the UK”, “add *Pro costs £49 a month* as an approved pricing claim”, “add #ad to every social post”, “switch this project to block”. The agent changes only what you asked and leaves the other rules alone. Claude and ChatGPT ask you to confirm first.
  • Only owners and admins can change guardrails, whether in settings or through an agent.
  • What agents may change is set in the Guardrails tab: All changes (the default), Stricter only, or Off. With *Stricter only*, agents can add banned phrases and disclosures and raise enforcement, but anything that loosens the rules (removing a rule, lowering enforcement, adding an approved claim) has to be done in settings.
  • Adding an approved claim counts as loosening, because it changes what the claim check accepts as true.
  • Every change is recorded in Agent activity, with what changed and whether it loosened the rules. A change the setting refused is recorded as *Blocked*.

For agents and MCP clients

  • get_project_guardrails: read the project’s rules before writing, with each rule’s id and what agents may change.
  • update_project_guardrails: add or remove rules (by id or exact text), change enforcement or sensitive claim types; dryRun previews without saving. Refused with agent_edit_not_allowed when the project’s setting does not allow the change.
  • check_draft_guardrails: check a draft (text, channel, optional market and language; the AI claim check is on unless claims is false). "Would be blocked" means publishing will be refused.
  • A publish refused by the rules returns the error guardrail_blocked with the list of findings.
  • The Reaudit desktop agent and assistant already follow these steps. With Claude or ChatGPT, ask them to "check this against our guardrails before publishing".

Usage

Banned phrases and disclosures cost nothing to check. The claim check makes two small AI calls per draft and is counted in your project’s AI usage (Project settings → Usage). Each project can run up to 50 AI claim checks a day; after that, drafts are still checked for banned phrases and disclosures, and claim checks start again the next day (UTC).
guardrailscomplianceupdate guardrailsagent changeslegalapproved claimsbanned phrasesdisclosuresfact checkclaimsbrand safetyblock publishingdesktop agentmcpclaudechatgpt