nqzai / Capabilities / AI Agent Guardrails

The agent that questions the job before it bills you for it

Most AI tools do exactly what you asked and charge you for it. nqzai reads the request first, and when it can already tell the answer will disappoint you, it says so — before the spend, not in the post-mortem.

TL;DR
  • 10 tools ask a question before any money moves — on the same approval card you were already reading.
  • 9 of the 10 use no model at all. They are arithmetic on rows we already hold, so the check itself costs nothing.
  • Answers mark claims they could not verify and end with what went unchecked.
  • It never blocks you. Confirm always still means run it anyway.
  • One check was built and then deleted because it could never fire. That is on the record below.
An nqzai approval card showing a guardrail line above the confirm button: three of twelve recipients will not send

A guardrail is one sentence, in front of the button it affects.

The expensive failure is not a wrong answer

Most writing about AI guardrails is about safety — stopping an agent from saying something harmful. That matters, and it is not where the money goes.

The expensive failure is an agent that does exactly what you asked, correctly, and produces something nobody wanted. It is invisible in every log. Retrieval worked. The tool returned. The bill is real.

The case that started this: a search for dentists returned 25 real dentists, every record correct — and the fit column on each one read “no relevant fit” for what that account actually sells. Nothing failed. One question up front would have caught it.

So the guardrails here are not a content filter. They are a layer that reads the request and the account’s own history, and speaks up when those two disagree.

One promise, two halves

Everything below exists for a single outcome: you are never surprised. Not by a result, not by a bill, not by a constraint that got quietly dropped, and not by a confident sentence resting on nothing.

  • Before the work — the tool about to spend asks whether it will actually help. This happens before the approval card, because a clarification that arrives after you click Confirm has cost you a decision you already made.
  • While the answer is written — every claim is checked against what actually executed. A sentence about a page nobody read is marked as unverified, and the answer ends with a list of what could not be checked.

Same principle at both ends. The system knows what it actually did, so it does not have to be trusted to remember.

Before the spend: what each tool asks

Ten tools carry a check today. Each one asks a different question, and each question is the one a careful colleague would ask if they were watching over your shoulder. Click a row's kind below to see just the arithmetic checks or the one that needs judgement.

ToolWhat it asks before it spendsHow it decides
Lead search Does this audience actually fit what you sell? If not, here is a better one, in one click. Judgement
Send emails How many of these will actually leave? Unsubscribed and verification-failed addresses are subtracted before you approve a number, not after. Arithmetic — no model call
Verify contacts How many of these already carry a verdict from this week? You would be buying answers you hold. Arithmetic — no model call
AI visibility How many answers will this sample actually produce? One answer is not a trend, and the card says so before you read it as one. Arithmetic — no model call
Keyword enrichment How many of these were refreshed recently? Only the stale ones need paying for. Arithmetic — no model call
Backlink gap analysis Is your site URL set? Without it there is nothing of yours to subtract, so the result would be their whole backlink list, not a gap. Arithmetic — no model call
Backlink deep scan How many links are already on file, and how recently? A rescan only pays for what changed, and backlinks accrue over weeks, not hours. Arithmetic — no model call
On-page audit When did you last crawl this site? If nothing changed since, this returns the same findings. Arithmetic — no model call
Index comparison When did you last compare sitemap against index? Same question, same restraint. Arithmetic — no model call
Full Google diagnostic The heaviest run in the product, so the freshness question matters most here. Arithmetic — no model call

Every other outward-acting tool is listed in an exemption map with a written reason. A build check fails if a tool appears in neither — because “nobody thought about it” and “we decided no” look identical in an empty registry.

What it looks like when one fires

Not a modal, not a warning triangle. One sentence on the card you were already reading, in front of the button it affects.

Send 12 emails

Estimated cost shown in tokens before you confirm

Before you confirm — 3 of these 12 won’t go out because they unsubscribed or failed verification, so 9 will actually send.

ConfirmSend to the 9 verified onlyCancel

The number 12 was ours. We put it on the screen, and the send loop would have skipped three of them — correctly, for compliance and deliverability — and told you afterwards. Reporting it afterwards is the defect. You approved a number that was never achievable.

Same shape everywhere else in the layer: a quiet under-delivery against a number we put on the screen ourselves.

Why most of it is arithmetic, not AI

Nine of the ten checks call no model. That is a design rule, not a cost saving.

A model-based check that fires wrongly argues with you about your own market — and that failure is worse than the one it fixes. Users who get argued with once learn to click past the line forever, and then the layer protects nobody. So the rule is: where a check can be deterministic, it must be.

  • “How many of these were verified this week?” — a count.
  • “How many days since the last crawl?” — a subtraction.
  • “How many answers will this sample produce?” — prompts multiplied by engines.

Only one question genuinely needs judgement — whether a requested audience fits what the account sells — and only that one calls a model.

One check was built and then deleted. A contact-enrichment check turned out to sit behind a threshold it could never reach: the call costs about 22,400 tokens against a 75,000-token approval threshold, so the card it would speak on never opens. Running the check would have cost more than the spend it protected. Measuring before building is now the rule.

While the answer is written: grounding the claim

The second half runs after the work starts. nqzai searches the web and reads live pages mid-answer — the research loop covers that in full — and the guardrail on top of it is about what the answer is then allowed to assert.

  • Claims are marked, not just caveated. A sentence describing a page nobody read carries “(unverified — we could not read this page)” right there in the sentence. A disclaimer three lines below an assertion does not retract the assertion; readers keep the specific sentence and drop the general warning.
  • A search result does not count as reading the page. A search returns a title and a snippet. That grounds “this page exists”, never “this page is better optimised”.
  • Blocked sites are named. Some domains refuse automated readers. That is a permanent boundary on the analysis, not a glitch worth retrying, and the answer says which domain and why.
  • Every answer that hit a wall ends with what it could not check — written from what actually ran, so it cannot be quietly omitted.

The restraint is the feature

An assistant that questions everything is an assistant you learn to click past. Four rules keep this layer quiet enough to be worth reading:

  • Silence is the default. A check speaks only when the answer is clearly no, never when it is merely uncertain.
  • It never blocks. The Confirm button that was already there always still means run it anyway. It is your budget and your market.
  • One question, with buttons. Not a questionnaire. A better option you can take in one click, or carry on.
  • It rides the card you were already reading. No extra step, no new screen.

The measure of this layer is not how often it fires. It is whether the times it does fire were worth your attention.

How it compares

ToolQuestions the request firstMarks unverified claimsBlocks you
nqzai Yes — 10 tools, on the approval card, before the spend Yes — ungrounded claims marked inline, gaps listed Never — Confirm always still runs it
Frontier chat assistants n/a — nothing is being spent on your behalf Hedges in prose, when it chooses to Refuses, occasionally
Classic SEO suites Credit counter only — how much, never whether n/a Hard quota stop
Outbound / sales tools Credit counter only n/a Hard quota stop

Comparison reflects nqzai’s own measured behaviour as of August 2026; verify competitor capabilities against their current documentation.

Reports tell you what happened. Ask nqzai why.

A credit counter tells you how much a run costs. It never tells you whether the run is worth making. That is the question a guardrail answers, and it is the only one that changes what you do next.

Every tool“Find me 25 leads”
nqzai“Find me 25 leads and tell me if they actually fit what we sell”
Compares the audience you asked for against what your product brief says you sell, and when they clearly conflict it says so on the approval card — with a better audience you can take in one click.
Every tool“Send these emails”
nqzai“How many of these will actually land in an inbox?”
Subtracts unsubscribed and verification-failed recipients before you approve a number, so the count on the button is the count that sends.
Every tool“Audit my site”
nqzai“Has anything changed since the last crawl?”
Names the date of your last audit and says plainly that an unchanged site returns the same findings — so you skip the spend rather than repeat it.

The report is still produced and saved — it rides along as the evidence for the answer, instead of being handed over in place of one. You can open it, copy it, or push it to your CMS. You just don’t have to read it to find out what changed.

“Find me 25 leads and tell me if they actually fit what we sell”

Questions

What are AI agent guardrails?

Guardrails are the checks that sit between an agent deciding to act and the action actually happening. Most discussion of them is about safety — stopping an agent saying something harmful. The expensive gap in practice is different: an agent that does exactly what it was told, correctly, and produces a result nobody wanted. In nqzai the guardrails are a preflight layer on the tools that spend money, plus a grounding layer on the sentences the answer makes.

What was the point of building this?

One promise: you are never surprised. Not by a result, not by a bill, not by a constraint that got quietly dropped, and not by a confident sentence resting on nothing. Everything in the layer serves that. Before work starts, the tool that is about to spend asks whether it will actually help — and it asks before the approval card, because a clarification arriving after you click Confirm has cost you a decision you already made. While the answer is written, any claim about a page nobody read is marked as unverified.

Does it block me from running things?

No. Not once, anywhere in the layer. Every check adds a line and usually a one-click better option to the card you were already reading; the existing Confirm button always still means "run it anyway". This is deliberate: it is your budget and your market, and an agent that overrules you on either is worse than one that stays quiet.

Is this just another LLM call before every action?

The opposite, and this is the part most people get wrong. Nine of the ten checks call no model at all — they are arithmetic on rows already in hand: how many of these contacts were verified this week, how many days since the last crawl, how many links are already on file. A model is used only where the question is genuinely a judgement call, like whether a requested audience fits the product. A heuristic that fires wrongly argues with you about your own market, and that failure is worse than the one it fixes.

What does a guardrail actually look like when it fires?

A sentence on the approval card, in front of the Confirm button. "Before you confirm — 3 of these 12 will not go out because they unsubscribed or failed verification, so 9 will actually send." Or: "Before you spend — you already crawled this site 4 days ago. Unless the site changed since, this will mostly return the same findings." Short, specific, countable, and attached to the decision it affects.

How is a claim marked as unverified?

The system records which pages were actually read while answering. A sentence that describes a page nobody read is marked inline — "(unverified — we could not read this page)" — and the answer ends with a list of what went unchecked. A web search result does not count as having read the page: a search gives a title and a snippet, which grounds "this page exists", never "this page is better optimised".

Has a guardrail ever been wrong?

One was built and then deleted. A check for contact enrichment turned out to sit behind a threshold it could never reach — the call costs roughly 22,400 tokens against a 75,000-token approval threshold, so the card it would have spoken on never opens, and running the check would have cost more than the spend it protected. Measuring before building is now the rule. It is recorded rather than quietly removed, because "we tried and it was wrong" and "we never looked" are different facts.

Which tools have guardrails today?

Ten: lead search, email sending, contact verification, AI-visibility runs, keyword enrichment, backlink gap analysis, backlink deep scans, and three site audits (on-page crawl, index comparison, and the full Google diagnostic). Every other outward-acting tool is listed in an exemption map with a written reason, because "nobody thought about it" and "we decided no" look identical in an empty registry — a build check fails if a tool appears in neither.

See a guardrail fire on your own account

The fastest one to trigger: ask for a lead list that does not match what you sell, and watch the card answer back before it charges you.

Find me 25 leads and tell me if they actually fit what we sell
Start free → 1 million tokens included on signup. No card.

New here? How nqzai works covers the whole platform, and the research agent covers how answers get their evidence.