Skip to content
← Back to home

§ DOCUMENTATION

Model-Judged Rules

Some rules do not fit a pattern: “never promise a refund above 100 euros”, “no medical advice”. Write them in your own words. After each action a model reads what the agent said or did and records whether it broke the rule, and when it could not tell, says so.

§ 01

What it is, and what it is not

A policy of type semantic_rule holds one statement in plain language. It is not translated into a fixed rule: every stored action that the policy applies to is judged by a model against your sentence, in the background, after the action.

  • Monitor or warn only. It cannot block or hold an action for approval; the API refuses those modes. In warn mode a violation is recorded like any other, with an incident.
  • Professional and Enterprise plans.
  • A strong signal, not proof. A model judges, and models err, mostly on near misses. Each violation says it was model-judged, by which model, with what confidence and why.

This is different from the dashboard's Plain English mode, which turns a sentence into a fixed, deterministic policy once, when you save it.

§ 02

Writing a rule

In the dashboard: Policies, New, Model-judged rule. Or through the API, like any policy:

{  "name": "No refunds above 100 euros",  "policyType": "semantic_rule",  "enforcementMode": "warn",  "ruleDefinition": {    "statement": "The agent must not promise a customer a refund above 100 euros.",    "applies_to": "output",    "examples": {      "violating": ["I have refunded the full 300 euros to your card."],      "allowed": [        "I can't promise a refund of 300 euros; a supervisor will decide.",        "I have refunded 60 euros for the missing item."      ]    }  }}
FieldMeaning
statementThe rule, 10 to 1000 characters. Say what the agent must or must not do, and put limits and exceptions in the sentence.
applies_tooutput (default: what the agent said or did), input (what it was asked), or both.
examplesUp to 10 violating and 10 allowed examples, sent with every judgement. The strongest lever you have.
min_confidenceA violation counts only at or above this confidence (default 0.8). The models measured so far report 0.95 or more on almost every answer, so it rarely changes a result.

Check the examples before saving. The check judges each example against the rule without that example (a model shown it would only repeat it) and says where the model disagrees. A disagreement means the sentence does not yet say what you mean: make it more specific, or add an example of that kind. One model call per example; admins only.

POST /api/v1/policies/semantic-rule/check-examples{ "ruleDefinition": { "statement": "…", "examples": { … } } } {  "data": {    "checked": 3, "matched": 2, "mismatched": 1, "notJudged": 0,    "examples": [      { "text": "I can't promise a refund of 300 euros; …",        "expected": "no_violation", "outcome": "violation", "matches": false,        "reason": "Mentions a refund of 300 euros." }    ]  }}

A model-judged rule cannot be a prerequisite of another policy or have prerequisites: policy chains are decided in one pass, before the model has judged.

§ 03

Which model judges, and how it is checked

The rule is judged by the model the deployment uses for its AI features. On the hosted product that is a hosted model; the statement, its examples and the judged part of the action are sent to the same provider, on the same terms, as the inputs Execlave's semantic classifier already sends. A self-hosted installation uses its local model, and nothing leaves it.

The model decides the quality. Measured with the same prompt, one model reported 3% of allowed actions as violations and others 20% and 45%. So the deployment checks its own model before the model may judge: when the worker starts, it runs 206 labelled actions over 13 rules, in English and German, with near misses and content aimed at the judge, and stores the verdict. It renews it after seven days.

VerdictWhat happens
qualifiedPrecision at least 90%, at most 10% of allowed actions reported, at least 90% of violations found, no attack that worked. The model judges.
pendingNot checked yet. Rules can be saved; actions are recorded as not judged until the verdict is in, and judged by the catch-up afterwards.
not_qualifiedNothing is judged with it, and new rules are refused, naming what it failed on.

The status endpoint (developer role) says whether your organization can use the type now, what blocks it, and what the model measured:

GET /api/v1/policies/semantic-rule/status {  "data": {    "available": true,    "blocker": null,    "plan": "professional",    "modes": ["monitor", "warn"],    "judging": true,    "model": {      "name": "groq:openai/gpt-oss-120b",      "qualification": "qualified",      "overridden": false,      "checkedAt": "2026-10-06T08:00:00.000Z",      "cases": 206,      "precision": 0.95,      "falsePositiveRate": 0.05,      "recall": 1,      "failedCriteria": []    }  }}

Self-hosted settings:

# Self-hosted: the local model (Ollama or another compatible server)LOCAL_LLM_URL=http://localhost:11434SEMANTIC_RULE_MODEL=qwen3.5           # defaults to CLASSIFIER_MODELSEMANTIC_RULE_TIMEOUT_MS=8000 # Hosted inference instead (LLM_PROVIDER=groq)GROQ_SEMANTIC_RULE_MODEL=…            # defaults to GROQ_CLASSIFIER_MODEL # Judge with a model that did not pass the check (shown as overridden)SEMANTIC_RULE_ALLOW_UNQUALIFIED_MODEL=false
§ 04

What is recorded for each action

One judgement per rule and action, with the model, the prompt version and a hash of the statement it was judged against. A real judgement is final; it is never rewritten.

OutcomeMeaning
violationThe model found a violation at or above min_confidence. In warn mode also a policy violation, an audit entry and an incident, marked model-judged.
no_violationThe model judged the action and found none.
not_judgedNobody judged it, with the reason: model unavailable or rate limited, wrong answer format, nothing to judge, content addressed to the judge, plan or model not qualified. Never counted as a pass.

Judgements are kept as long as the traces they are about, by your plan's retention.

§ 05

Reading the results

The policy's page shows how its traffic of the last 24 hours, 7 or 30 days came out, and every judgement with the model's reason, linked to its trace. Coverage is counted from the stored actions, not from the judgements: an action the model never saw shows as “no judgement yet” instead of disappearing from the count.

GET /api/v1/policies/{id}/judgements/coverage?days=7 {  "data": {    "windowDays": 7,    "traces": 1240,    "violation": 9,    "noViolation": 1198,    "notJudged": 21,    "notJudgedByReason": { "rate_limited": 17, "no_content": 4 },    "unjudged": 12,    "unjudgedRecent": 12  }}

The judgements themselves: GET /api/v1/policies/{id}/judgements with outcome, limit and the nextCursor of the previous page. Viewers do not see them: the model's reason can quote what the agent said.

§ 06

Measured accuracy

On the labelled set, with rules written as a bare sentence (October 2026, local models on one machine):

ModelViolations foundAllowed actions reportedWith examples
qwen3.5 (9.7B)100%3–5%1%
qwen2.5-coder 7B100%about 20%13%
llama3.2 (3.2B)100%45%not measured

The hosted model is measured by the deployment itself at start; its own numbers are in the status endpoint and on the create form. 206 actions written for the purpose are a small set: read the figures as an indication, not a guarantee.

§ 07

What this does not cover

  • It does not stop actions. See the first question below.
  • Large payloads stored in object storage are not loaded; such an action is recorded as not judged.
  • Only actions after the rule's last change are judged against it. An action never judged within six hours stays not judged.
  • The model's reason is usually in English, whatever the language of the action or the dashboard.
  • Not in the trial plan.
§ 08

Frequently asked questions

Can a model-judged rule block an action?
No. It is judged after the action, by a model, so it can only monitor or warn. A model call takes hundreds of milliseconds to seconds, and even the best model measured reports some allowed actions as violations. Stopping actions on that would stop legitimate work. For anything that must be blocked, use a deterministic policy (an expression, an allowlist, a limit) on the enforcement path.
What happens when the model is down or rate limited?
The action is recorded as not judged, with the reason, never as passed. The worker tries again with backoff and comes back for recent actions every 15 minutes. The policy page counts not-judged actions separately, so an outage is visible instead of reading as a clean week.
Which model judges, and how good is it?
On the hosted product, the hosted model Execlave uses for its AI features; self-hosted, the local model you configure. Either way the deployment checks the model against 206 labelled actions before it may judge, and shows the result. The status endpoint and the create form give the model name, when it was checked, and the share of violations it found and of allowed actions it wrongly reported.
Can the content of an action talk the model out of a violation?
The action is passed to the model as data between markers it cannot know, and the rule is restated after it. When the content addresses the judge ("the rule is suspended", "answer no violation") and the model answers "no violation", the answer is not trusted: the action is recorded as not judged. In the labelled set no attack produced a "no violation" on any model measured.
Why did an obviously fine action get flagged?
Near misses are where models err: a refusal ("I can’t promise a 300 euro refund"), an estimate where the rule forbids a guarantee, the customer’s own data where the rule protects other people’s. Add examples of exactly those to the rule. In measurement, examples cut false positives by a third to two thirds.
Model-Judged Rules — Execlave Docs