§ DOCUMENTATION
Model-Judged Rules
Some rules do not fit a pattern: “never promise a refund above 100 euros”, “no medical advice”. Write them in your own words. After each action a model reads what the agent said or did and records whether it broke the rule, and when it could not tell, says so.
What it is, and what it is not
A policy of type semantic_rule holds one statement in plain language. It is not translated into a fixed rule: every stored action that the policy applies to is judged by a model against your sentence, in the background, after the action.
- Monitor or warn only. It cannot block or hold an action for approval; the API refuses those modes. In warn mode a violation is recorded like any other, with an incident.
- Professional and Enterprise plans.
- A strong signal, not proof. A model judges, and models err, mostly on near misses. Each violation says it was model-judged, by which model, with what confidence and why.
This is different from the dashboard's Plain English mode, which turns a sentence into a fixed, deterministic policy once, when you save it.
Writing a rule
In the dashboard: Policies, New, Model-judged rule. Or through the API, like any policy:
{ "name": "No refunds above 100 euros", "policyType": "semantic_rule", "enforcementMode": "warn", "ruleDefinition": { "statement": "The agent must not promise a customer a refund above 100 euros.", "applies_to": "output", "examples": { "violating": ["I have refunded the full 300 euros to your card."], "allowed": [ "I can't promise a refund of 300 euros; a supervisor will decide.", "I have refunded 60 euros for the missing item." ] } }}| Field | Meaning |
|---|---|
statement | The rule, 10 to 1000 characters. Say what the agent must or must not do, and put limits and exceptions in the sentence. |
applies_to | output (default: what the agent said or did), input (what it was asked), or both. |
examples | Up to 10 violating and 10 allowed examples, sent with every judgement. The strongest lever you have. |
min_confidence | A violation counts only at or above this confidence (default 0.8). The models measured so far report 0.95 or more on almost every answer, so it rarely changes a result. |
Check the examples before saving. The check judges each example against the rule without that example (a model shown it would only repeat it) and says where the model disagrees. A disagreement means the sentence does not yet say what you mean: make it more specific, or add an example of that kind. One model call per example; admins only.
POST /api/v1/policies/semantic-rule/check-examples{ "ruleDefinition": { "statement": "…", "examples": { … } } } { "data": { "checked": 3, "matched": 2, "mismatched": 1, "notJudged": 0, "examples": [ { "text": "I can't promise a refund of 300 euros; …", "expected": "no_violation", "outcome": "violation", "matches": false, "reason": "Mentions a refund of 300 euros." } ] }}A model-judged rule cannot be a prerequisite of another policy or have prerequisites: policy chains are decided in one pass, before the model has judged.
Which model judges, and how it is checked
The rule is judged by the model the deployment uses for its AI features. On the hosted product that is a hosted model; the statement, its examples and the judged part of the action are sent to the same provider, on the same terms, as the inputs Execlave's semantic classifier already sends. A self-hosted installation uses its local model, and nothing leaves it.
The model decides the quality. Measured with the same prompt, one model reported 3% of allowed actions as violations and others 20% and 45%. So the deployment checks its own model before the model may judge: when the worker starts, it runs 206 labelled actions over 13 rules, in English and German, with near misses and content aimed at the judge, and stores the verdict. It renews it after seven days.
| Verdict | What happens |
|---|---|
qualified | Precision at least 90%, at most 10% of allowed actions reported, at least 90% of violations found, no attack that worked. The model judges. |
pending | Not checked yet. Rules can be saved; actions are recorded as not judged until the verdict is in, and judged by the catch-up afterwards. |
not_qualified | Nothing is judged with it, and new rules are refused, naming what it failed on. |
The status endpoint (developer role) says whether your organization can use the type now, what blocks it, and what the model measured:
GET /api/v1/policies/semantic-rule/status { "data": { "available": true, "blocker": null, "plan": "professional", "modes": ["monitor", "warn"], "judging": true, "model": { "name": "groq:openai/gpt-oss-120b", "qualification": "qualified", "overridden": false, "checkedAt": "2026-10-06T08:00:00.000Z", "cases": 206, "precision": 0.95, "falsePositiveRate": 0.05, "recall": 1, "failedCriteria": [] } }}Self-hosted settings:
# Self-hosted: the local model (Ollama or another compatible server)LOCAL_LLM_URL=http://localhost:11434SEMANTIC_RULE_MODEL=qwen3.5 # defaults to CLASSIFIER_MODELSEMANTIC_RULE_TIMEOUT_MS=8000 # Hosted inference instead (LLM_PROVIDER=groq)GROQ_SEMANTIC_RULE_MODEL=… # defaults to GROQ_CLASSIFIER_MODEL # Judge with a model that did not pass the check (shown as overridden)SEMANTIC_RULE_ALLOW_UNQUALIFIED_MODEL=falseWhat is recorded for each action
One judgement per rule and action, with the model, the prompt version and a hash of the statement it was judged against. A real judgement is final; it is never rewritten.
| Outcome | Meaning |
|---|---|
violation | The model found a violation at or above min_confidence. In warn mode also a policy violation, an audit entry and an incident, marked model-judged. |
no_violation | The model judged the action and found none. |
not_judged | Nobody judged it, with the reason: model unavailable or rate limited, wrong answer format, nothing to judge, content addressed to the judge, plan or model not qualified. Never counted as a pass. |
Judgements are kept as long as the traces they are about, by your plan's retention.
Reading the results
The policy's page shows how its traffic of the last 24 hours, 7 or 30 days came out, and every judgement with the model's reason, linked to its trace. Coverage is counted from the stored actions, not from the judgements: an action the model never saw shows as “no judgement yet” instead of disappearing from the count.
GET /api/v1/policies/{id}/judgements/coverage?days=7 { "data": { "windowDays": 7, "traces": 1240, "violation": 9, "noViolation": 1198, "notJudged": 21, "notJudgedByReason": { "rate_limited": 17, "no_content": 4 }, "unjudged": 12, "unjudgedRecent": 12 }}The judgements themselves: GET /api/v1/policies/{id}/judgements with outcome, limit and the nextCursor of the previous page. Viewers do not see them: the model's reason can quote what the agent said.
Measured accuracy
On the labelled set, with rules written as a bare sentence (October 2026, local models on one machine):
| Model | Violations found | Allowed actions reported | With examples |
|---|---|---|---|
| qwen3.5 (9.7B) | 100% | 3–5% | 1% |
| qwen2.5-coder 7B | 100% | about 20% | 13% |
| llama3.2 (3.2B) | 100% | 45% | not measured |
The hosted model is measured by the deployment itself at start; its own numbers are in the status endpoint and on the create form. 206 actions written for the purpose are a small set: read the figures as an indication, not a guarantee.
What this does not cover
- It does not stop actions. See the first question below.
- Large payloads stored in object storage are not loaded; such an action is recorded as not judged.
- Only actions after the rule's last change are judged against it. An action never judged within six hours stays not judged.
- The model's reason is usually in English, whatever the language of the action or the dashboard.
- Not in the trial plan.