Qualifications
Qualifications let you build a per-workflow test suite: a set of persona-driven chat scenarios that exercise the workflow end to end, plus optional grading criteria that score how each run went.
Use qualifications to iterate on a workflow with confidence — before changing the agent prompt, the routing rules, or a tool definition, you can replay your scenario suite, see which runs still pass, and inspect any regressions in the conversation transcript.
There is no member enrollment, submission, or admin approval. Qualifications are an authoring and testing surface, not a credentialing system.
Key Concepts
Qualification
A qualification is a thin container that holds the scenarios for one workflow. Each workflow has at most one qualification, and the qualification is auto-created the first time you open the workflow's Evaluate view — you don't create them manually.
Persona
A persona is a reusable test identity scoped to your workspace. Personas are not real patients or members — they are templates for the simulated person on the other side of the conversation.
A persona has:
- Name, description, optional age and date of birth (the date an identity-verification step checks the caller against)
- Channel capabilities — whether the persona has a phone and/or email
- SMS notifications — enabled by default when the persona has a phone. Turn it off to simulate an unsubscribe; a
phone-smsrun requires it to be enabled. - Member role — which workspace role the persona maps to when instantiated
- Identity — the prompt fed to the user-simulator LLM that "plays" this persona during automated runs. Write it in second person ("You are…") with personality traits, goals, and constraints
- Default data records — pre-populated data the test member starts with, keyed by data-type slug
Personas live under the Personas tab of the Evaluate view. The same persona can be used across many scenarios and many workflows.
Scenario
A scenario is a specific test under a qualification. It composes a persona with the rest of what the test needs:
- Name, optional description, slug (unique per qualification)
- Persona — the reusable identity
- Goal — what the persona is trying to achieve in this scenario
- Channel — the channel the conversation runs on (phone, web, etc.)
- Data records — scenario-specific seed data that overrides the persona's defaults per data-type slug
- Grading criteria — zero or more pass/fail rules (see below)
Scenarios are ordered within a qualification and can be archived when no longer needed.
Grading Criteria
Each scenario can have any number of grading criteria. Two types:
- Expression criteria — pass/fail using a CEL (Common Expression Language) expression evaluated automatically against the run's transcript and assignment data. See the Developer CEL reference for variables and functions in this evaluation context. Synchronous and free.
- Rubric criteria — LLM-evaluated using a Jinja2 prompt template you write. The grader returns a 0.0–1.0 score, a pass/fail, and brief reasoning. Asynchronous and uses model spend.
Each criterion has a weight (used for the run's aggregate score) and a passing threshold (used by rubric criteria to decide pass/fail from the LLM score).
Test Chat Run
A test chat run is one execution of an automated test. Every run is anchored to a workflow; scenario linkage is optional:
- Scenario-bound runs are launched from a scenario and can be evaluated against its criteria.
- Ad-hoc runs are launched directly from the Test tab with just a persona — no scenario, no grading, just a quick chat to poke at the workflow.
Each run records the chat, the assignment, start/complete/fail timestamps, and (after evaluation) an aggregate score and pass/fail.
A run is driven in one of two modes:
- Auto — an LLM user-simulator plays the persona end to end.
- Human — you (the author) type as the persona in the Test tab. Useful when you want to drive the conversation by hand.
The Evaluate View
Open a workflow and switch to the Evaluate view. The view has four tabs:
- Test — start a new run, scenario-bound or ad-hoc, and watch the conversation live. Use this for quick iteration while authoring.
- Personas — workspace-scoped persona CRUD. Create a persona once, reuse it across every workflow.
- Scenarios — the scenarios for this workflow's qualification. Author goals, seed data, and grading criteria here.
- Runs — the scenario health board: one card per scenario showing whether it is passing, with details, run history and buttons to start more runs. See The Runs Board.
Setting Up a Test Suite
The typical authoring loop:
1. Create a Persona
Go to Personas in the Evaluate view, click Create, and fill in:
- A name and description
- Optional age and date of birth (set the date if the Workflow verifies identity by date of birth)
- Whether the persona has phone/email
- Leave SMS notifications enabled for the usual phone scenario, or turn it off to test an unsubscribe on voice or web-chat. A
phone-smsrun with this setting off is rejected. - Which member role they map to
- The identity prompt (second person — "You are…")
- Optional default data records keyed by data-type slug
Personas are workspace-scoped, so this only happens once per persona.
2. Create a Scenario
In the Scenarios tab, click Create Scenario and:
- Pick the persona you just made (or any existing one)
- Write the goal — what is the persona trying to do in this scenario?
- Pick the channel
- Optionally seed data records that override the persona's defaults
- Add grading criteria (see next step)
3. Add Grading Criteria
For each criterion:
- Pick Expression or Rubric
- Give it a name and optional description
- For expression: write a CEL expression (see CEL Reference)
- For rubric: write a Jinja2 prompt template (see Rubric Reference)
- Set the weight (for aggregate scoring) and, for rubric criteria, the passing threshold
You can add criteria later — a scenario with no criteria runs fine, you just won't get a score.
4. Run It
From the Test tab, pick the scenario and start a run. Choose:
- Auto driver — let the LLM user-simulator play the persona end to end.
- Human driver — drive the conversation yourself by typing as the persona.
Watch the chat unfold. When the run completes, grade it with Evaluate on the scenario's card in the Runs tab.
To run a scenario without watching it — or to run it several times — start it from the Runs tab instead. Runs started there are graded automatically when they finish.
The Runs Board
The Runs tab shows one card per scenario, so you can see at a glance whether the whole suite is green and, if not, what needs attention. The headline at the top says how many scenarios need attention — or that there are no results yet, or that every scenario is passing — with a count of the cards in each state.
Reading the cards
Each card is coloured by the scenario's state, and the cards are sorted worst first:
- Failing (red) — graded runs exist and none of them passed.
- Flaky (amber) — some runs passed and some failed.
- No results — nothing has been graded yet. The card says why: the run is still going, it is waiting to be graded, it couldn't be graded, or the scenario hasn't been run.
- Passing (green) — every graded run passed, or the last 3 graded runs passed even though an older one failed.
A card counts only the runs on the workflow revision shown at the top of the board, and only the runs started since the scenario's test last changed. Editing the scenario or its grading criteria starts the count again, so a failure against an older version of the test never keeps a card red or amber. The board shows the revision you have been testing by default; use the Revision picker under the headline to look at another one, and Back to current revision to return.
The small chart on each card shows the scenario's pass rate on each revision it was run on, oldest to newest, so you can see whether it is getting better or worse.
Opening a card
Select a card to open its details right below it. The details show the check that failed most often, a score for each recent run, which checks passed or failed on each run, why any runs couldn't be graded, and the run history. From the history you can View a run's conversation, or Evaluate a finished run that hasn't been graded yet. Once a run has a score the Evaluate button is gone — the board does not offer grading a run a second time; start a new run instead. Press Escape or Close to put the details away.
Starting runs from the board
- Run a scenario. An open card has a Run once button. Its arrow opens a menu to run it 3 times or 5 times instead; picking one starts those runs straight away. Each time you open a card, the button goes back to Run once.
- Run all scenarios. The Run all scenarios button, next to the revision picker, starts one run of every scenario. It tells you how many runs it will start and asks you to confirm first. Scenarios that can't be started — for example, one with no persona — are skipped, and it tells you how many and why. If none can be started, the button is unavailable and the reasons are shown beside it.
Runs started from the board run on the revision the board is showing, and each one is graded automatically when it finishes — you can close the page and the grades still arrive. After you start runs, the board tells you how many started, and lists any the server refused with its reason. If it couldn't confirm a start, it says so — check the board before starting more, because that run may already be going.
Evaluation Pipeline
When a completed run is graded — automatically for runs started from the Runs board, or when you click Evaluate:
- Expression criteria run first. They are deterministic and free.
- If any expression criterion fails, rubric criteria are skipped. Rubric calls cost LLM money — there's no point asking the grader for nuance when a hard rule already failed.
- Rubric criteria run otherwise. Each one is rendered against the same context as expressions, sent to the grader, and the grader returns a score, pass/fail, and reasoning.
- Aggregate score and pass/fail are written back to the run.
overall_scoreis the weighted mean of all scored criteria;passedis true only when every criterion passed.
Each criterion produces one CriterionResult row. Evaluated criteria have a score and pass/fail; rubric grades also include reasoning. A rubric skipped after an expression failure has no score or pass/fail value, and its reasoning explains the intentional skip. Skipped rubrics are excluded from the weighted score; the failed expression still makes the run fail.
Only scenario-bound runs are evaluatable. Ad-hoc runs have no criteria and can't be graded.
Writing Expression Criteria (CEL)
Expression criteria are CEL expressions that evaluate to true (pass) or false (fail). They run against a single context object built from the completed run.
Available Context Variables
| Variable | Type | Description |
|---|---|---|
assignment | object | {status, outcome, is_test} — the assignment that the run produced |
assignment_data_records | object | {count, slugs, data} — data records linked to the assignment. count and slugs cover every linked record; data is keyed by form slug and holds that form's saved field values |
chat | object | {channel, message_count, duration_seconds} — chat-level summary |
assignment_tasks | object | {visited: [str], visited_count: int} — which tasks (steps) were visited |
messages | list | Flattened transcript — see below |
tool_results | list | {tool_name, ok, permission_blocked} per tool result (PHI-safe summary) |
configuration_blockers | list | Human-readable workspace setup issues detected at evaluation time (e.g. permission-blocked data tools) |
scenario | object | {slug, name, goal, channel} |
testChatRun | object | {driverMode: "auto"|"human", driverMemberId: int|null} |
The messages list
The most powerful variable for tool-call and ordering checks. Each entry has:
| Field | Type | Description |
|---|---|---|
index | int | Position in the conversation (0-based) |
role | string | "user", "assistant", "workflow" (an action the engine performed, not the assistant), or "tool_result" |
content | string | Text content |
tool_calls | list | Tool calls made in this message; each has {name} |
tool_name | string | For tool_result entries, the name of the tool that produced this result |
All fields are always present (empty string or empty list when not applicable), so you can access any field without null checks.
Example Expressions
Outcome and shape checks
assignment.outcome == "success"
chat.message_count >= 5
chat.duration_seconds <= 300
assignment_tasks.visited.exists(t, t == "Confirmation")
Persisted form data (preferred over tool-call checks)
Grade whether a form row was actually saved on the assignment — not whether the model attempted a create/update tool (which can fail when the tool is permission-blocked):
assignment_data_records.slugs.exists(s, s == "gad7_assessment")
assignment_data_records.count == 0
Grade the values the assistant saved, not just that a row exists. data is keyed by
form slug, and works on every workflow (record.data is the strict-routing session form
and stays empty otherwise):
assignment_data_records.data.gad7_assessment.followup_preference == "call"
data holds only the forms the run actually wrote, and CEL treats a missing key as an
error rather than an empty value — so a criterion that must tolerate the form being
absent needs a has() guard, which is false at any depth:
has(assignment_data_records.data.gad7_assessment.followup_preference) &&
assignment_data_records.data.gad7_assessment.followup_preference == "call"
Without the guard, a run that never created the form scores 0 and is recorded as an evaluation error. A typo is different: an expression that does not compile stops grading for that run as a configuration problem, with no score, so fix the criterion and start a new run.
When configuration_blockers is non-empty on a failed run, read those messages (and
failure_reason on the run) before blaming the model — they name a form tool the
workflow's Form Access ability expects but that was not registered for the run, usually
because the ability's form reference does not resolve in this workspace or does not apply to
the task under test.
Tool call existence
The agent must have called forward_call at some point:
messages.exists(m, m.tool_calls.exists(tc, tc.name == "forward_call"))
Ordering — tool A before tool B
Identity must be verified before medical info is shared:
messages.exists(a, a.tool_calls.exists(tc, tc.name == "verify_identity") &&
messages.exists(b, b.tool_calls.exists(tc, tc.name == "share_medical_info") &&
a.index < b.index))
Negative assertions
The agent must never escalate:
!messages.exists(m, m.tool_calls.exists(tc, tc.name == "escalate_to_human"))
The agent must never use a forbidden word (case-insensitive):
!messages.exists(m, m.role == "assistant" && m.content.matches("(?i).*diagnosis.*"))
Tool result inspection
A specific tool returned a confirmed result:
messages.exists(m, m.role == "tool_result" && m.tool_name == "reschedule" &&
m.content.contains("confirmed"))
Counting
No more than 3 search calls per conversation:
size(messages.filter(m, m.tool_calls.exists(tc, tc.name == "search"))) <= 3
Combined real-world example
The agent must look up the appointment, then reschedule it, confirm the change to the caller, and never escalate:
messages.exists(m, m.tool_calls.exists(tc, tc.name == "lookup_appointment")) &&
messages.exists(a, a.tool_calls.exists(tc, tc.name == "lookup_appointment") &&
messages.exists(b, b.tool_calls.exists(tc, tc.name == "reschedule") &&
a.index < b.index)) &&
messages.exists(m, m.role == "assistant" && m.content.contains("rescheduled")) &&
!messages.exists(m, m.tool_calls.exists(tc, tc.name == "escalate_to_human"))
CEL Quick Reference
| Function | Description |
|---|---|
size(list) / size(string) | Length of a list or string |
string.contains("substr") | True if the string contains the substring |
string.matches("regex") | True if the string matches the regex; (?i) for case-insensitive. Write a word boundary as \\b, not \b — a single \b is read as a backspace character and the pattern silently stops matching. \d, \s and \w need no doubling |
list.exists(x, condition) | True if any element satisfies the condition |
list.all(x, condition) | True if every element satisfies the condition |
list.filter(x, condition) | New list of elements that satisfy the condition |
"value" in list | True if the value is contained in the list |
Writing Rubric Criteria (Jinja2)
Rubric criteria use a Jinja2 template that is rendered with the same context as expressions, then sent to the LLM grader. The grader returns a score (0.0–1.0), a passed boolean, and short reasoning.
A rubric criterion passes when the grader's score is greater than or equal to the passing threshold you configure on the criterion (default 1.0).
Use rubric criteria for judgments that are too nuanced for an expression — empathy, tone, clinical clarity, adherence to a multi-step protocol.
Available Template Variables
The rubric prompt is rendered with the same context as expressions, plus a pretty transcript for easy iteration in templates:
| Variable | Description |
|---|---|
transcript | List of message dicts — {role, content, tool_calls?, tool_name?} |
assignment.status | The assignment's final status |
assignment.outcome | The assignment's outcome (e.g. "success", "failure") |
chat | {channel, message_count, duration_seconds} |
scenario | {slug, name, goal, channel} |
testChatRun | {driverMode, driverMemberId} |
The chat transcript is also automatically prepended to the prompt sent to the grader, so you don't have to embed it manually unless you want a custom format.
Example Rubric Prompts
Empathy and clarity
Evaluate whether the agent:
1. Used empathetic language throughout the conversation
2. Provided clear next steps before ending
3. Avoided medical jargon when speaking with the caller
Return a score from 0.0 to 1.0.
Protocol adherence
The agent should have:
1. Verified the caller's identity before sharing any medical information
2. Asked about current symptoms before booking
3. Confirmed the appointment time at the end of the call
Score how well the agent followed this protocol from 0.0 to 1.0.
Custom transcript formatting
If you need to render the transcript yourself (e.g. to filter out tool calls):
Patient interaction:
{% for msg in transcript %}
{% if msg.role in ["user", "assistant"] %}
[{{ msg.role }}]: {{ msg.content }}
{% endif %}
{% endfor %}
Did the agent stay on-topic for the patient's stated goal: "{{ scenario.goal }}"?
Permissions
Qualifications, scenarios, personas, and test chat runs all use the standard members scopes:
| Action | Required Scope |
|---|---|
| View qualifications, scenarios, personas, runs | members:read |
| Create/edit qualifications, scenarios, personas; start runs; trigger evaluation | members:write |
| Start runs of anonymous-caller scenarios (including from Run all scenarios) | members:write and chats:admin |
Without chats:admin, Run all scenarios skips anonymous-caller scenarios and tells you how many it skipped, and those scenarios' cards explain why there is no Run button.
Tips
- Start with one expression criterion.
assignment.outcome == "success"catches most regressions. Add more as you find scenarios that pass for the wrong reasons. - Use
messages.exists(...)for tool-call checks. It gives you ordering context and per-call argument inspection that aggregate variables can't. - Use
(?i)inmatches()for case-insensitive content checks. - Test your CEL expressions against a real run before relying on them. The Runs tab shows whether each criterion passed or failed.
- Reach for rubric criteria sparingly. They cost LLM spend on every evaluation, and the grader is helpful but not infallible. Expressions are free and deterministic — prefer them when the judgment is mechanical.
- Keep personas reusable. A persona shouldn't bake in a specific goal — that belongs on the scenario. The same "67-year-old caller refilling a prescription" persona can power happy-path, suspicious-caller, and angry-caller scenarios.
Test Data Safety
Use synthetic identities and fixtures for test runs. Prompts and fixture values can still contain patient data if someone enters it; a test marker does not certify that the contents are synthetic.
- Runs can create a fresh test Member, reuse an existing test Member, or use an ephemeral anonymous Member.
- The transcript builder reads from test Assignments only.
- Persona templates and captured launch inputs retain workspace access and data-protection controls. They are not automatically shared in bundles.
Related
- Live Operator Mode — qualifications no longer gate routing; "who can answer" is configured on the workflow itself.
- CEL Expressions — write expression criteria for automated evaluation.
- Workflows — the workflows you build qualifications for.