Skip to main content

Qualifications

Qualifications let you build a per-workflow test suite: a set of persona-driven chat scenarios that exercise the workflow end to end, plus optional grading criteria that score how each run went.

Use qualifications to iterate on a workflow with confidence — before changing the agent prompt, the routing rules, or a tool definition, you can replay your scenario suite, see which runs still pass, and inspect any regressions in the conversation transcript.

There is no member enrollment, submission, or admin approval. Qualifications are an authoring and testing surface, not a credentialing system.

Key Concepts​

Qualification​

A qualification is a thin container that holds the scenarios for one workflow. Each workflow has at most one qualification, and the qualification is auto-created the first time you open the workflow's Evaluate view — you don't create them manually.

Persona​

A persona is a reusable test identity scoped to your workspace. Personas are not real patients or members — they are templates for the simulated person on the other side of the conversation.

A persona has:

  • Name, description, optional age and date of birth (the date an identity-verification step checks the caller against)
  • Channel capabilities — whether the persona has a phone and/or email
  • SMS notifications — enabled by default when the persona has a phone. Turn it off to simulate an unsubscribe; a phone-sms run requires it to be enabled.
  • Member role — which workspace role the persona maps to when instantiated
  • Identity — the prompt fed to the user-simulator LLM that "plays" this persona during automated runs. Write it in second person ("You are…") with personality traits, goals, and constraints
  • Default data records — pre-populated data the test member starts with, keyed by data-type slug

Personas live under the Personas tab of the Evaluate view. The same persona can be used across many scenarios and many workflows.

Scenario​

A scenario is a specific test under a qualification. It composes a persona with the rest of what the test needs:

  • Name, optional description, slug (unique per qualification)
  • Persona — the reusable identity
  • Goal — what the persona is trying to achieve in this scenario
  • Channel — the channel the conversation runs on (phone, web, etc.)
  • Data records — scenario-specific seed data that overrides the persona's defaults per data-type slug
  • Grading criteria — zero or more pass/fail rules (see below)

Scenarios are ordered within a qualification and can be archived when no longer needed.

Grading Criteria​

Each scenario can have any number of grading criteria. Two types:

  • Expression criteria — pass/fail using a CEL (Common Expression Language) expression evaluated automatically against the run's transcript and assignment data. See the Developer CEL reference for variables and functions in this evaluation context. Synchronous and free.
  • Rubric criteria — LLM-evaluated using a Jinja2 prompt template you write. The grader returns a 0.0–1.0 score, a pass/fail, and brief reasoning. Asynchronous and uses model spend.

Each criterion has a weight (used for the run's aggregate score) and a passing threshold (used by rubric criteria to decide pass/fail from the LLM score).

Test Chat Run​

A test chat run is one execution of an automated test. Every run is anchored to a workflow; scenario linkage is optional:

  • Scenario-bound runs are launched from a scenario and can be evaluated against its criteria.
  • Ad-hoc runs are launched directly from the Test tab with just a persona — no scenario, no grading, just a quick chat to poke at the workflow.

Each run records the chat, the assignment, start/complete/fail timestamps, and (after evaluation) an aggregate score and pass/fail.

A run is driven in one of two modes:

  • Auto — an LLM user-simulator plays the persona end to end.
  • Human — you (the author) type as the persona in the Test tab. Useful when you want to drive the conversation by hand.

The Evaluate View​

Open a workflow and switch to the Evaluate view. The view has four tabs:

  • Test — start a new run, scenario-bound or ad-hoc, and watch the conversation live. Use this for quick iteration while authoring.
  • Personas — workspace-scoped persona CRUD. Create a persona once, reuse it across every workflow.
  • Scenarios — the scenarios for this workflow's qualification. Author goals, seed data, and grading criteria here.
  • Runs — the scenario health board: one card per scenario showing whether it is passing, with details, run history and buttons to start more runs. See The Runs Board.

Setting Up a Test Suite​

The typical authoring loop:

1. Create a Persona​

Go to Personas in the Evaluate view, click Create, and fill in:

  • A name and description
  • Optional age and date of birth (set the date if the Workflow verifies identity by date of birth)
  • Whether the persona has phone/email
  • Leave SMS notifications enabled for the usual phone scenario, or turn it off to test an unsubscribe on voice or web-chat. A phone-sms run with this setting off is rejected.
  • Which member role they map to
  • The identity prompt (second person — "You are…")
  • Optional default data records keyed by data-type slug

Personas are workspace-scoped, so this only happens once per persona.

2. Create a Scenario​

In the Scenarios tab, click Create Scenario and:

  1. Pick the persona you just made (or any existing one)
  2. Write the goal — what is the persona trying to do in this scenario?
  3. Pick the channel
  4. Optionally seed data records that override the persona's defaults
  5. Add grading criteria (see next step)

3. Add Grading Criteria​

For each criterion:

  1. Pick Expression or Rubric
  2. Give it a name and optional description
  3. For expression: write a CEL expression (see CEL Reference)
  4. For rubric: write a Jinja2 prompt template (see Rubric Reference)
  5. Set the weight (for aggregate scoring) and, for rubric criteria, the passing threshold

You can add criteria later — a scenario with no criteria runs fine, you just won't get a score.

4. Run It​

From the Test tab, pick the scenario and start a run. Choose:

  • Auto driver — let the LLM user-simulator play the persona end to end.
  • Human driver — drive the conversation yourself by typing as the persona.

Watch the chat unfold. When the run completes, grade it with Evaluate on the scenario's card in the Runs tab.

To run a scenario without watching it — or to run it several times — start it from the Runs tab instead. Runs started there are graded automatically when they finish.

The Runs Board​

The Runs tab shows one card per scenario, so you can see at a glance whether the whole suite is green and, if not, what needs attention. The headline at the top says how many scenarios need attention — or that there are no results yet, or that every scenario is passing — with a count of the cards in each state.

Reading the cards​

Each card is coloured by the scenario's state, and the cards are sorted worst first:

  • Failing (red) — graded runs exist and none of them passed.
  • Flaky (amber) — some runs passed and some failed.
  • No results — nothing has been graded yet. The card says why: the run is still going, it is waiting to be graded, it couldn't be graded, or the scenario hasn't been run.
  • Passing (green) — every graded run passed, or the last 3 graded runs passed even though an older one failed.

A card counts only the runs on the workflow revision shown at the top of the board, and only the runs started since the scenario's test last changed. Editing the scenario or its grading criteria starts the count again, so a failure against an older version of the test never keeps a card red or amber. The board shows the revision you have been testing by default; use the Revision picker under the headline to look at another one, and Back to current revision to return.

The small chart on each card shows the scenario's pass rate on each revision it was run on, oldest to newest, so you can see whether it is getting better or worse.

Opening a card​

Select a card to open its details right below it. The details show the check that failed most often, a score for each recent run, which checks passed or failed on each run, why any runs couldn't be graded, and the run history. From the history you can View a run's conversation, or Evaluate a finished run that hasn't been graded yet. Once a run has a score the Evaluate button is gone — the board does not offer grading a run a second time; start a new run instead. Press Escape or Close to put the details away.

Starting runs from the board​

  • Run a scenario. An open card has a Run once button. Its arrow opens a menu to run it 3 times or 5 times instead; picking one starts those runs straight away. Each time you open a card, the button goes back to Run once.
  • Run all scenarios. The Run all scenarios button, next to the revision picker, starts one run of every scenario. It tells you how many runs it will start and asks you to confirm first. Scenarios that can't be started — for example, one with no persona — are skipped, and it tells you how many and why. If none can be started, the button is unavailable and the reasons are shown beside it.

Runs started from the board run on the revision the board is showing, and each one is graded automatically when it finishes — you can close the page and the grades still arrive. After you start runs, the board tells you how many started, and lists any the server refused with its reason. If it couldn't confirm a start, it says so — check the board before starting more, because that run may already be going.

Evaluation Pipeline​

When a completed run is graded — automatically for runs started from the Runs board, or when you click Evaluate:

  1. Expression criteria run first. They are deterministic and free.
  2. If any expression criterion fails, rubric criteria are skipped. Rubric calls cost LLM money — there's no point asking the grader for nuance when a hard rule already failed.
  3. Rubric criteria run otherwise. Each one is rendered against the same context as expressions, sent to the grader, and the grader returns a score, pass/fail, and reasoning.
  4. Aggregate score and pass/fail are written back to the run. overall_score is the weighted mean of all scored criteria; passed is true only when every criterion passed.

Each criterion produces one CriterionResult row. Evaluated criteria have a score and pass/fail; rubric grades also include reasoning. A rubric skipped after an expression failure has no score or pass/fail value, and its reasoning explains the intentional skip. Skipped rubrics are excluded from the weighted score; the failed expression still makes the run fail.

Only scenario-bound runs are evaluatable. Ad-hoc runs have no criteria and can't be graded.

Writing Expression Criteria (CEL)​

Expression criteria are CEL expressions that evaluate to true (pass) or false (fail). They run against a single context object built from the completed run.

Available Context Variables​

VariableTypeDescription
assignmentobject{status, outcome, is_test} — the assignment that the run produced
assignment_data_recordsobject{count, slugs, data} — data records linked to the assignment. count and slugs cover every linked record; data is keyed by form slug and holds that form's saved field values
chatobject{channel, message_count, duration_seconds} — chat-level summary
assignment_tasksobject{visited: [str], visited_count: int} — which tasks (steps) were visited
messageslistFlattened transcript — see below
tool_resultslist{tool_name, ok, permission_blocked} per tool result (PHI-safe summary)
configuration_blockerslistHuman-readable workspace setup issues detected at evaluation time (e.g. permission-blocked data tools)
scenarioobject{slug, name, goal, channel}
testChatRunobject{driverMode: "auto"|"human", driverMemberId: int|null}

The messages list​

The most powerful variable for tool-call and ordering checks. Each entry has:

FieldTypeDescription
indexintPosition in the conversation (0-based)
rolestring"user", "assistant", "workflow" (an action the engine performed, not the assistant), or "tool_result"
contentstringText content
tool_callslistTool calls made in this message; each has {name}
tool_namestringFor tool_result entries, the name of the tool that produced this result

All fields are always present (empty string or empty list when not applicable), so you can access any field without null checks.

Example Expressions​

Outcome and shape checks​

assignment.outcome == "success"
chat.message_count >= 5
chat.duration_seconds <= 300
assignment_tasks.visited.exists(t, t == "Confirmation")

Persisted form data (preferred over tool-call checks)​

Grade whether a form row was actually saved on the assignment — not whether the model attempted a create/update tool (which can fail when the tool is permission-blocked):

assignment_data_records.slugs.exists(s, s == "gad7_assessment")
assignment_data_records.count == 0

Grade the values the assistant saved, not just that a row exists. data is keyed by form slug, and works on every workflow (record.data is the strict-routing session form and stays empty otherwise):

assignment_data_records.data.gad7_assessment.followup_preference == "call"

data holds only the forms the run actually wrote, and CEL treats a missing key as an error rather than an empty value — so a criterion that must tolerate the form being absent needs a has() guard, which is false at any depth:

has(assignment_data_records.data.gad7_assessment.followup_preference) &&
assignment_data_records.data.gad7_assessment.followup_preference == "call"

Without the guard, a run that never created the form scores 0 and is recorded as an evaluation error. A typo is different: an expression that does not compile stops grading for that run as a configuration problem, with no score, so fix the criterion and start a new run.

When configuration_blockers is non-empty on a failed run, read those messages (and failure_reason on the run) before blaming the model — they name a form tool the workflow's Form Access ability expects but that was not registered for the run, usually because the ability's form reference does not resolve in this workspace or does not apply to the task under test.

Tool call existence​

The agent must have called forward_call at some point:

messages.exists(m, m.tool_calls.exists(tc, tc.name == "forward_call"))

Ordering — tool A before tool B​

Identity must be verified before medical info is shared:

messages.exists(a, a.tool_calls.exists(tc, tc.name == "verify_identity") &&
messages.exists(b, b.tool_calls.exists(tc, tc.name == "share_medical_info") &&
a.index < b.index))

Negative assertions​

The agent must never escalate:

!messages.exists(m, m.tool_calls.exists(tc, tc.name == "escalate_to_human"))

The agent must never use a forbidden word (case-insensitive):

!messages.exists(m, m.role == "assistant" && m.content.matches("(?i).*diagnosis.*"))

Tool result inspection​

A specific tool returned a confirmed result:

messages.exists(m, m.role == "tool_result" && m.tool_name == "reschedule" &&
m.content.contains("confirmed"))

Counting​

No more than 3 search calls per conversation:

size(messages.filter(m, m.tool_calls.exists(tc, tc.name == "search"))) <= 3

Combined real-world example​

The agent must look up the appointment, then reschedule it, confirm the change to the caller, and never escalate:

messages.exists(m, m.tool_calls.exists(tc, tc.name == "lookup_appointment")) &&
messages.exists(a, a.tool_calls.exists(tc, tc.name == "lookup_appointment") &&
messages.exists(b, b.tool_calls.exists(tc, tc.name == "reschedule") &&
a.index < b.index)) &&
messages.exists(m, m.role == "assistant" && m.content.contains("rescheduled")) &&
!messages.exists(m, m.tool_calls.exists(tc, tc.name == "escalate_to_human"))

CEL Quick Reference​

FunctionDescription
size(list) / size(string)Length of a list or string
string.contains("substr")True if the string contains the substring
string.matches("regex")True if the string matches the regex; (?i) for case-insensitive. Write a word boundary as \\b, not \b — a single \b is read as a backspace character and the pattern silently stops matching. \d, \s and \w need no doubling
list.exists(x, condition)True if any element satisfies the condition
list.all(x, condition)True if every element satisfies the condition
list.filter(x, condition)New list of elements that satisfy the condition
"value" in listTrue if the value is contained in the list

Writing Rubric Criteria (Jinja2)​

Rubric criteria use a Jinja2 template that is rendered with the same context as expressions, then sent to the LLM grader. The grader returns a score (0.0–1.0), a passed boolean, and short reasoning.

A rubric criterion passes when the grader's score is greater than or equal to the passing threshold you configure on the criterion (default 1.0).

Use rubric criteria for judgments that are too nuanced for an expression — empathy, tone, clinical clarity, adherence to a multi-step protocol.

Available Template Variables​

The rubric prompt is rendered with the same context as expressions, plus a pretty transcript for easy iteration in templates:

VariableDescription
transcriptList of message dicts — {role, content, tool_calls?, tool_name?}
assignment.statusThe assignment's final status
assignment.outcomeThe assignment's outcome (e.g. "success", "failure")
chat{channel, message_count, duration_seconds}
scenario{slug, name, goal, channel}
testChatRun{driverMode, driverMemberId}

The chat transcript is also automatically prepended to the prompt sent to the grader, so you don't have to embed it manually unless you want a custom format.

Example Rubric Prompts​

Empathy and clarity​

Evaluate whether the agent:
1. Used empathetic language throughout the conversation
2. Provided clear next steps before ending
3. Avoided medical jargon when speaking with the caller

Return a score from 0.0 to 1.0.

Protocol adherence​

The agent should have:
1. Verified the caller's identity before sharing any medical information
2. Asked about current symptoms before booking
3. Confirmed the appointment time at the end of the call

Score how well the agent followed this protocol from 0.0 to 1.0.

Custom transcript formatting​

If you need to render the transcript yourself (e.g. to filter out tool calls):

Patient interaction:

{% for msg in transcript %}
{% if msg.role in ["user", "assistant"] %}
[{{ msg.role }}]: {{ msg.content }}
{% endif %}
{% endfor %}

Did the agent stay on-topic for the patient's stated goal: "{{ scenario.goal }}"?

Permissions​

Qualifications, scenarios, personas, and test chat runs all use the standard members scopes:

ActionRequired Scope
View qualifications, scenarios, personas, runsmembers:read
Create/edit qualifications, scenarios, personas; start runs; trigger evaluationmembers:write
Start runs of anonymous-caller scenarios (including from Run all scenarios)members:write and chats:admin

Without chats:admin, Run all scenarios skips anonymous-caller scenarios and tells you how many it skipped, and those scenarios' cards explain why there is no Run button.

Tips​

  • Start with one expression criterion. assignment.outcome == "success" catches most regressions. Add more as you find scenarios that pass for the wrong reasons.
  • Use messages.exists(...) for tool-call checks. It gives you ordering context and per-call argument inspection that aggregate variables can't.
  • Use (?i) in matches() for case-insensitive content checks.
  • Test your CEL expressions against a real run before relying on them. The Runs tab shows whether each criterion passed or failed.
  • Reach for rubric criteria sparingly. They cost LLM spend on every evaluation, and the grader is helpful but not infallible. Expressions are free and deterministic — prefer them when the judgment is mechanical.
  • Keep personas reusable. A persona shouldn't bake in a specific goal — that belongs on the scenario. The same "67-year-old caller refilling a prescription" persona can power happy-path, suspicious-caller, and angry-caller scenarios.

Test Data Safety​

Use synthetic identities and fixtures for test runs. Prompts and fixture values can still contain patient data if someone enters it; a test marker does not certify that the contents are synthetic.

  • Runs can create a fresh test Member, reuse an existing test Member, or use an ephemeral anonymous Member.
  • The transcript builder reads from test Assignments only.
  • Persona templates and captured launch inputs retain workspace access and data-protection controls. They are not automatically shared in bundles.
  • Live Operator Mode — qualifications no longer gate routing; "who can answer" is configured on the workflow itself.
  • CEL Expressions — write expression criteria for automated evaluation.
  • Workflows — the workflows you build qualifications for.