Academy · Platform · Agents

Crew & evaluations

In one line. Watch your agents work in real time on the collection’s Crew tab, and prove they still work — against curated ground truth, on a schedule — with Evaluations. You’ll be able to. Read the live crew view during a run, build a golden dataset, run an evaluation against any crew (agent, XFlow or mesh), and use baselines to catch regressions before your users do. Where this lives. The Crew tab is on the collection itself, in the app your users work in — Manage in Studio on that tab jumps to Studio > Data > Collections. Evaluations live at Studio > Agents > Quality > Evaluations.

Crew — the live view

A crew is the platform’s umbrella word for whatever workforce is attached to a collection: a single agent, an XFlow, or a mesh. The Crew tab on the collection workspace merges “what is this collection?” and “who works it, right now?” into one screen:

  1. Identity strip — the collection’s title, a live Ready / Crew working status pill, and one row of vitals: documents, fields, stages, intake. A Manage in Studio link jumps to configuration.
  2. Pipeline rail — the collection’s stages as numbered nodes, tinting while work is in flight.
  3. The live crew — KPI tiles, the crew stage with each member lighting up as it works, recent decisions, and recent runs.

If you don’t see a Crew tab on your collection, your workspace still shows the older Overview and Agents tabs instead — ask your administrator.

This is the screen to keep open during a demo or a go-live: documents land, lanes light left to right, decisions accumulate — and when something stalls, you can see which member stalled instead of inferring it from a silent queue.

The Crew tab also hosts the per-document step-log timeline: click a document row and every ingestion step appears in order — including each retry attempt as its own row, attempt-numbered and carrying the full per-attempt error, so an admin can read “failed 1/3 → retrying → failed after 3” and tell a transient blip from a broken model config. A Knowledge collection’s Index Health View logs link jumps straight here.

Evaluations — the proof

The live view tells you the crew is working. Evaluations tell you it is working correctly — and still working correctly after every prompt tweak, model upgrade, or upstream change.

Studio > Agents > Quality > Evaluations

The page opens on four tabs — Overview, Datasets, Runs, Drift — and one button, + New evaluation, which walks you through five steps: what you’re testing, which dataset, which metrics, which judge model, and when it should run.

Watch out. Evaluations is on by default, but an administrator can switch it off for a workspace. When it’s off you lose three things together: this menu entry, the Dataset purpose when you create a collection, and the Evaluate tab in the flow designer. If you can’t find them, ask your administrator.

The system-under-test is any crew: a single agent, an XFlow, or a mesh. The loop:

  1. Golden dataset. Curate documents whose correct outputs you know — drawn from human-verified annotations and reviewed corrections, not generated by a model. This is the asset that makes every later step honest.
  2. Evaluation run. The crew re-processes the golden documents as if they were new, and the results are scored two ways. Answers that have one right value are scored by exact comparison — how many fields matched, how many documents were classified correctly. Written answers are scored by a second model acting as judge — is this claim supported by the documents, does the answer cover what was asked — and every judge verdict is a plain yes or no rather than a mark out of ten, which keeps the scoring stable run to run. You choose the judge model when you set the evaluation up; pick one different from the model under test, so no agent marks its own homework.
  3. Baseline & regression. A run can be pinned as the baseline; later runs compare against it, so “did last week’s prompt change hurt extraction?” has a measured answer instead of an argument.
  4. Schedule it. Evaluations run on a schedule, which turns them from a release ritual into a canary: quality drops show up in the report, not in a customer escalation.

What each word means here

Word Meaning Caught by
Accuracy Output matches curated ground truth Field and document scoring, classification metrics
Inconsistency Same input produces different outputs run to run Stability scoring over repeated runs
Drift Quality or data distributions move over time on live traffic Scheduled runs and monitoring over live telemetry

The habit that makes this pay

Treat the golden dataset like code: it grows with every incident. When a reviewer catches a wrong extraction, that document — with its corrected values — belongs in the dataset, so the mistake can never come back silently. Twenty well-chosen documents that cover your ugly cases beat two hundred easy ones.

Run an evaluation before every meaningful change ships: persona edits, model swaps, schema changes, new pipeline steps. The comparison against baseline is your regression gate.

Where to go next

Prefer learning inside the product? The same academy lives in the platform's Learn menu — every screen links to the chapter that explains it.

See the platform live