Academy · Platform · Governance

Observability

In one line. Every run leaves a trace — what ran, what it read, what it called, what it cost, what it decided — and knowing where each trace lives turns “the AI did something weird” into a five-minute diagnosis. You’ll be able to. Answer the three operational questions on any live solution: is it working now, what exactly happened on that run, and is quality holding over time. Where this lives. The collection’s Agents tab and each agent’s Runs for the live picture, Studio ▸ Governance ▸ Oversight ▸ Audit & Usage for the numbers, and Account menu ▸ Admin ▸ Governance ▸ Observability for the whole environment.

Is it working right now?

Start where the work is:

  • The collection’s Agents tab (Studio ▸ Data ▸ Collections ▸ <collection> ▸ Agents) — the agents wired to that collection and what they are doing with it. If a demo or go-live has a single pane of glass, this is it.
  • The document list (E2) — lifecycle stages are a queue-depth readout: a growing In review pile or documents stuck in New tells you where the process is jammed.
  • Dashboards (E5) — throughput, exceptions and aging over time, for the people who own the process rather than the plumbing.

What exactly happened on that run?

Every execution is recorded, at the level you need:

Trace Where What it answers
Agent run trace The agent’s Runs (and the playground, live) The full loop: inputs, tool calls, output, tokens in/out, estimated cost, latency
Flow runs Studio ▸ Agents ▸ Agents — filter the list to flows, open one, then its Runs Which step ran, with what parameters, what each produced, where it failed
Multi-agent pipeline trace The collection’s Agents tab Which member is running now, how long each took, and where the run stopped
Document history The document’s detail pane (E3) What was extracted, corrected, decided — and by whom or what
Chat & search audit Studio ▸ Experience ▸ Channels ▸ Search ▸ Audit and Studio ▸ Experience ▸ Channels ▸ Chat ▸ Audit What users asked, what was answered, what was retrieved

The habit that pays: when an output looks wrong, open the run trace first. Nine times out of ten the trace shows the cause directly — a missing attachment, a tool that returned nothing, a low-confidence field that should have routed to review.

The numbers: Audit & Usage

Studio ▸ Governance ▸ Oversight ▸ Audit & Usage is the day-to-day observability page for whoever owns a project. One page, a two-way switch under the title: Audit answers who did what; Usage answers what is being consumed. The switch is part of the address, so either view can be bookmarked or shared.

The Audit side is covered in full by G5 · Audit & compliance. The Usage side is the one you read when someone asks how much the solution is doing, and what it is costing:

  • Period (top right) — Last 7 days, Last 30 days (where it starts) or Last 90 days. Everything below it except the plan counters follows this window.
  • AI usage — four tiles: Requests (with the passed / failed split), Agents (how many were active in the period), Tokens (in and out), and Cost (est.) with average latency underneath.
  • Requests per day — a bar chart of daily volume across the period; hover a bar for that day’s request count and estimated cost. It is only drawn when the period has more than one day of data.
  • By agent — one row per agent: Requests, Failed, Tokens, Cost (est.), Avg latency, Last run. Click a row and that agent opens, so you can go straight from “this one is expensive” to its runs.
  • People activity — reading sessions and active time, by document and by user. Empty until people actually start opening documents.
  • Plan counters — live project totals for Documents, Pages and Annotations. These are running totals, not period figures, so changing Period does not move them.

Cost figures are estimates. Every cost number on this page is derived from tokens and model price and is labelled (est.). The billing authority is the Managed Gateway’s ledger, surfaced through the Billing meters at the foot of the workspace and platform Usage views. Never hand an estimate to finance as an invoice.

Empty is not the same as broken. Each section loads on its own. A section that failed says “Couldn’t load…” — most also offer Retry, though the plan counters just report the failure — while a section that says there was no activity in the period means there genuinely was none.

The same page exists one and two levels up. Workspace settings ▸ Audit & Usage covers every project in the workspace plus sign-ins, and adds a per-project breakdown and the billing meters. Account menu ▸ Admin ▸ Governance ▸ Audit & Usage covers every workspace, with a cross-workspace leaderboard and a workspace picker on the meters. Neither has the Period dropdown — both report on a fixed recent window, printed next to the first section heading. Both are limited to workspace admins and platform admins respectively; if you don’t see the entry, you don’t have that scope.

Is quality holding over time?

Point-in-time traces don’t catch slow drift. Two instruments do:

  • Evaluations (A12) — scheduled scored runs against golden datasets, compared to a pinned baseline. This is your regression gate for instruction, model and pipeline changes.
  • Review-rate trends — the share of documents routing to human review is a quality signal in itself: rising review rates mean confidence is dropping somewhere upstream, and the review queue tells you which fields.

Cost is part of observability

Run traces carry token usage and cost, and the By agent table on Usage ranks them for you. When a solution gets expensive, the two together show which agent, which step, and usually why — an over-attached knowledge tab or a heavier model than the job needs (A1 has the diagnosis table). Measure before switching models; the playground makes the comparison cheap. Environment-wide spend lives in the admin console below.

The platform-level console

Everything above is per-project. Platform admins get a whole-environment console at Account menu ▸ Admin ▸ Governance ▸ Observability — and Administration itself opens on the Action Center, a self-clearing to-do list of setup gaps, incidents, and cost breaches. The pages that matter most:

Page What it answers
Pulse Live platform health at a glance.
Pages · APIs · DB Ops · Jobs · Exceptions Performance and error drill-downs — from a slow route down to the offending query or stack trace.
Cost Model spend across the platform: period totals, a daily-spend trend, and a per-workspace table carrying each workspace’s top model — estimates, on the same caveat as everywhere else.
Infra What the environment actually runs on: cloud inventory with live health, node/database/cache metrics, and the real cloud bill (month-to-date + forecast), read-only.
Budgets · Retention Spend envelopes that warn before the invoice, and the data-retention dials.

The escalation path, in order

  1. The collection’s Agents tab — is work moving?
  2. The stuck document’s history — where did it stop?
  3. The run trace — what did the worker actually do?
  4. Audit & Usage — who did what, and how much has been consumed?
  5. The chat and search audit tabs — what have users been experiencing?
  6. Evaluations — did a recent change move quality?

Learn this ladder once and most “something’s wrong” reports resolve without an engineer.

Where to go next

Prefer learning inside the product? The same academy lives in the platform's Learn menu — every screen links to the chapter that explains it.

See the platform live