Academy · Platform · Data

Collections & schema

In one line. What a Collection is in practice, how to create one, and how to give it a taxonomy — the fields the platform will extract from every document. You’ll be able to. Create a Processing collection from scratch and define its fields, ready for documents to flow in. Where this lives. Studio ▸ Data ▸ Collections

Why it matters

A Collection is the operational home for one type of record — “Vendor Invoices”, “HR Policies”, “Support Tickets”. Everything downstream hangs off it: documents land in a collection, an agent works it, your end-users read it. And before any of that, you give it a taxonomy — the list of fields the platform pulls out of each document. The taxonomy is the single most important thing you author here, because your fields drive extraction: the platform’s extraction agents read your field list to know what to pull out. No field, no extracted value.

Vocabulary. Collection is the product word — the container your documents and records live in. A Taxonomy is a named group of fields the platform extracts from a document. On screen one field is called a Label — that is the word on the buttons (+ Label, Label Name), and your workspace can rename it. This page says “field” for the idea and quotes the on-screen word for the controls. Keep the Glossary open.

Two things define a collection:

  • Purpose — why it exists: Processing (decide on each document — review stages, field extraction, an agent acts on each one), Knowledge (a passive corpus, auto-indexed so agents can read and cite it), or Structured (typed rows rather than documents). You choose it once, on the create form, and the three have a dedicated page: Collection types.
  • Who works it: most collections are worked by an agent. A human collection is a people-run work queue instead — the Inbox is the one you get without setting anything up.

Availability. Knowledge collections — the auto-indexing behaviour and the Index Health tab — are on by default. An administrator can switch them off for a workspace, in which case the Knowledge option does not appear when you create a collection. If you don’t see it, ask your administrator. A fourth purpose, Dataset (Evaluation), appears only where evaluations are switched on. Processing collections are always available, and that is what the walkthrough below builds.

The Collections workspace

Menu path: Studio ▸ Data ▸ Collections.

Collections is the first entry in Data, one of the four pillars across the top of Studio (Agents, Data, Experience, Governance). The Data pane is one flat list with no group headings, ordered the way data actually flows — define it, bring it in, look at it, send it out: Collections, Ingestion, Input Form, Drive, Datasheets, Views, Export.

This is a master-detail surface: a list of every collection on the left, and a detail pane on the right that shows the selected collection’s tabs.

  • Left rail lists collections by name. Parent/child collections are shown as indented nesting (a child sits under its parent, with a coloured left border by depth). A star marks the Primary/Default collection — the project’s default ingestion target. Click a row to load it on the right.
  • Right pane shows the selected collection’s name, a row-action menu, and the tab strip.

What every control does

Control What it does Notes
+ Collection (rail footer) Opens the Create Collection form Sits at the foot of the rail even when the list is empty. If you don’t see it, ask your administrator.
Row ⋮ ▸ Delete Deletes the collection (confirm) Warns by name if the collection has child collections.
+ Taxonomy (pane header) Opens Create Taxonomy scoped to the selected collection A shortcut to add a taxonomy without opening the Taxonomy tab first.
Empty state “No Collection found in the project” + + Collection Shown before you create your first.

Tip. An “Edit in Studio” link from elsewhere in the product opens Collections with the right collection already selected — you don’t need to hunt for it in the list.

The right-pane tabs — what each one is and where it’s covered

Every collection has seven tabs. A Knowledge collection adds Index Health; a Structured collection adds Schema. Each tab is a full page in its own right; here they open inside the collection.

Tab What it’s for Covered in
General The collection’s settings — this is also the create/edit form This page, below
Taxonomy Define the fields to extract This page
Lifecycle The review stages documents move through Taxonomy, lifecycle, tags & events
Tags The project’s free-form tag set applied to records Taxonomy, lifecycle, tags & events
Events Webhooks / emails / Slack fired on document or stage events Taxonomy, lifecycle, tags & events
Ingestion Where this collection’s documents come from Ingestion & connectors
Agents Assign the agent that works this collection, and watch it run Your first agent
Index Health Indexing health, per-cause failure groups, retries and reindex — Knowledge collections only Search, indexes & vectorization
Schema The typed columns of a table of records — Structured collections only Collection types

These are tabs inside a collection — you will not find Taxonomy, Lifecycle, Tags or Events in the Studio menu.

Two ways an agent meets a collection. An agent can work it — the collection’s intake agent runs on every document that lands, and you assign it on the Agents tab. Or an agent can read it — open the agent at Studio ▸ Agents ▸ Agents and attach the collection on the agent’s Knowledge tab. Different job, different screen; don’t confuse the two.

Creating a collection — the General tab

Click + Collection (or open the General tab of an existing one). The same form serves as both Create and Edit, and as the inline General tab. Save is disabled while the form is invalid.

You only need three decisions to get a working Processing collection — Purpose, name, description. The rest are optional accordions with sensible defaults.

The key settings groups (summary)

Group What you set Required?
Purpose Why the collection exists — Processing, Knowledge or Structured (see Collection types). Fixed once created — on edit it shows as a read-only chip. Pick one
Knowledge settings Search model (the model used to index text — required; the first available model is preselected on create and Save is blocked without one), chunk size (default 2000), chunk overlap (default 10). An Advanced ingestion settings foldout adds Index location (Shared vs Project index), page/document summaries, and OCR for scans. Expanded by default on create. Only when Purpose = Knowledge. Pick the Search model
Collection Name The display name Yes
Description What it holds Yes
Allowed Ingestion Types Whitelist which source types may feed this collection (Upload, Cloud, connectors, Form Templates) Optional
Default View A starting column view (only if any exist) Optional
Parent Collection Make this a child of another collection — reveals field-mapping accordions. Disabled once set (no re-parenting). Optional
Enrich / Dashboards / Independent Queue Advanced accordions (enrich triggers, per-collection dashboards, dedicated queue) Optional
Set as Primary Make this the project’s default ingestion target (exactly one per project) Optional

The full field-by-field table — every accordion, every child-mapping row, every default — is in the Schema field reference.

Watch out — Purpose is permanent. You choose it once, at create time. After that it shows as a read-only chip and cannot be changed. Pick deliberately.

The create form defines the collection’s shape and rules; you pick the agent that works it on the Agents tab (Your first agent).

The Taxonomy tab — defining what to extract

This is the heart of the page. Menu path: Studio ▸ Data ▸ Collections, select a collection, then Taxonomy.

How your fields drive extraction

When a document arrives, the platform’s extraction agents look at the collection’s taxonomy and extract a value for each field. A field called invoice_total (with a currency validation) tells the agent “find the total and pull it out as money.” Remove the field and that value is simply never captured. So authoring the taxonomy is authoring the extraction. (For a child collection, a parent’s extracted fields can auto-derive the child’s records via field mapping — see Taxonomy, lifecycle, tags & events.)

A taxonomy is a small tree:

  • Taxonomy — a named group of fields the platform extracts from the document
    • Field — one extractable value: a name plus behaviour
      • child field — a column, when the parent field is a table/record

Layout — master-detail

┌──────────────────────────────┐ ┌──────────────────────────────────────────────┐
│ Taxonomy   [type ▾] [search] │ │  Invoice Header        [+ Label]   [⋮ Edit]  │
│ ──────────────────────────── │ │  Labels | Details                            │
│  * Invoice Header (primary)  │ │  ─────────                                   │
│    Invoice Line Items        │ │   invoice_number      text                   │
│                              │ │   invoice_date        date                   │
│                              │ │   total_amount        currency               │
│  ──────────────────────────  │ │   line_items          ▸ table (3 child fields)│
│  [ Taxonomy ]  [ Import ]    │ │                                              │
└──────────────────────────────┘ └──────────────────────────────────────────────┘
  • Left list — every taxonomy (Name + Description); a star marks the Primary one. A type filter (All · Custom · Default · Imported) and a search icon narrow the list.
  • Footer — Taxonomy (create a new one) and Import (import one). Empty state: “No Taxonomies found. Create new Taxonomy.”
  • Right pane (a taxonomy selected) — + Label (top-right, on the Labels sub-tab) and an action menu with Edit. Two sub-tabs: Labels (a hierarchical field tree) and Details (the taxonomy’s own settings).
Control What it does
Taxonomy (footer) Opens Create Taxonomy — defines a new field group
Import (footer) Opens the import panel and brings a taxonomy in from a file
type ▾ filter Filters the list: All / Custom / Default / Imported
+ Label Opens the field editor — adds a field to the selected taxonomy
⋮ ▸ Edit Edits the selected taxonomy

Adding a taxonomy (Create Taxonomy)

A taxonomy groups the fields and is the thing that can be trained later. To create one, click Taxonomy and fill:

Field Notes
Taxonomy Name (required) Non-empty; duplicate names rejected
Description (required)
Type (required) What it classifies or extracts — for whole-document field extraction pick Document Classification
Entity (required) Which collection it belongs to (pre-filled to the one you’re on)

The full list of taxonomy Types (Document Classification, Table Classification, Image Annotation and more) is in the Schema field reference.

Adding a field — the field editor

Select a taxonomy, then + Label. The editor has two tabs — one to add a field by hand, one to import many at once.

  • Label Name and Description are the only required inputs — and for most fields they’re all you need.
  • Show Advanced Settings reveals six accordions (summarised below).
  • Add saves the field and keeps the dialog open (handy for entering several in a row); Add and exit saves and closes.

Field types — how a field becomes text, number, currency, date, boolean or a table

There is no single “type” dropdown. A field’s shape comes from a combination of a validation rule (in View Config) and Record Config. In practice you produce these shapes:

Shape How you get it
Text A field with no special validation (the default)
Number A numeric validation rule (View Config ▸ Validation)
Currency / amount Numeric validation, optionally Enable Total Field
Date A date validation rule
Boolean / yes-no A validation rule limiting to two values
Single-select / enum A validation rule listing allowed values
Table / line-items Is Record (Record Config) + child fields as columns
Masked / sensitive Mask Values toggle (Toggles/Features)

Tip. 90% of fields are “a field + a validation rule.” The one shape that’s different is a table/line-items field: turn on Is Record, then add child fields — each child becomes a column.

The exhaustive enumeration of every shape and every validation option is in the Schema field reference.

The six advanced accordions (summary)

You’ll rarely need these for a first build — here’s what each is for. The full control-by-control list is in the Schema field reference.

Accordion In one line
Table Styling Header/body background and text colours for a table field
Contextual Insights Real-time derivation and aggregation widgets
Record Config Make the field a table (Is Record), define columns, grouping, widths
Toggles/Features Mask Values, Expand Values, Enrich
Lookup Config Derive/refresh a value from another field or source
View Config Validation rules (this is where number/date/boolean/enum live), sort, totals

Bulk import of fields

The import tab takes .txt, .xlsx, .tsv or .csv. Parent/child nesting is expressed by indentation (text) or columns (spreadsheet), so you can stand up a whole table-with-columns in one go, then Upload & Create.

Tip. For a large taxonomy, build the field list in a spreadsheet first (one column per nesting level) and import it — far faster than clicking + Label dozens of times.

Behaviours to know

  • Required inputs: taxonomy Name + Description; Label Name + Description. Save/Submit/Add stay disabled until they’re filled.
  • Duplicate names are rejected at the taxonomy level.

Try it yourself

Build a Vendor Invoices Processing collection with four fields. (~5 minutes.)

  1. Go to Studio ▸ Data ▸ Collections. Click + Collection.
  2. Purpose: leave/select Processing. Collection Name: Vendor Invoices. Description: Incoming supplier invoices for approval. Leave everything else default. Click Save.
  3. Your collection appears in the left rail. Select it, open the Taxonomy tab.
  4. Click Taxonomy. Name: Invoice Header, Description: Top-level invoice fields, Type: Document Classification, Entity: Vendor Invoices. Submit.
  5. Select Invoice Header, click + Label. Add these four, using Add between each and Add and exit on the last:
    • invoice_number — Description: The supplier's invoice number. (Plain text — no advanced settings.)
    • invoice_date — open View Config, add a date validation rule.
    • vendor_name — plain text.
    • total_amount — open View Config, add a numeric validation rule (currency).
  6. Confirm all four show in the Labels tree under Invoice Header.

You now have a Processing collection whose extraction agent knows to pull four fields from every invoice. In Ingestion & connectors you’ll feed it documents; in Your first agent you’ll assign the agent that extracts them. The full end-to-end build is Build: Invoice settlement.

Bonus. Add a fifth field line_items as a table: in the editor open Record Config, turn on Is Record, then add child fields description, quantity, unit_price as its columns.

Recap

  • A Collection is the typed home for one kind of record, at Studio ▸ Data ▸ Collections — a master-detail workspace (collection rail + tabbed detail pane).
  • Create one from + Collection / the General tab: pick a Purpose — Processing, Knowledge or Structured (fixed once created) — name it, describe it. Everything else is optional accordions with defaults.
  • The detail pane has seven tabs — General + Taxonomy (this page), Lifecycle / Tags / Events, Ingestion, Agents — plus Index Health on a Knowledge collection and Schema on a Structured one. None of them appear in the Studio menu; they live inside the collection.
  • The Taxonomy tab defines field groups and the fields in them. Your fields drive extraction — the agent extracts exactly the fields you define.
  • A field’s type comes from its validation rule + Record Config, not a single dropdown. Tables use Is Record + child fields. Bulk import stands up a big taxonomy fast.

Where to go next

Prefer learning inside the product? The same academy lives in the platform's Learn menu — every screen links to the chapter that explains it.

See the platform live