Academy · Reference

Schema field reference

In one line. Every field on the collection form, the Taxonomy dialog and the Label editor, in one place. You’ll be able to. Fill in all three without guessing what a field wants — including the six advanced accordions and the bulk label import. Where this lives. Studio > Data > Collections — the + Collection button under the collection list opens the create form, and the Taxonomy tab on any collection is where Learners and Labels are built.

The concepts and the narrative walkthrough live in Collections & schema; open this page when you need the full list.

Vocabulary reminder. A Collection is a typed home for one kind of record. A Learner is the trainable extractor that owns a set of fields (the product also calls it a Taxonomy or a Data Capture group). A Label is one extractable field (some screens call it a Field). A Schema is “a Learner plus its Labels” — it is not a separate object.

Create-collection / General-tab form

The same form appears in the Create Collection dialog and on a collection’s General tab. Top to bottom:

# Field Type Required Notes
1 Purpose radio (create-only) — What kind of collection this is: Knowledge (a corpus agents learn from — every document is indexed for retrieval), Processing (documents flow through stages and agents act on them), Structured (typed records in a table you can query, sort and edit), Dataset (Evaluation) (documents with verified expected outputs, used only by evaluations). On edit it becomes a read-only chip — the purpose cannot be changed after create. Not every workspace offers all four; ask your administrator if one is missing.
2 Knowledge settings accordion — Only when Purpose = Knowledge, and open by default while you create. Search model (required — this is what makes documents searchable, and there is no silent default), Chunk size (number, default 2000), Chunk overlap (number, default 10), plus Advanced ingestion settings: index location (one shared index, or a dedicated index for this collection), per-page summaries, a whole-document summary, and reading scanned PDFs.
3 Schema designer (create-only) — Only when Purpose = Structured. Declare the fields — each becomes a typed column you can query and a label agents can read and write. After create, edit them on the collection’s Schema tab.
4 Collection Name text yes The display name.
5 Description textarea yes What this collection holds.
6 Allowed Ingestion Types multi-select (grouped) — Restricts which source types may feed this collection — Upload, cloud and connector kinds, and any Form Templates you have built.
7 Default View select — Appears only if column-view templates exist. None or a named column view.
8 Parent Collection select — None or another collection → makes this a child collection. Disabled once set (no re-parenting). Reveals the mapping accordions below. Only available if you can create child collections — ask your administrator if the field is missing.
9 Enrich Collection Configuration accordion — Enrich Xflow (select: None + the project’s flows); Enrich Labels (+ rows: Learner + Label); Flow Trigger Conditional (+ rows: Learner + Label + Value) — fires the enrich flow only when the label equals the value.
10 Collection Dashboards accordion — Per-collection dashboard configuration.
11 Independent Queue accordion (edit-only) — If not configured → Create IQ Request link; else shows status (Enabled / Disabled / Request Pending / Request Denied) with Enable/Disable buttons.
12 Set as Primary slide-toggle — Exactly one Primary collection per project — the default ingestion target.

The form is committed with Save; Close leaves without saving.

Where the intake agent is picked. Not here. Pick the agent that processes incoming documents on the collection’s Agents tab.

Structured collections. A collection that holds typed rows rather than documents keeps its columns on its own Schema tab. This page covers the Taxonomy side — Learners and Labels — which is what document collections use.

Child-collection mapping (revealed when a Parent Collection is chosen)

  • Dependent Configuration ▸ Primary Mapping — a Source (parent Learner + parent Label) => Target (this collection’s Learner + Label) row. Auto-populated and disabled once the collection is parented. This wires how a parent’s extracted Label becomes a child record.
  • Additional Label Mappings (sub-accordion, + adds rows) — more Source => Target Learner/Label pairs; each row deletable.

Create-Taxonomy (Learner) dialog

Opened from the Taxonomy button under the Learner list (the same button appears in the empty state when a collection has none yet).

Field Type Required Notes
Taxonomy Name text yes Non-empty; duplicate names are rejected with “Taxonomy name already exists”.
Description text yes
Type select yes What this Learner classifies or extracts — see below.
Collection select yes Which collection this Learner belongs to.

Submit stays disabled until every required field is filled.

Type — what the picker offers

Document Classification · Page Classification · Section Classification · Table Classification · Table Relation Extraction · Row Classification · Column Identification · Cell Identification · Relevant Block Identification — plus a few specialist image and cell types.

Your workspace may show a shorter list: an administrator can hide the types a project does not need.

Tip. For most extraction work you want Document Classification — one Learner per document type — and let its Labels do the field work.

Create-Label — the field editor

The dialog has two tabs: Add label and Import labels.

Core fields

Field Type Required Notes
Label Name text yes The field’s name.
Description text yes What this field captures (guides the extraction agent).

A Show Advanced Settings disclosure reveals six accordions (all optional). Footer: Add (saves, keeps the dialog open for the next field) and Add and exit (saves and closes).

Field “type” — how a Label becomes text vs number vs table

There is no single “data type” dropdown the way a spreadsheet has one. A Label’s behaviour is shaped by a combination of its validation rule (in View Config), its Record Config (whether it is a record / table), and Toggles. The common shapes a builder produces:

You want… Configure it as Where
Plain text a Label with no special validation default
Number Label + a numeric validation rule View Config ▸ Validation
Currency / amount numeric validation + (optionally) Enable Total Field View Config
Date a date validation rule View Config ▸ Validation
Boolean / yes-no a validation rule constraining to two values View Config ▸ Validation
Single-select / enum a validation rule listing allowed values View Config ▸ Validation
Table / line-items turn on Is Record and add child Labels as columns Record Config
Masked / sensitive Mask Values toggle Toggles
Derived / looked-up Lookup Config (Meta Label, Dependent Label Auto Update) Lookup Config

Tip. For a builder starting out, 90% of fields are “a Label + a validation rule”. Tables (line-items) are the one shape that needs Is Record plus child Labels. Everything in the six accordions below is refinement on top of those two ideas.

The six advanced accordions — every control

1. Table Styling — visual styling for table/record output. Appears only for a child Label or once Is Record is on:

  • Header Background, Header Text, Body Background, Body Text (each a colour-picker + hex field).

2. Contextual Insights — relationships and derived widgets:

  • Is Dependent Collection document (toggle)
  • Enable RTD operation for Document list page (toggle) — keeps the value derived in real time on the document list
  • Taxonomy + Label + Aggregation Type picker — what to aggregate and how
  • A label filter for the widget

3. Record Config — table / line-item behaviour:

  • Is Record (toggle) — makes this Label a table; its child Labels become columns
  • Freeze Rows
  • Label Grouping — appears once the field carries a grouping validation rule
  • Group By — how rows are grouped on the page
  • Minimum Column Width (in px)
  • Child-only toggles: Duplicate Label, Add Default Column, Unique Label for a record, Include Top Extraction for Record, Include Segment Grouping, Enable Filter

4. Toggles/Features:

  • Mask Values (hide sensitive values in the UI)
  • Expand Label Values
  • Enrich Label

5. Lookup Config — derive/refresh values from another source:

  • Meta Label
  • Skip Automation
  • Disable Manual Refresh
  • Dependent Label Auto Update (+ Taxonomy / Label picker)
  • Consolidated Derivation
  • Show Default Recordview

6. View Config — display, ordering, and validation (the most-used accordion):

  • Top Extraction Count
  • Sort Using (e.g. by confidence)
  • Doc Meta Property
  • Is Data Entry
  • Enable Total Field (requires Is Record)
  • Add / Edit Validation rules — this is where number/date/boolean/enum shapes are defined
  • Auto Resolve On Edit
  • Ignore Errors On Stage Movement

Import labels (bulk)

  • Tab Import labels: drag-and-drop or Choose File. Accepts .txt, .xlsx, .tsv, .csv.
  • Shows a text example and an Excel/CSV format example — parent/child nesting is expressed by indentation (a tab or four spaces in a text file) or by columns (Label Name, Description, Child Label Name, Child Description…), so you can import a whole table-with-columns in one go. In a text file, | separates each name from its description.
  • Upload & Create creates every Label in one action.

Tip. Importing is the fastest way to stand up a large schema. Build the field list in a spreadsheet (one column per nesting level), then Import labels ▸ Upload & Create.

Taxonomy tab — Learner list controls

The Taxonomy tab is a master-detail surface:

  • Left list — every Learner with Name + Description; a star marks the Primary Learner. A type filter (All · Custom · Default · Imported) and a search icon narrow the list. Footer buttons Taxonomy (create a Learner) and Import (import a Learner).
  • Right pane (a Learner selected) — top-right + Label (Labels sub-tab only) and an action menu with Edit (edit the Learner). Two sub-tabs: Labels (a hierarchical label tree) and Details.

Prefer learning inside the product? The same academy lives in the platform's Learn menu — every screen links to the chapter that explains it.

See the platform live